system
The voice adjustment system addresses hearing difficulties in the elderly by digitizing, noise-reducing, and adjusting sound quality while providing visual aids, enhancing communication clarity and comfort.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-15
- Publication Date
- 2026-04-27
AI Technical Summary
As people age, their hearing deteriorates, making it difficult to hear conversations over the phone, especially in environments with background noise, leading to social isolation and communication challenges.
A voice adjustment system that digitizes voice signals, reduces background noise, adjusts volume and sound quality using AI, and displays important phrases to enhance understanding, tailored to individual hearing characteristics.
Enables clearer and more comfortable telephone conversations for the elderly by optimizing sound quality and providing visual aids, supporting independent living and reducing communication stress.
Smart Images

Figure 2026070290000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] As people age, their hearing deteriorates, and it often becomes difficult to hear conversations over the phone. As a result, important communication in daily life becomes difficult, which may lead to social isolation. In addition, in an environment with background noise, it becomes even more difficult to accurately understand the content of a conversation. Conventional technologies have not provided sufficient solutions to such problems, and thus there is a need for a technology that efficiently supports the voice communication of the elderly.
Means for Solving the Problems
[0005] This invention provides a voice adjustment system for the elderly. First, a voice input means digitizes the voice signal when a user converses and captures it as voice data. This voice data is transmitted to a server via a network and analyzed in real time by a voice analysis means. Using an artificial intelligence model, background noise is reduced from the voice, and the volume and sound quality are adjusted to be easily heard based on the analysis results. This adjusted voice is provided to the user through a voice output means. In addition, a text analysis means extracts important words and phrases and displays them on the user interface to help the user understand the conversation more accurately. Furthermore, this system can learn the user's auditory characteristics and individually optimize the voice adjustment. This helps the elderly to have comfortable telephone conversations.
[0006] "Voice input means" refers to a device and function for converting user speech into digital data and processing the signal.
[0007] "Voice analysis means" refers to a device and function that uses an artificial intelligence model to analyze voice data in real time and extract information from the voice.
[0008] "Adjustment means" refers to devices and functions that adjust the volume and sound quality of audio based on the results of audio analysis to make it easier for the user to hear.
[0009] "Audio output means" refers to a device and function used to provide a controlled audio to the user.
[0010] "Text analysis means" refers to a device and function for extracting important words or phrases from audio data or the results of its analysis and displaying them in text format.
[0011] "Display means" refers to devices and functions for visually presenting textual information or important phrases to a user through a user interface. [Brief explanation of the drawing]
[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0013] An example of an embodiment of the system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0014] First, the terms used in the following description will be explained.
[0015] In the following embodiments, a tagged processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0016] In the following embodiments, a tagged RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0017] In the following embodiments, a tagged storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0018] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0020] [First Embodiment]
[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0033] This invention is a system that supports voice communication for the elderly and consists of multiple components. Specifically, it comprises a terminal held by the user, a server that processes the voice, and a network connecting the two.
[0034] The user initiates a call using the device. The device has a voice input mechanism and captures the user's voice as soon as they begin speaking, converting it into digital data. The converted voice data is then sent to the server using a protocol optimized to maintain low latency.
[0035] Upon receiving this audio data, the server first analyzes the data in real time using an audio analysis tool. A generative AI model is utilized to extract information on important frequency bands and background noise from the audio. Based on the analysis results, the server adjusts the volume and sound quality to make it easier for elderly people to hear. This adjusted audio is then returned to the user's terminal through an audio output tool and played back as clear audio.
[0036] In addition, the server uses text analysis to convert the audio data into text. From the transcribed data, important keywords are identified and displayed on the terminal's display through the user interface. This allows for visual supplementation of parts that were difficult to understand through auditory means.
[0037] As a concrete example, consider a case where user A uses the device to make a call with family members. The family members' voices are captured by the device's microphone and immediately sent to the server. The server analyzes the audio in real time, adjusts the sound quality to match user A's hearing characteristics, and then returns the adjusted audio to the device. Simultaneously, the conversation is transcribed into text, and particularly important keywords are displayed on the screen, allowing user A to confirm the conversation content not only aurally but also visually.
[0038] This system makes telephone communication more enjoyable and stress-free for users with hearing impairments. This invention could be a significant support for elderly people in maintaining independent living.
[0039] The following describes the processing flow.
[0040] Step 1:
[0041] The device captures the user's voice conversation using a microphone. The audio signal is converted from analog to digital and processed as audio data.
[0042] Step 2:
[0043] The terminal transmits the converted audio data to the server over the network. The audio data is transferred using an efficient protocol to maintain low latency.
[0044] Step 3:
[0045] The server analyzes the received audio data. It applies a generative AI model to analyze the frequency characteristics and noise components of the audio in real time.
[0046] Step 4:
[0047] Based on the results of the audio analysis, the server optimizes the volume and sound quality to make it easier for the user to hear. It emphasizes specific frequency bands and reduces background noise.
[0048] Step 5:
[0049] The server re-encodes the optimized audio and sends it to the terminal. The adjusted audio is then supplied to the user via the audio output device, resulting in clear playback.
[0050] Step 6:
[0051] The server simultaneously converts the audio data into text. It uses speech recognition technology to identify important keywords and extract them as text.
[0052] Step 7:
[0053] The terminal receives text data sent from the server and displays it on the screen. Important keywords are presented in a visually recognizable format within the user interface.
[0054] Step 8:
[0055] Users can review the audio-text displayed on their device, supplementing any parts of the conversation that were difficult to understand. This process makes it easier to accurately understand the content of the conversation.
[0056] (Example 1)
[0057] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0058] Providing voice communication tailored to the hearing characteristics of the elderly is challenging, and conventional communication devices may not allow for sufficient comprehension of conversations. Furthermore, in environments with significant background noise, hearing difficulties increase, potentially leading to the loss of important information. This can result in users experiencing stress and discomfort during communication.
[0059] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0060] In this invention, the server includes an input device that collects voice information and converts it into digital data, an analysis device that uses an artificial intelligence model to perform real-time analysis, and an adjustment device that adjusts the volume and sound quality to convert it into easily audible voice. This enables clear voice communication tailored to individual hearing characteristics.
[0061] "Audio information" refers to human speech and sounds, which are signals collected for digital processing.
[0062] "Digital data" refers to information obtained by converting analog signals into a format that can be processed by computers and digital devices.
[0063] An "input device" is a hardware or software component that collects audio information and converts it into digital data.
[0064] An "artificial intelligence model" is an algorithm that learns patterns and features from large amounts of data and performs analysis and prediction.
[0065] "Real-time analysis" is a process that processes input data immediately and outputs results quickly.
[0066] An "analysis device" is a component used to process received digital data using an artificial intelligence model.
[0067] "Volume and sound quality adjustment" is the process of adapting the volume and tone of received audio to the individual characteristics of each device.
[0068] A "regulating device" is a device that has the function of optimizing volume and sound quality.
[0069] "Easy-to-listen audio" refers to audio signals that are easy to understand and less burdensome to the listener, adjusted according to the user's auditory characteristics.
[0070] An "output device" is a device or interface for providing a pre-tuned audio signal to the user.
[0071] "Key terms" are words or phrases in audio data that contain particularly important information.
[0072] An "analysis device" is a device or software that has the function of extracting key words from audio data.
[0073] A "display function" refers to an interface or device that visually shows the extracted main terms to the user.
[0074] This invention is a system that supports voice communication for the elderly, and its main components include an input device, an analysis device, an adjustment device, an output device, an analytical device, and a display function.
[0075] Input device:
[0076] The user uses a device to acquire voice information and converts it into digital data. This process utilizes a microphone and a digital signal processing chip.
[0077] Analysis equipment:
[0078] The server analyzes the received digital data in real time using an artificial intelligence model. This AI model performs noise reduction and frequency enhancement. For example, it is possible to set prompts such as "reduce ambient noise and highlight important conversations."
[0079] Adjustment device:
[0080] The server adjusts the volume and sound quality based on the analysis results. This adjustment process uses equalization and sound pressure optimization techniques to adapt to the user's auditory characteristics.
[0081] Output device:
[0082] The device plays pre-tuned audio transmitted from the server, ensuring the user receives clear sound. Playback utilizes the device's speakers or headphones.
[0083] Analytical device and display function:
[0084] The server converts the audio data into text and extracts key phrases. This process identifies important information and displays it on the terminal's display. This allows the user to visually confirm the content of the conversation.
[0085] As a concrete example, when user A uses the device to make a call with family members, the family members' voices are captured by the device's microphone and immediately transmitted to the server as digital data. The server analyzes the audio in real time, removes noise, and sends the newly adjusted audio back to the device. At the same time, important keywords are displayed as text on the device's screen to complement the conversation. This allows user A to communicate smoothly within the family using both their ears and eyes.
[0086] This system provides hearing aids tailored to the hearing characteristics of elderly individuals, supporting verbal communication in daily life. This allows users to enjoy a more secure and independent life.
[0087] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0088] Step 1:
[0089] The terminal acquires the user's voice information and converts it from analog to digital. The input is the user's voice, captured by the terminal's microphone. This voice signal is converted into digital data by a digital signal processing chip. The output is digital voice data, which is then used for further processing.
[0090] Step 2:
[0091] The terminal sends the converted digital audio data to the server using an optimized low-latency protocol. The input is the digital audio data from step 1. Based on the transmission protocol (e.g., WebSocket or UDP), the data is transported to the server, minimizing network latency. The output is the completion of the transfer of the audio data to the server.
[0092] Step 3:
[0093] The server analyzes received digital audio data in real time using a generating AI model. The input is digital audio data sent from the terminal. Based on the prompt "remove noise and emphasize important frequencies," the AI model analyzes the audio data, extracts important frequency bands, and reduces noise. The output is the analyzed audio information.
[0094] Step 4:
[0095] The server adjusts the volume and sound quality based on the analysis results. This process uses analyzed audio information from an AI model as input. Equalization and compression are used for adjustment to optimize the audio to suit the user's auditory characteristics. The output is the adjusted audio data.
[0096] Step 5:
[0097] The server sends the adjusted audio data to the terminal. The input is the adjusted audio data obtained in step 4. The data is sent to the terminal again using the optimized protocol, maintaining low latency. The output is the reception of the adjusted audio data on the terminal.
[0098] Step 6:
[0099] The terminal plays the pre-tuned audio received from the server. The input is the pre-tuned audio data sent from the server. Clear audio is provided to the user using the terminal's speaker or headphones. The output is audio playback in an audible format.
[0100] Step 7:
[0101] The server transcribes the audio data into text and extracts key keywords through analysis. The input is the audio data obtained in step 3. Using speech recognition technology, it generates transcribed data and extracts key phrases from the conversation. The output is the transcribed data and key phrases.
[0102] Step 8:
[0103] The terminal displays the extracted keywords on its screen. Input consists of transcribed data and key phrases from the server. The user interface complements auditory input by visually displaying the main points of the conversation. Output is a display of keywords as visual information on the screen.
[0104] (Application Example 1)
[0105] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0106] For the elderly, the insufficient accuracy of voice recognition when giving voice instructions for electronic transactions, and the inability to visually confirm the content of the voice instructions, are major obstacles. Furthermore, sound quality issues caused by background noise and other noises affect the accuracy of voice instructions, highlighting the need for effective solutions.
[0107] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0108] In this invention, the server includes voice acquisition means, voice analysis means, voice adjustment means, and electronic transaction support means. This makes it possible to instantly adjust the sound quality in response to voice instructions and provide voice feedback that is easy for the elderly to hear. Furthermore, by visually presenting important transaction details, it is possible to prevent misunderstandings of transactions based solely on voice instructions and to support the elderly in confidently conducting electronic transactions by voice.
[0109] A "voice acquisition means" is a device that captures a user's speech as an audio signal and converts it into audio information.
[0110] A "voice analysis device" is a device that has the function of analyzing acquired voice information in real time using a generation AI model and extracting important content.
[0111] A "sound adjustment means" is a device that, based on analyzed sound information, adjusts the volume and sound quality to match the user's auditory characteristics, converting it into easily understandable sound.
[0112] "Audio playback means" refers to a device for outputting adjusted audio to the user.
[0113] A "text analysis device" is a device that extracts important words and phrases from audio information and processes them as text information.
[0114] An "information presentation means" is a device that displays important keywords extracted by a text analysis means on a user interface.
[0115] An "electronic transaction support device" is a device that receives voice instructions from a user and performs the necessary information processing to facilitate electronic transactions.
[0116] A "display support device" is a device that has the function of displaying information so that transaction details can be visually confirmed.
[0117] This invention is a system that supports electronic transactions via voice commands for the elderly. The system consists of a user terminal, a server, and a network connecting them. When a user gives voice commands for an electronic transaction, the voice is captured by a microphone on the terminal and transmitted to the server as voice data.
[0118] The server uses the Google® Cloud Speech-to-Text API to perform speech analysis, converting the audio data into text. Furthermore, it extracts important keywords from the analyzed text information using a natural language processing (NLP) library and presents them visually to the user.
[0119] To reduce audio noise and adjust sound quality to suit the hearing characteristics of elderly individuals, a generative AI model using TENSORFLOW® further processes the audio data. The processed audio is returned to the device and played back in a format that is easy for the user to understand.
[0120] The terminal's display shows important information about electronic transactions conducted via voice commands, such as the trading partner and the amount. This display helps the user confirm the transaction.
[0121] For example, if a user gives a voice command such as, "I want to buy a 500 yen item at the nearby supermarket," the command is instantly converted into text, and the necessary information is displayed. Furthermore, the AI model analyzes the voice command and returns feedback with appropriately adjusted volume and sound quality. A confirmation prompt appears asking, "Do you want to proceed with purchasing the 500 yen item at the supermarket based on this voice command?"
[0122] An example of a prompt is, "If a user indicates their intention to make an electronic transaction via voice, explain in detail how to process this intention." This prompt is useful for confirming and understanding the steps involved in voice analysis and transaction processing using a generative AI model.
[0123] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0124] Step 1:
[0125] The user issues voice commands to the terminal. The terminal captures the voice signal using its built-in microphone and converts it into digital audio data. This data is then sent to the server.
[0126] Step 2:
[0127] The server converts the received digital audio data into text using the Google Cloud Speech-to-Text API. The input is audio data, and the output is text data. The server extracts the converted text data and passes it on to the next processing step.
[0128] Step 3:
[0129] The server analyzes text data using generative AI models and natural language processing (NLP) techniques. The analysis extracts and identifies key keywords. The input is text data, and the output is a list of key keywords. This makes important instructions regarding transactions easier to understand.
[0130] Step 4:
[0131] The server organizes the extracted keywords and constructs the information necessary to visualize the details of electronic transactions. This information is sent to the user's terminal and displayed on the screen. The input is a list of important keywords, and the output is transaction information for display.
[0132] Step 5:
[0133] The server analyzes the audio data using an AI model and adjusts the sound quality to suit the hearing characteristics of elderly individuals. This adjusted audio data is then sent to the terminal and played back. The input is the original audio data, and the output is optimized audio data. This allows the user to receive clear audio feedback.
[0134] Step 6:
[0135] The terminal simultaneously presents the user with transaction information for display and pre-recorded audio. The user can verify transaction details both visually and aurally. Input consists of transaction information for display and pre-recorded audio data, while the visual display and audio output constitute the output.
[0136] Step 7:
[0137] Once the user confirms the details and indicates their intention to execute the transaction, the details are sent back from the terminal to the server, and the electronic transaction is completed. The input is the confirmed transaction information, and the output is a transaction success message. This ensures that electronic transactions are conducted safely and smoothly.
[0138] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0139] This invention provides a voice communication support system for the elderly that incorporates an emotion engine that recognizes the user's emotions in addition to voice analysis. This system consists of a user terminal and a server, which are connected via a network.
[0140] The user makes a call using their device, and the audio is captured by a voice input device. The captured audio data is digitized and sent to a server. The server analyzes the audio data using a voice analysis device and understands the characteristics of the audio using a generative AI model. Based on this, the server adjusts the volume and sound quality to make it easy for the user to hear.
[0141] This system is equipped with an emotion engine, and the server analyzes the user's emotional state from their voice in real time. The emotion engine analyzes changes in intonation, tempo, and volume of the voice to infer the user's emotions. The resulting emotional state is then used for further optimization in voice adjustment.
[0142] For example, if the emotion engine determines that the user is stressed, the server adjusts the voice to a calmer tone and further reduces background noise. The adjusted voice is then sent back from the server to the terminal and delivered to the user as clear audio.
[0143] Furthermore, along with the audio data transcribed into text by the text analysis method, emotional information analyzed by the emotion engine is displayed on the user interface. Specifically, simple feedback corresponding to the emotional state is displayed on the screen, allowing the user to obtain information based on their own emotional state. For example, if the user's emotions are analyzed as being calm, a message such as "Relaxed" will be displayed.
[0144] This system will allow elderly individuals to enjoy more fulfilling communication not only through voice but also through emotional feedback. The introduction of an emotion engine will enable individualized adaptation that takes into account the user's emotions, which is expected to further improve the quality of communication.
[0145] The following describes the processing flow.
[0146] Step 1:
[0147] The device captures the user's call audio using a microphone and converts the analog signal into digital audio data. The audio data is then encoded into a format that allows for real-time processing.
[0148] Step 2:
[0149] The terminal sends the converted audio data to the server. The transmission is performed using a secure protocol to minimize latency.
[0150] Step 3:
[0151] When the server receives audio data, it first uses an audio analysis tool to analyze the audio frequency and volume. This analysis provides the basic data needed to adjust the sound quality to make it easier for the user to hear.
[0152] Step 4:
[0153] The server uses an emotion engine to analyze the emotional nuances contained in the speech. It analyzes the intonation, tempo, and volume fluctuations of the speech to infer the user's emotional state.
[0154] Step 5:
[0155] The server adjusts the volume and frequency of the audio based on the results of voice analysis and emotion analysis. The adjustments are adapted to the user's auditory characteristics and estimated emotional state.
[0156] Step 6:
[0157] The server sends back audio, adjusted to a sound quality that allows the user to relax, to the terminal via an audio output device. The terminal then plays the adjusted audio clearly.
[0158] Step 7:
[0159] The server converts the audio data into text, identifying and extracting important keywords. The extracted text information, along with emotional feedback, is sent to the terminal.
[0160] Step 8:
[0161] The device displays received text and sentiment information on its screen. Users can visually confirm key points of the conversation and receive feedback on their own emotional state.
[0162] Step 9:
[0163] By reviewing the displayed information and understanding parts that were difficult to hear or their own emotional state, users can facilitate smoother communication.
[0164] (Example 2)
[0165] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0166] When elderly people communicate with others via voice, the sound quality and volume may be insufficient, or the user's emotional state may not be understood, leading to communication that proceeds without proper understanding. This can hinder smooth communication and result in an unsatisfactory experience. Therefore, there is a need for technology that improves sound quality and provides feedback that responds to the user's emotions, thereby enriching the communication experience.
[0167] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0168] In this invention, the server includes means for receiving digitized speech and analyzing speech characteristics in real time using a generation AI model; a speech adjustment device that adjusts the volume and sound quality based on the analyzed speech characteristics and emotional state to convert it into an easily audible sound; and a text processing device that generates text from the speech and extracts important information. As a result, users can receive clearer and easier-to-understand speech and also receive feedback tailored to their emotional state.
[0169] "Elderly people" generally refers to people who are older, and in this invention, they are a group that requires special consideration when communication is primarily conducted via voice.
[0170] "Digitalization" is the process of sampling analog audio signals and converting them into digital signals, and it is a fundamental step that enables data processing for the entire system.
[0171] A "generative AI model" is a type of artificial intelligence used to analyze speech characteristics, designed to perform intelligible speech conversion and emotion analysis in real time.
[0172] "Sound characteristics" refer to the physical and sensory characteristics of an audio signal, and specifically include volume, sound quality, frequency spectrum, intonation, tempo, etc.
[0173] A "voice adjustment device" is a device that adjusts voice to an optimal state based on analyzed voice characteristics and emotional state, thereby allowing users to receive voices that are easier to hear.
[0174] A "text processing device" is a device that converts audio signals into text and extracts important information, making it useful for providing information to users.
[0175] "Feedback" is information provided based on the user's voice and emotional state, and is intended to help users understand their own expressions and facilitate smooth communication.
[0176] This system is designed to help elderly people communicate smoothly and comfortably using voice. The system mainly consists of terminals, servers, and various software components.
[0177] The user initiates a voice call using their device. The device has a built-in microphone that captures the user's voice as an analog signal. The captured audio is digitized by the device, sampled, and converted into a digital audio format. This digitized audio data is then transmitted to the server via the network.
[0178] The server processes the received audio data using a generative AI model. Specifically, it analyzes the characteristics of the audio in real time and performs adjustments such as noise reduction and volume adjustment. Furthermore, the server analyzes the intonation and tempo of the audio using its built-in emotion engine and infers the user's emotional state. The results of this emotion analysis are also fed back into the audio adjustment process, generating the optimal audio for the user.
[0179] The server also converts the audio data into text and extracts important information. This text, along with feedback on the emotional state, is displayed in the user interface, allowing the user to understand their own emotional state while communicating.
[0180] As a concrete example, here is an example of a prompt sentence to give to a generative AI model: "Based on the audio data the user is speaking, use the emotion engine to analyze the emotions and provide appropriate feedback."
[0181] This system is expected to improve the quality of communication among the elderly by enabling higher-quality voice communication and providing more comprehensive feedback based on emotional understanding.
[0182] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0183] Step 1:
[0184] The user initiates a call using the device. The device's microphone captures the user's speech and recognizes it as an analog audio signal. The device then samples and quantizes this input audio, converting it into a digital audio format. In this process, the analog signal is transformed into digital data. Digital audio data is generated and prepared for the next processing step.
[0185] Step 2:
[0186] The terminal transmits digitized audio data to the server via the network. The server uses the received audio data as input to run a generating AI model, which analyzes the audio characteristics in real time. Data calculations such as volume, sound quality, and frequency spectrum are performed, and parts that require noise reduction or volume adjustment are identified. The analysis results output the points that need adjustment and the adjusted audio characteristics.
[0187] Step 3:
[0188] The server works in conjunction with a generative AI model to activate the emotion engine. It analyzes changes in intonation, tempo, and volume from the speech to estimate the user's emotional state. The analyzed emotional state is detected as input, and the emotion is output as a state such as stress or relaxation. This information is used for subsequent speech adjustments.
[0189] Step 4:
[0190] The server activates an audio adjustment device based on the analyzed voice characteristics and emotional state. Specifically, it adjusts the voice to make it easier to hear by applying noise cancellation and volume adjustments. For example, if stress is detected, it processes the voice to calm it down. This process generates a clear voice, which is then prepared for transmission.
[0191] Step 5:
[0192] The server sends the adjusted audio data back to the terminal. The terminal provides the received audio to the user through an audio output device. Here, it is output as the audio the user actually hears. The server also transcribes the audio data into text, extracts important information, and supplies it to the user interface display as feedback information. This allows the user to access information including emotional feedback.
[0193] Step 6:
[0194] Users refer to text information and sentiment feedback displayed on the device's screen. This allows them to understand their own expressions, obtain the information necessary for communication, and use it for the next step. By reviewing the displayed feedback information, they can gain insights to communicate effectively.
[0195] (Application Example 2)
[0196] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0197] Improving the quality of voice communication is crucial for the elderly and those with limited communication abilities. However, in addition to ensuring clear voice quality, there is a lack of systems that can accurately grasp the emotional state of the other party and respond appropriately as needed. Furthermore, real-time emotion recognition is required to detect and respond to potential security risks arising from emotional shifts at an early stage. Solving these challenges is essential.
[0198] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0199] In this invention, the server includes an acoustic input means, an acoustic analysis means, an adjustment means, an emotion recognition means, and an alarm means. This makes it easy for the elderly to use and enables real-time detection and response to security risks based on emotions.
[0200] An "acoustic input device" is a device that acquires voice signals emitted by elderly people and converts them into voice information.
[0201] An "acoustic analysis device" is a device that analyzes received audio information in real time using a generation model to understand acoustic characteristics and emotional states.
[0202] A "adjustment device" is a device that adjusts the volume and sound quality based on the results of acoustic analysis, converting the sound into one that is easy for elderly people to hear.
[0203] "Audio output means" refers to a device for providing adjusted sound to the user.
[0204] A "language analysis device" is a device that has the function of extracting important words from audio information and displaying that information on the user's screen.
[0205] A "display device" is a device that visually shows analyzed emotional information and important words on the user's screen.
[0206] An "emotion recognition device" is a device that analyzes the characteristics of speech from speech information obtained by an acoustic analysis device to estimate an emotional state.
[0207] An "alarm device" is a device that has the function of detecting potential security risks and providing necessary notifications based on an estimated emotional state.
[0208] The system of this invention supports voice communication for the elderly and provides emotion-based adaptive feedback and security features. A detailed embodiment of this system is shown below.
[0209] First, the terminal uses an acoustic input device to acquire an audio signal from the user and converts it into digital data as audio information. This audio information is then sent to the server.
[0210] The server uses acoustic analysis tools and generative AI models to analyze speech information in real time. Specifically, the acoustic analysis engine analyzes changes in speech intonation, tempo, and volume, and estimates the user's emotional state through emotion recognition tools. As generative AI models, for example, Google Cloud's Speech-to-Text API and Emotion Recognition API can be used.
[0211] Next, the adjustment means adjusts the volume and sound quality based on the analyzed information to create a sound that is easy for the user to hear. This adjusted sound is then provided to the user by the sound output means.
[0212] Furthermore, the language analysis tool extracts important words from the audio information and displays them visually on the user's screen via the display tool. Simultaneously, emotional information is also displayed on the screen, allowing the user to understand their own emotional state.
[0213] The alarm system activates when the estimated emotional state indicates a certain level of security risk. Specifically, when the emotional state is determined to be stressful, it sends a notification to registered contacts using communication methods such as Twilio or SendGrid.
[0214] This system allows elderly individuals to safely enjoy communication not only through voice but also through emotional feedback. For example, when an elderly person feels lonely, their family is notified, prompting them to take care of them.
[0215] Examples of prompt messages include the following:
[0216] "Analyze the current emotional state and determine whether the elderly person's emotions are calm."
[0217] "Use the user's voice data to detect their stress levels and send alerts if necessary."
[0218] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0219] Step 1:
[0220] The terminal acquires the user's voice using an acoustic input device. The input is the user's voice signal, which the terminal converts into digital data. The converted digital audio data is then sent to the server.
[0221] Step 2:
[0222] The server uses speech analysis tools to analyze digital speech data through a generative AI model. The input is digital speech data, and the server analyzes changes in intonation, tempo, and volume. As a result of the analysis, acoustic characteristics and emotional states are extracted.
[0223] Step 3:
[0224] The server adjusts the volume and sound quality based on the analysis results using adjustment mechanisms. The input is acoustic characteristics, and the server adjusts the sound to be easily audible to the user, generating clear, adjusted audio data. The adjusted audio data is then transmitted to the terminal.
[0225] Step 4:
[0226] The terminal outputs adjusted audio to the user using an audio output means. The input is adjusted audio data, allowing the user to hear clear audio.
[0227] Step 5:
[0228] The server extracts important words from audio information using language analysis tools. The input is audio information, and the server performs language analysis to extract important words. Once important words are extracted, the data is transmitted to the terminal via a display device.
[0229] Step 6:
[0230] The terminal uses a display mechanism to show important words and estimated sentiment information on the user's screen. The input consists of extracted important words and sentiment information, allowing the user to visually confirm their emotional state.
[0231] Step 7:
[0232] The server utilizes emotion recognition mechanisms to detect potential security risks based on estimated emotional states. If the emotion indicates stress or danger, it uses that state as input to send a notification command to the alarm system.
[0233] Step 8:
[0234] The server uses alarm mechanisms to notify pre-configured emergency contacts via communication means as needed. The input is information about detected security risks, a notification is output, and relevant parties are prompted to take prompt action.
[0235] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0236] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0237] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0238] [Second Embodiment]
[0239] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0240] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0241] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0242] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0243] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0244] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0245] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0246] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0247] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0248] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0249] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0250] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0251] This invention is a system that supports voice communication for the elderly and consists of multiple components. Specifically, it comprises a terminal held by the user, a server that processes the voice, and a network connecting the two.
[0252] The user initiates a call using the device. The device has a voice input mechanism and captures the user's voice as soon as they begin speaking, converting it into digital data. The converted voice data is then sent to the server using a protocol optimized to maintain low latency.
[0253] Upon receiving this audio data, the server first analyzes the data in real time using an audio analysis tool. A generative AI model is utilized to extract information on important frequency bands and background noise from the audio. Based on the analysis results, the server adjusts the volume and sound quality to make it easier for elderly people to hear. This adjusted audio is then returned to the user's terminal through an audio output tool and played back as clear audio.
[0254] In addition, the server uses text analysis to convert the audio data into text. From the transcribed data, important keywords are identified and displayed on the terminal's display through the user interface. This allows for visual supplementation of parts that were difficult to understand through auditory means.
[0255] As a concrete example, consider a case where user A uses the device to make a call with family members. The family members' voices are captured by the device's microphone and immediately sent to the server. The server analyzes the audio in real time, adjusts the sound quality to match user A's hearing characteristics, and then returns the adjusted audio to the device. Simultaneously, the conversation is transcribed into text, and particularly important keywords are displayed on the screen, allowing user A to confirm the conversation content not only aurally but also visually.
[0256] This system makes telephone communication more enjoyable and stress-free for users with hearing impairments. This invention could be a significant support for elderly people in maintaining independent living.
[0257] The following describes the processing flow.
[0258] Step 1:
[0259] The device captures the user's voice conversation using a microphone. The audio signal is converted from analog to digital and processed as audio data.
[0260] Step 2:
[0261] The terminal transmits the converted audio data to the server over the network. The audio data is transferred using an efficient protocol to maintain low latency.
[0262] Step 3:
[0263] The server analyzes the received audio data. It applies a generative AI model to analyze the frequency characteristics and noise components of the audio in real time.
[0264] Step 4:
[0265] Based on the results of the audio analysis, the server optimizes the volume and sound quality to make it easier for the user to hear. It emphasizes specific frequency bands and reduces background noise.
[0266] Step 5:
[0267] The server re-encodes the optimized audio and sends it to the terminal. The adjusted audio is then supplied to the user via the audio output device, resulting in clear playback.
[0268] Step 6:
[0269] The server simultaneously converts the audio data into text. It uses speech recognition technology to identify important keywords and extract them as text.
[0270] Step 7:
[0271] The terminal receives text data sent from the server and displays it on the screen. Important keywords are presented in a visually recognizable format within the user interface.
[0272] Step 8:
[0273] Users can review the audio-text displayed on their device, supplementing any parts of the conversation that were difficult to understand. This process makes it easier to accurately understand the content of the conversation.
[0274] (Example 1)
[0275] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0276] Providing voice communication tailored to the hearing characteristics of the elderly is challenging, and conventional communication devices may not allow for sufficient comprehension of conversations. Furthermore, in environments with significant background noise, hearing difficulties increase, potentially leading to the loss of important information. This can result in users experiencing stress and discomfort during communication.
[0277] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0278] In this invention, the server includes an input device that collects voice information and converts it into digital data, an analysis device that uses an artificial intelligence model to perform real-time analysis, and an adjustment device that adjusts the volume and sound quality to convert it into easily audible voice. This enables clear voice communication tailored to individual hearing characteristics.
[0279] "Audio information" refers to human speech and sounds, which are signals collected for digital processing.
[0280] "Digital data" refers to information obtained by converting analog signals into a format that can be processed by computers and digital devices.
[0281] An "input device" is a hardware or software component that collects audio information and converts it into digital data.
[0282] An "artificial intelligence model" is an algorithm that learns patterns and features from large amounts of data and performs analysis and prediction.
[0283] "Real-time analysis" is a process that immediately processes the input data and quickly outputs the results.
[0284] "Analysis device" is a component for processing the received digital data using an artificial intelligence model.
[0285] "Adjustment of volume and sound quality" is a process of adapting the received voice's loudness and timbre to individual characteristics.
[0286] "Adjustment device" is a device having a function for optimizing volume and sound quality.
[0287] "Easily audible voice" is an easily understandable and less burdensome voice signal adjusted according to the user's auditory characteristics.
[0288] "Output device" is a device or interface for providing the adjusted voice signal to the user.
[0289] "Main phrases" are words or phrases containing particularly important information in the voice data.
[0290] "Analysis device" is a device for extracting main phrases from voice data or software having its function.
[0291] "Display function" is an interface or device for visually showing the extracted main phrases to the user.
[0292] This invention is a system for assisting the voice communication of the elderly, and includes an input device, an analysis device, an adjustment device, an output device, an analysis device, and a display function as main components.
[0293] Input device:
[0294] The user uses a device to acquire voice information and converts it into digital data. This process utilizes a microphone and a digital signal processing chip.
[0295] Analysis equipment:
[0296] The server analyzes the received digital data in real time using an artificial intelligence model. This AI model performs noise reduction and frequency enhancement. For example, it is possible to set prompts such as "reduce ambient noise and highlight important conversations."
[0297] Adjustment device:
[0298] The server adjusts the volume and sound quality based on the analysis results. This adjustment process uses equalization and sound pressure optimization techniques to adapt to the user's auditory characteristics.
[0299] Output device:
[0300] The device plays pre-tuned audio transmitted from the server, ensuring the user receives clear sound. Playback utilizes the device's speakers or headphones.
[0301] Analytical device and display function:
[0302] The server converts the audio data into text and extracts key phrases. This process identifies important information and displays it on the terminal's display. This allows the user to visually confirm the content of the conversation.
[0303] As a specific example, when user A uses the terminal to talk to their family, the voices of the family are captured by the microphone of the terminal and immediately transmitted to the server as digital data. The server analyzes the voice in real time, removes noise, and sends the newly adjusted voice back to the terminal. At the same time, important keywords are displayed as text on the display of the terminal to complement the conversation. As a result, user A can smoothly communicate within the family using their ears and eyes.
[0304] This system provides a hearing-aid function according to the hearing characteristics of the elderly and supports voice communication in daily life. As a result, users can enjoy a more secure and independent life.
[0305] The flow of the specific process in Example 1 will be described using FIG. 11.
[0306] Step 1:
[0307] The terminal acquires the user's voice information and converts it from analog to digital. As input, the user's voice is captured by the microphone of the terminal. This voice signal is converted into digital data by a digital signal processing chip. The output is digital-form voice data, and this data is then used for subsequent processing.
[0308] Step 2:
[0309] The terminal transmits the converted digital voice data to the server using an optimized low-latency protocol. The input is the digital voice data from Step 1. Based on the transmission protocol (e.g., WebSocket or UDP), the data is transported to the server, minimizing network latency. The output is the completion of the transfer of voice data to the server.
[0310] Step 3:
[0311] The server analyzes received digital audio data in real time using a generating AI model. The input is digital audio data sent from the terminal. Based on the prompt "remove noise and emphasize important frequencies," the AI model analyzes the audio data, extracts important frequency bands, and reduces noise. The output is the analyzed audio information.
[0312] Step 4:
[0313] The server adjusts the volume and sound quality based on the analysis results. This process uses analyzed audio information from an AI model as input. Equalization and compression are used for adjustment to optimize the audio to suit the user's auditory characteristics. The output is the adjusted audio data.
[0314] Step 5:
[0315] The server sends the adjusted audio data to the terminal. The input is the adjusted audio data obtained in step 4. The data is sent to the terminal again using the optimized protocol, maintaining low latency. The output is the reception of the adjusted audio data on the terminal.
[0316] Step 6:
[0317] The terminal plays the pre-tuned audio received from the server. The input is the pre-tuned audio data sent from the server. Clear audio is provided to the user using the terminal's speaker or headphones. The output is audio playback in an audible format.
[0318] Step 7:
[0319] The server transcribes the audio data into text and extracts key keywords through analysis. The input is the audio data obtained in step 3. Using speech recognition technology, it generates transcribed data and extracts key phrases from the conversation. The output is the transcribed data and key phrases.
[0320] Step 8:
[0321] The terminal displays the extracted keywords on its screen. Input consists of transcribed data and key phrases from the server. The user interface complements auditory input by visually displaying the main points of the conversation. Output is a display of keywords as visual information on the screen.
[0322] (Application Example 1)
[0323] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0324] For the elderly, the insufficient accuracy of voice recognition when giving voice instructions for electronic transactions, and the inability to visually confirm the content of the voice instructions, are major obstacles. Furthermore, sound quality issues caused by background noise and other noises affect the accuracy of voice instructions, highlighting the need for effective solutions.
[0325] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0326] In this invention, the server includes voice acquisition means, voice analysis means, voice adjustment means, and electronic transaction support means. This makes it possible to instantly adjust the sound quality in response to voice instructions and provide voice feedback that is easy for the elderly to hear. Furthermore, by visually presenting important transaction details, it is possible to prevent misunderstandings of transactions based solely on voice instructions and to support the elderly in confidently conducting electronic transactions by voice.
[0327] A "voice acquisition means" is a device that captures a user's speech as an audio signal and converts it into audio information.
[0328] A "voice analysis device" is a device that has the function of analyzing acquired voice information in real time using a generation AI model and extracting important content.
[0329] A "sound adjustment means" is a device that, based on analyzed sound information, adjusts the volume and sound quality to match the user's auditory characteristics, converting it into easily understandable sound.
[0330] "Audio playback means" refers to a device for outputting adjusted audio to the user.
[0331] A "text analysis device" is a device that extracts important words and phrases from audio information and processes them as text information.
[0332] An "information presentation means" is a device that displays important keywords extracted by a text analysis means on a user interface.
[0333] An "electronic transaction support device" is a device that receives voice instructions from a user and performs the necessary information processing to facilitate electronic transactions.
[0334] A "display support device" is a device that has the function of displaying information so that transaction details can be visually confirmed.
[0335] This invention is a system that supports electronic transactions via voice commands for the elderly. The system consists of a user terminal, a server, and a network connecting them. When a user gives voice commands for an electronic transaction, the voice is captured by a microphone on the terminal and transmitted to the server as voice data.
[0336] The server uses the Google Cloud Speech-to-Text API to perform speech analysis, converting the audio data into text. Furthermore, it extracts important keywords from the analyzed text using a natural language processing (NLP) library and presents them visually to the user.
[0337] To reduce audio noise and adjust sound quality to suit the hearing characteristics of elderly individuals, a generative AI model using TensorFlow further processes the audio data. The processed audio is returned to the device and played back in a format that is easy for the user to understand.
[0338] The terminal's display shows important information about electronic transactions conducted via voice commands, such as the trading partner and the amount. This display helps the user confirm the transaction.
[0339] For example, if a user gives a voice command such as, "I want to buy a 500 yen item at the nearby supermarket," the command is instantly converted into text, and the necessary information is displayed. Furthermore, the AI model analyzes the voice command and returns feedback with appropriately adjusted volume and sound quality. A confirmation prompt appears asking, "Do you want to proceed with purchasing the 500 yen item at the supermarket based on this voice command?"
[0340] An example of a prompt is, "If a user indicates their intention to make an electronic transaction via voice, explain in detail how to process this intention." This prompt is useful for confirming and understanding the steps involved in voice analysis and transaction processing using a generative AI model.
[0341] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0342] Step 1:
[0343] The user issues voice commands to the terminal. The terminal captures the voice signal using its built-in microphone and converts it into digital audio data. This data is then sent to the server.
[0344] Step 2:
[0345] The server converts the received digital audio data into text using the Google Cloud Speech-to-Text API. The input is audio data, and the output is text data. The server extracts the converted text data and passes it on to the next processing step.
[0346] Step 3:
[0347] The server analyzes text data using generative AI models and natural language processing (NLP) techniques. The analysis extracts and identifies key keywords. The input is text data, and the output is a list of key keywords. This makes important instructions regarding transactions easier to understand.
[0348] Step 4:
[0349] The server organizes the extracted keywords and constructs the information necessary to visualize the details of electronic transactions. This information is sent to the user's terminal and displayed on the screen. The input is a list of important keywords, and the output is transaction information for display.
[0350] Step 5:
[0351] The server analyzes the audio data using an AI model and adjusts the sound quality to suit the hearing characteristics of elderly individuals. This adjusted audio data is then sent to the terminal and played back. The input is the original audio data, and the output is optimized audio data. This allows the user to receive clear audio feedback.
[0352] Step 6:
[0353] The terminal simultaneously presents the user with transaction information for display and pre-recorded audio. The user can verify transaction details both visually and aurally. Input consists of transaction information for display and pre-recorded audio data, while the visual display and audio output constitute the output.
[0354] Step 7:
[0355] Once the user confirms the details and indicates their intention to execute the transaction, the details are sent back from the terminal to the server, and the electronic transaction is completed. The input is the confirmed transaction information, and the output is a transaction success message. This ensures that electronic transactions are conducted safely and smoothly.
[0356] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0357] This invention provides a voice communication support system for the elderly that incorporates an emotion engine that recognizes the user's emotions in addition to voice analysis. This system consists of a user terminal and a server, which are connected via a network.
[0358] The user makes a call using their device, and the audio is captured by a voice input device. The captured audio data is digitized and sent to a server. The server analyzes the audio data using a voice analysis device and understands the characteristics of the audio using a generative AI model. Based on this, the server adjusts the volume and sound quality to make it easy for the user to hear.
[0359] This system is equipped with an emotion engine, and the server analyzes the user's emotional state from their voice in real time. The emotion engine analyzes changes in intonation, tempo, and volume of the voice to infer the user's emotions. The resulting emotional state is then used for further optimization in voice adjustment.
[0360] For example, if the emotion engine determines that the user is stressed, the server adjusts the voice to a calmer tone and further reduces background noise. The adjusted voice is then sent back from the server to the terminal and delivered to the user as clear audio.
[0361] Furthermore, along with the audio data transcribed into text by the text analysis method, emotional information analyzed by the emotion engine is displayed on the user interface. Specifically, simple feedback corresponding to the emotional state is displayed on the screen, allowing the user to obtain information based on their own emotional state. For example, if the user's emotions are analyzed as being calm, a message such as "Relaxed" will be displayed.
[0362] This system will allow elderly individuals to enjoy more fulfilling communication not only through voice but also through emotional feedback. The introduction of an emotion engine will enable individualized adaptation that takes into account the user's emotions, which is expected to further improve the quality of communication.
[0363] The following describes the processing flow.
[0364] Step 1:
[0365] The device captures the user's call audio using a microphone and converts the analog signal into digital audio data. The audio data is then encoded into a format that allows for real-time processing.
[0366] Step 2:
[0367] The terminal sends the converted audio data to the server. The transmission is performed using a secure protocol to minimize latency.
[0368] Step 3:
[0369] When the server receives audio data, it first uses an audio analysis tool to analyze the audio frequency and volume. This analysis provides the basic data needed to adjust the sound quality to make it easier for the user to hear.
[0370] Step 4:
[0371] The server uses an emotion engine to analyze the emotional nuances contained in the speech. It analyzes the intonation, tempo, and volume fluctuations of the speech to infer the user's emotional state.
[0372] Step 5:
[0373] The server adjusts the volume and frequency of the audio based on the results of voice analysis and emotion analysis. The adjustments are adapted to the user's auditory characteristics and estimated emotional state.
[0374] Step 6:
[0375] The server sends back audio, adjusted to a sound quality that allows the user to relax, to the terminal via an audio output device. The terminal then plays the adjusted audio clearly.
[0376] Step 7:
[0377] The server converts the audio data into text, identifying and extracting important keywords. The extracted text information, along with emotional feedback, is sent to the terminal.
[0378] Step 8:
[0379] The device displays received text and sentiment information on its screen. Users can visually confirm key points of the conversation and receive feedback on their own emotional state.
[0380] Step 9:
[0381] By reviewing the displayed information and understanding parts that were difficult to hear or their own emotional state, users can facilitate smoother communication.
[0382] (Example 2)
[0383] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0384] When elderly people communicate with others via voice, the sound quality and volume may be insufficient, or the user's emotional state may not be understood, leading to communication that proceeds without proper understanding. This can hinder smooth communication and result in an unsatisfactory experience. Therefore, there is a need for technology that improves sound quality and provides feedback that responds to the user's emotions, thereby enriching the communication experience.
[0385] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0386] In this invention, the server includes means for receiving digitized speech and analyzing speech characteristics in real time using a generation AI model; a speech adjustment device that adjusts the volume and sound quality based on the analyzed speech characteristics and emotional state to convert it into an easily audible sound; and a text processing device that generates text from the speech and extracts important information. As a result, users can receive clearer and easier-to-understand speech and also receive feedback tailored to their emotional state.
[0387] "Elderly people" generally refers to people who are older, and in this invention, they are a group that requires special consideration when communication is primarily conducted via voice.
[0388] "Digitalization" is the process of sampling analog audio signals and converting them into digital signals, and it is a fundamental step that enables data processing for the entire system.
[0389] A "generative AI model" is a type of artificial intelligence used to analyze speech characteristics, designed to perform intelligible speech conversion and emotion analysis in real time.
[0390] "Sound characteristics" refer to the physical and sensory characteristics of an audio signal, and specifically include volume, sound quality, frequency spectrum, intonation, tempo, etc.
[0391] A "voice adjustment device" is a device that adjusts voice to an optimal state based on analyzed voice characteristics and emotional state, thereby allowing users to receive voices that are easier to hear.
[0392] A "text processing device" is a device that converts audio signals into text and extracts important information, making it useful for providing information to users.
[0393] "Feedback" is information provided based on the user's voice and emotional state, and is intended to help users understand their own expressions and facilitate smooth communication.
[0394] This system is designed to help elderly people communicate smoothly and comfortably using voice. The system mainly consists of terminals, servers, and various software components.
[0395] The user initiates a voice call using their device. The device has a built-in microphone that captures the user's voice as an analog signal. The captured audio is digitized by the device, sampled, and converted into a digital audio format. This digitized audio data is then transmitted to the server via the network.
[0396] The server processes the received audio data using a generative AI model. Specifically, it analyzes the characteristics of the audio in real time and performs adjustments such as noise reduction and volume adjustment. Furthermore, the server analyzes the intonation and tempo of the audio using its built-in emotion engine and infers the user's emotional state. The results of this emotion analysis are also fed back into the audio adjustment process, generating the optimal audio for the user.
[0397] The server also converts the audio data into text and extracts important information. This text, along with feedback on the emotional state, is displayed in the user interface, allowing the user to understand their own emotional state while communicating.
[0398] As a concrete example, here is an example of a prompt sentence to give to a generative AI model: "Based on the audio data the user is speaking, use the emotion engine to analyze the emotions and provide appropriate feedback."
[0399] This system is expected to improve the quality of communication among the elderly by enabling higher-quality voice communication and providing more comprehensive feedback based on emotional understanding.
[0400] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0401] Step 1:
[0402] The user initiates a call using the device. The device's microphone captures the user's speech and recognizes it as an analog audio signal. The device then samples and quantizes this input audio, converting it into a digital audio format. In this process, the analog signal is transformed into digital data. Digital audio data is generated and prepared for the next processing step.
[0403] Step 2:
[0404] The terminal transmits digitized audio data to the server via the network. The server uses the received audio data as input to run a generating AI model, which analyzes the audio characteristics in real time. Data calculations such as volume, sound quality, and frequency spectrum are performed, and parts that require noise reduction or volume adjustment are identified. The analysis results output the points that need adjustment and the adjusted audio characteristics.
[0405] Step 3:
[0406] The server works in conjunction with a generative AI model to activate the emotion engine. It analyzes changes in intonation, tempo, and volume from the speech to estimate the user's emotional state. The analyzed emotional state is detected as input, and the emotion is output as a state such as stress or relaxation. This information is used for subsequent speech adjustments.
[0407] Step 4:
[0408] The server activates an audio adjustment device based on the analyzed voice characteristics and emotional state. Specifically, it adjusts the voice to make it easier to hear by applying noise cancellation and volume adjustments. For example, if stress is detected, it processes the voice to calm it down. This process generates a clear voice, which is then prepared for transmission.
[0409] Step 5:
[0410] The server sends the adjusted audio data back to the terminal. The terminal provides the received audio to the user through an audio output device. Here, it is output as the audio the user actually hears. The server also transcribes the audio data into text, extracts important information, and supplies it to the user interface display as feedback information. This allows the user to access information including emotional feedback.
[0411] Step 6:
[0412] Users refer to text information and sentiment feedback displayed on the device's screen. This allows them to understand their own expressions, obtain the information necessary for communication, and use it for the next step. By reviewing the displayed feedback information, they can gain insights to communicate effectively.
[0413] (Application Example 2)
[0414] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0415] Improving the quality of voice communication is crucial for the elderly and those with limited communication abilities. However, in addition to ensuring clear voice quality, there is a lack of systems that can accurately grasp the emotional state of the other party and respond appropriately as needed. Furthermore, real-time emotion recognition is required to detect and respond to potential security risks arising from emotional shifts at an early stage. Solving these challenges is essential.
[0416] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0417] In this invention, the server includes an acoustic input means, an acoustic analysis means, an adjustment means, an emotion recognition means, and an alarm means. This makes it easy for the elderly to use and enables real-time detection and response to security risks based on emotions.
[0418] An "acoustic input device" is a device that acquires voice signals emitted by elderly people and converts them into voice information.
[0419] An "acoustic analysis device" is a device that analyzes received audio information in real time using a generation model to understand acoustic characteristics and emotional states.
[0420] A "adjustment device" is a device that adjusts the volume and sound quality based on the results of acoustic analysis, converting the sound into one that is easy for elderly people to hear.
[0421] "Audio output means" refers to a device for providing adjusted sound to the user.
[0422] A "language analysis device" is a device that has the function of extracting important words from audio information and displaying that information on the user's screen.
[0423] A "display device" is a device that visually shows analyzed emotional information and important words on the user's screen.
[0424] An "emotion recognition device" is a device that analyzes the characteristics of speech from speech information obtained by an acoustic analysis device to estimate an emotional state.
[0425] An "alarm device" is a device that has the function of detecting potential security risks and providing necessary notifications based on an estimated emotional state.
[0426] The system of this invention supports voice communication for the elderly and provides emotion-based adaptive feedback and security features. A detailed embodiment of this system is shown below.
[0427] First, the terminal uses an acoustic input device to acquire an audio signal from the user and converts it into digital data as audio information. This audio information is then sent to the server.
[0428] The server uses acoustic analysis tools and generative AI models to analyze speech information in real time. Specifically, the acoustic analysis engine analyzes changes in speech intonation, tempo, and volume, and estimates the user's emotional state through emotion recognition tools. As generative AI models, for example, Google Cloud's Speech-to-Text API and Emotion Recognition API can be used.
[0429] Next, the adjustment means adjusts the volume and sound quality based on the analyzed information to create a sound that is easy for the user to hear. This adjusted sound is then provided to the user by the sound output means.
[0430] Furthermore, the language analysis tool extracts important words from the audio information and displays them visually on the user's screen via the display tool. Simultaneously, emotional information is also displayed on the screen, allowing the user to understand their own emotional state.
[0431] The alarm system activates when the estimated emotional state indicates a certain level of security risk. Specifically, when the emotional state is determined to be stressful, it sends a notification to registered contacts using communication methods such as Twilio or SendGrid.
[0432] This system allows elderly individuals to safely enjoy communication not only through voice but also through emotional feedback. For example, when an elderly person feels lonely, their family is notified, prompting them to take care of them.
[0433] Examples of prompt messages include the following:
[0434] "Analyze the current emotional state and determine whether the elderly person's emotions are calm."
[0435] "Use the user's voice data to detect their stress levels and send alerts if necessary."
[0436] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0437] Step 1:
[0438] The terminal acquires the user's voice using an acoustic input device. The input is the user's voice signal, which the terminal converts into digital data. The converted digital audio data is then sent to the server.
[0439] Step 2:
[0440] The server uses speech analysis tools to analyze digital speech data through a generative AI model. The input is digital speech data, and the server analyzes changes in intonation, tempo, and volume. As a result of the analysis, acoustic characteristics and emotional states are extracted.
[0441] Step 3:
[0442] The server adjusts the volume and sound quality based on the analysis results using adjustment mechanisms. The input is acoustic characteristics, and the server adjusts the sound to be easily audible to the user, generating clear, adjusted audio data. The adjusted audio data is then transmitted to the terminal.
[0443] Step 4:
[0444] The terminal outputs adjusted audio to the user using an audio output means. The input is adjusted audio data, allowing the user to hear clear audio.
[0445] Step 5:
[0446] The server extracts important words from audio information using language analysis tools. The input is audio information, and the server performs language analysis to extract important words. Once important words are extracted, the data is transmitted to the terminal via a display device.
[0447] Step 6:
[0448] The terminal uses a display mechanism to show important words and estimated sentiment information on the user's screen. The input consists of extracted important words and sentiment information, allowing the user to visually confirm their emotional state.
[0449] Step 7:
[0450] The server utilizes emotion recognition mechanisms to detect potential security risks based on estimated emotional states. If the emotion indicates stress or danger, it uses that state as input to send a notification command to the alarm system.
[0451] Step 8:
[0452] The server uses alarm mechanisms to notify pre-configured emergency contacts via communication means as needed. The input is information about detected security risks, a notification is output, and relevant parties are prompted to take prompt action.
[0453] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0454] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0455] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0456] [Third Embodiment]
[0457] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0458] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0459] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0460] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0461] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0462] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0463] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0464] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0465] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0466] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0467] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0468] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0469] This invention is a system that supports voice communication for the elderly and consists of multiple components. Specifically, it comprises a terminal held by the user, a server that processes the voice, and a network connecting the two.
[0470] The user initiates a call using the device. The device has a voice input mechanism and captures the user's voice as soon as they begin speaking, converting it into digital data. The converted voice data is then sent to the server using a protocol optimized to maintain low latency.
[0471] Upon receiving this audio data, the server first analyzes the data in real time using an audio analysis tool. A generative AI model is utilized to extract information on important frequency bands and background noise from the audio. Based on the analysis results, the server adjusts the volume and sound quality to make it easier for elderly people to hear. This adjusted audio is then returned to the user's terminal through an audio output tool and played back as clear audio.
[0472] In addition, the server uses text analysis to convert the audio data into text. From the transcribed data, important keywords are identified and displayed on the terminal's display through the user interface. This allows for visual supplementation of parts that were difficult to understand through auditory means.
[0473] As a concrete example, consider a case where user A uses the device to make a call with family members. The family members' voices are captured by the device's microphone and immediately sent to the server. The server analyzes the audio in real time, adjusts the sound quality to match user A's hearing characteristics, and then returns the adjusted audio to the device. Simultaneously, the conversation is transcribed into text, and particularly important keywords are displayed on the screen, allowing user A to confirm the conversation content not only aurally but also visually.
[0474] This system makes telephone communication more enjoyable and stress-free for users with hearing impairments. This invention could be a significant support for elderly people in maintaining independent living.
[0475] The following describes the processing flow.
[0476] Step 1:
[0477] The device captures the user's voice conversation using a microphone. The audio signal is converted from analog to digital and processed as audio data.
[0478] Step 2:
[0479] The terminal transmits the converted audio data to the server over the network. The audio data is transferred using an efficient protocol to maintain low latency.
[0480] Step 3:
[0481] The server analyzes the received audio data. It applies a generative AI model to analyze the frequency characteristics and noise components of the audio in real time.
[0482] Step 4:
[0483] Based on the results of the audio analysis, the server optimizes the volume and sound quality to make it easier for the user to hear. It emphasizes specific frequency bands and reduces background noise.
[0484] Step 5:
[0485] The server re-encodes the optimized audio and sends it to the terminal. The adjusted audio is then supplied to the user via the audio output device, resulting in clear playback.
[0486] Step 6:
[0487] The server simultaneously converts the audio data into text. It uses speech recognition technology to identify important keywords and extract them as text.
[0488] Step 7:
[0489] The terminal receives text data sent from the server and displays it on the screen. Important keywords are presented in a visually recognizable format within the user interface.
[0490] Step 8:
[0491] Users can review the audio-text displayed on their device, supplementing any parts of the conversation that were difficult to understand. This process makes it easier to accurately understand the content of the conversation.
[0492] (Example 1)
[0493] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0494] Providing voice communication tailored to the hearing characteristics of the elderly is challenging, and conventional communication devices may not allow for sufficient comprehension of conversations. Furthermore, in environments with significant background noise, hearing difficulties increase, potentially leading to the loss of important information. This can result in users experiencing stress and discomfort during communication.
[0495] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0496] In this invention, the server includes an input device that collects voice information and converts it into digital data, an analysis device that uses an artificial intelligence model to perform real-time analysis, and an adjustment device that adjusts the volume and sound quality to convert it into easily audible voice. This enables clear voice communication tailored to individual hearing characteristics.
[0497] "Audio information" refers to human speech and sounds, which are signals collected for digital processing.
[0498] "Digital data" refers to information obtained by converting analog signals into a format that can be processed by computers and digital devices.
[0499] An "input device" is a hardware or software component that collects audio information and converts it into digital data.
[0500] An "artificial intelligence model" is an algorithm that learns patterns and features from large amounts of data and performs analysis and prediction.
[0501] "Real-time analysis" is a process that processes input data immediately and outputs results quickly.
[0502] An "analysis device" is a component used to process received digital data using an artificial intelligence model.
[0503] "Volume and sound quality adjustment" is the process of adapting the volume and tone of received audio to the individual characteristics of each device.
[0504] A "regulating device" is a device that has the function of optimizing volume and sound quality.
[0505] "Easy-to-listen audio" refers to audio signals that are easy to understand and less burdensome to the listener, adjusted according to the user's auditory characteristics.
[0506] An "output device" is a device or interface for providing a pre-tuned audio signal to the user.
[0507] "Key terms" are words or phrases in audio data that contain particularly important information.
[0508] An "analysis device" is a device or software that has the function of extracting key words from audio data.
[0509] A "display function" refers to an interface or device that visually shows the extracted main terms to the user.
[0510] This invention is a system that supports voice communication for the elderly, and its main components include an input device, an analysis device, an adjustment device, an output device, an analytical device, and a display function.
[0511] Input device:
[0512] The user uses a device to acquire voice information and converts it into digital data. This process utilizes a microphone and a digital signal processing chip.
[0513] Analysis equipment:
[0514] The server analyzes the received digital data in real time using an artificial intelligence model. This AI model performs noise reduction and frequency enhancement. For example, it is possible to set prompts such as "reduce ambient noise and highlight important conversations."
[0515] Adjustment device:
[0516] The server adjusts the volume and sound quality based on the analysis results. This adjustment process uses equalization and sound pressure optimization techniques to adapt to the user's auditory characteristics.
[0517] Output device:
[0518] The device plays pre-tuned audio transmitted from the server, ensuring the user receives clear sound. Playback utilizes the device's speakers or headphones.
[0519] Analytical device and display function:
[0520] The server converts the audio data into text and extracts key phrases. This process identifies important information and displays it on the terminal's display. This allows the user to visually confirm the content of the conversation.
[0521] As a concrete example, when user A uses the device to make a call with family members, the family members' voices are captured by the device's microphone and immediately transmitted to the server as digital data. The server analyzes the audio in real time, removes noise, and sends the newly adjusted audio back to the device. At the same time, important keywords are displayed as text on the device's screen to complement the conversation. This allows user A to communicate smoothly within the family using both their ears and eyes.
[0522] This system provides hearing aids tailored to the hearing characteristics of elderly individuals, supporting verbal communication in daily life. This allows users to enjoy a more secure and independent life.
[0523] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0524] Step 1:
[0525] The terminal acquires the user's voice information and converts it from analog to digital. The input is the user's voice, captured by the terminal's microphone. This voice signal is converted into digital data by a digital signal processing chip. The output is digital voice data, which is then used for further processing.
[0526] Step 2:
[0527] The terminal sends the converted digital audio data to the server using an optimized low-latency protocol. The input is the digital audio data from step 1. Based on the transmission protocol (e.g., WebSocket or UDP), the data is transported to the server, minimizing network latency. The output is the completion of the transfer of the audio data to the server.
[0528] Step 3:
[0529] The server analyzes received digital audio data in real time using a generating AI model. The input is digital audio data sent from the terminal. Based on the prompt "remove noise and emphasize important frequencies," the AI model analyzes the audio data, extracts important frequency bands, and reduces noise. The output is the analyzed audio information.
[0530] Step 4:
[0531] The server adjusts the volume and sound quality based on the analysis results. This process uses analyzed audio information from an AI model as input. Equalization and compression are used for adjustment to optimize the audio to suit the user's auditory characteristics. The output is the adjusted audio data.
[0532] Step 5:
[0533] The server sends the adjusted audio data to the terminal. The input is the adjusted audio data obtained in step 4. The data is sent to the terminal again using the optimized protocol, maintaining low latency. The output is the reception of the adjusted audio data on the terminal.
[0534] Step 6:
[0535] The terminal plays the pre-tuned audio received from the server. The input is the pre-tuned audio data sent from the server. Clear audio is provided to the user using the terminal's speaker or headphones. The output is audio playback in an audible format.
[0536] Step 7:
[0537] The server transcribes the audio data into text and extracts key keywords through analysis. The input is the audio data obtained in step 3. Using speech recognition technology, it generates transcribed data and extracts key phrases from the conversation. The output is the transcribed data and key phrases.
[0538] Step 8:
[0539] The terminal displays the extracted keywords on its screen. Input consists of transcribed data and key phrases from the server. The user interface complements auditory input by visually displaying the main points of the conversation. Output is a display of keywords as visual information on the screen.
[0540] (Application Example 1)
[0541] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0542] For the elderly, the insufficient accuracy of voice recognition when giving voice instructions for electronic transactions, and the inability to visually confirm the content of the voice instructions, are major obstacles. Furthermore, sound quality issues caused by background noise and other noises affect the accuracy of voice instructions, highlighting the need for effective solutions.
[0543] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0544] In this invention, the server includes voice acquisition means, voice analysis means, voice adjustment means, and electronic transaction support means. This makes it possible to instantly adjust the sound quality in response to voice instructions and provide voice feedback that is easy for the elderly to hear. Furthermore, by visually presenting important transaction details, it is possible to prevent misunderstandings of transactions based solely on voice instructions and to support the elderly in confidently conducting electronic transactions by voice.
[0545] A "voice acquisition means" is a device that captures a user's speech as an audio signal and converts it into audio information.
[0546] A "voice analysis device" is a device that has the function of analyzing acquired voice information in real time using a generation AI model and extracting important content.
[0547] A "sound adjustment means" is a device that, based on analyzed sound information, adjusts the volume and sound quality to match the user's auditory characteristics, converting it into easily understandable sound.
[0548] "Audio playback means" refers to a device for outputting adjusted audio to the user.
[0549] A "text analysis device" is a device that extracts important words and phrases from audio information and processes them as text information.
[0550] An "information presentation means" is a device that displays important keywords extracted by a text analysis means on a user interface.
[0551] An "electronic transaction support device" is a device that receives voice instructions from a user and performs the necessary information processing to facilitate electronic transactions.
[0552] A "display support device" is a device that has the function of displaying information so that transaction details can be visually confirmed.
[0553] This invention is a system that supports electronic transactions via voice commands for the elderly. The system consists of a user terminal, a server, and a network connecting them. When a user gives voice commands for an electronic transaction, the voice is captured by a microphone on the terminal and transmitted to the server as voice data.
[0554] The server uses the Google Cloud Speech-to-Text API to perform speech analysis, converting the audio data into text. Furthermore, it extracts important keywords from the analyzed text using a natural language processing (NLP) library and presents them visually to the user.
[0555] To reduce audio noise and adjust sound quality to suit the hearing characteristics of elderly individuals, a generative AI model using TensorFlow further processes the audio data. The processed audio is returned to the device and played back in a format that is easy for the user to understand.
[0556] The terminal's display shows important information about electronic transactions conducted via voice commands, such as the trading partner and the amount. This display helps the user confirm the transaction.
[0557] For example, if a user gives a voice command such as, "I want to buy a 500 yen item at the nearby supermarket," the command is instantly converted into text, and the necessary information is displayed. Furthermore, the AI model analyzes the voice command and returns feedback with appropriately adjusted volume and sound quality. A confirmation prompt appears asking, "Do you want to proceed with purchasing the 500 yen item at the supermarket based on this voice command?"
[0558] An example of a prompt is, "If a user indicates their intention to make an electronic transaction via voice, explain in detail how to process this intention." This prompt is useful for confirming and understanding the steps involved in voice analysis and transaction processing using a generative AI model.
[0559] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0560] Step 1:
[0561] The user issues voice commands to the terminal. The terminal captures the voice signal using its built-in microphone and converts it into digital audio data. This data is then sent to the server.
[0562] Step 2:
[0563] The server converts the received digital audio data into text using the Google Cloud Speech-to-Text API. The input is audio data, and the output is text data. The server extracts the converted text data and passes it on to the next processing step.
[0564] Step 3:
[0565] The server analyzes text data using generative AI models and natural language processing (NLP) techniques. The analysis extracts and identifies key keywords. The input is text data, and the output is a list of key keywords. This makes important instructions regarding transactions easier to understand.
[0566] Step 4:
[0567] The server organizes the extracted keywords and constructs the information necessary to visualize the details of electronic transactions. This information is sent to the user's terminal and displayed on the screen. The input is a list of important keywords, and the output is transaction information for display.
[0568] Step 5:
[0569] The server analyzes the audio data using an AI model and adjusts the sound quality to suit the hearing characteristics of elderly individuals. This adjusted audio data is then sent to the terminal and played back. The input is the original audio data, and the output is optimized audio data. This allows the user to receive clear audio feedback.
[0570] Step 6:
[0571] The terminal simultaneously presents the user with transaction information for display and pre-recorded audio. The user can verify transaction details both visually and aurally. Input consists of transaction information for display and pre-recorded audio data, while the visual display and audio output constitute the output.
[0572] Step 7:
[0573] Once the user confirms the details and indicates their intention to execute the transaction, the details are sent back from the terminal to the server, and the electronic transaction is completed. The input is the confirmed transaction information, and the output is a transaction success message. This ensures that electronic transactions are conducted safely and smoothly.
[0574] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0575] This invention provides a voice communication support system for the elderly that incorporates an emotion engine that recognizes the user's emotions in addition to voice analysis. This system consists of a user terminal and a server, which are connected via a network.
[0576] The user makes a call using their device, and the audio is captured by a voice input device. The captured audio data is digitized and sent to a server. The server analyzes the audio data using a voice analysis device and understands the characteristics of the audio using a generative AI model. Based on this, the server adjusts the volume and sound quality to make it easy for the user to hear.
[0577] This system is equipped with an emotion engine, and the server analyzes the user's emotional state from their voice in real time. The emotion engine analyzes changes in intonation, tempo, and volume of the voice to infer the user's emotions. The resulting emotional state is then used for further optimization in voice adjustment.
[0578] For example, if the emotion engine determines that the user is stressed, the server adjusts the voice to a calmer tone and further reduces background noise. The adjusted voice is then sent back from the server to the terminal and delivered to the user as clear audio.
[0579] Furthermore, along with the audio data transcribed into text by the text analysis method, emotional information analyzed by the emotion engine is displayed on the user interface. Specifically, simple feedback corresponding to the emotional state is displayed on the screen, allowing the user to obtain information based on their own emotional state. For example, if the user's emotions are analyzed as being calm, a message such as "Relaxed" will be displayed.
[0580] This system will allow elderly individuals to enjoy more fulfilling communication not only through voice but also through emotional feedback. The introduction of an emotion engine will enable individualized adaptation that takes into account the user's emotions, which is expected to further improve the quality of communication.
[0581] The following describes the processing flow.
[0582] Step 1:
[0583] The device captures the user's call audio using a microphone and converts the analog signal into digital audio data. The audio data is then encoded into a format that allows for real-time processing.
[0584] Step 2:
[0585] The terminal sends the converted audio data to the server. The transmission is performed using a secure protocol to minimize latency.
[0586] Step 3:
[0587] When the server receives audio data, it first uses an audio analysis tool to analyze the audio frequency and volume. This analysis provides the basic data needed to adjust the sound quality to make it easier for the user to hear.
[0588] Step 4:
[0589] The server uses an emotion engine to analyze the emotional nuances contained in the speech. It analyzes the intonation, tempo, and volume fluctuations of the speech to infer the user's emotional state.
[0590] Step 5:
[0591] The server adjusts the volume and frequency of the audio based on the results of voice analysis and emotion analysis. The adjustments are adapted to the user's auditory characteristics and estimated emotional state.
[0592] Step 6:
[0593] The server sends back audio, adjusted to a sound quality that allows the user to relax, to the terminal via an audio output device. The terminal then plays the adjusted audio clearly.
[0594] Step 7:
[0595] The server converts the audio data into text, identifying and extracting important keywords. The extracted text information, along with emotional feedback, is sent to the terminal.
[0596] Step 8:
[0597] The device displays received text and sentiment information on its screen. Users can visually confirm key points of the conversation and receive feedback on their own emotional state.
[0598] Step 9:
[0599] By reviewing the displayed information and understanding parts that were difficult to hear or their own emotional state, users can facilitate smoother communication.
[0600] (Example 2)
[0601] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0602] When elderly people communicate with others via voice, the sound quality and volume may be insufficient, or the user's emotional state may not be understood, leading to communication that proceeds without proper understanding. This can hinder smooth communication and result in an unsatisfactory experience. Therefore, there is a need for technology that improves sound quality and provides feedback that responds to the user's emotions, thereby enriching the communication experience.
[0603] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0604] In this invention, the server includes means for receiving digitized speech and analyzing speech characteristics in real time using a generation AI model; a speech adjustment device that adjusts the volume and sound quality based on the analyzed speech characteristics and emotional state to convert it into an easily audible sound; and a text processing device that generates text from the speech and extracts important information. As a result, users can receive clearer and easier-to-understand speech and also receive feedback tailored to their emotional state.
[0605] "Elderly people" generally refers to people who are older, and in this invention, they are a group that requires special consideration when communication is primarily conducted via voice.
[0606] "Digitalization" is the process of sampling analog audio signals and converting them into digital signals, and it is a fundamental step that enables data processing for the entire system.
[0607] A "generative AI model" is a type of artificial intelligence used to analyze speech characteristics, designed to perform intelligible speech conversion and emotion analysis in real time.
[0608] "Sound characteristics" refer to the physical and sensory characteristics of an audio signal, and specifically include volume, sound quality, frequency spectrum, intonation, tempo, etc.
[0609] A "voice adjustment device" is a device that adjusts voice to an optimal state based on analyzed voice characteristics and emotional state, thereby allowing users to receive voices that are easier to hear.
[0610] A "text processing device" is a device that converts audio signals into text and extracts important information, making it useful for providing information to users.
[0611] "Feedback" is information provided based on the user's voice and emotional state, and is intended to help users understand their own expressions and facilitate smooth communication.
[0612] This system is designed to help elderly people communicate smoothly and comfortably using voice. The system mainly consists of terminals, servers, and various software components.
[0613] The user initiates a voice call using their device. The device has a built-in microphone that captures the user's voice as an analog signal. The captured audio is digitized by the device, sampled, and converted into a digital audio format. This digitized audio data is then transmitted to the server via the network.
[0614] The server processes the received audio data using a generative AI model. Specifically, it analyzes the characteristics of the audio in real time and performs adjustments such as noise reduction and volume adjustment. Furthermore, the server analyzes the intonation and tempo of the audio using its built-in emotion engine and infers the user's emotional state. The results of this emotion analysis are also fed back into the audio adjustment process, generating the optimal audio for the user.
[0615] The server also converts the audio data into text and extracts important information. This text, along with feedback on the emotional state, is displayed in the user interface, allowing the user to understand their own emotional state while communicating.
[0616] As a concrete example, here is an example of a prompt sentence to give to a generative AI model: "Based on the audio data the user is speaking, use the emotion engine to analyze the emotions and provide appropriate feedback."
[0617] This system is expected to improve the quality of communication among the elderly by enabling higher-quality voice communication and providing more comprehensive feedback based on emotional understanding.
[0618] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0619] Step 1:
[0620] The user initiates a call using the device. The device's microphone captures the user's speech and recognizes it as an analog audio signal. The device then samples and quantizes this input audio, converting it into a digital audio format. In this process, the analog signal is transformed into digital data. Digital audio data is generated and prepared for the next processing step.
[0621] Step 2:
[0622] The terminal transmits digitized audio data to the server via the network. The server uses the received audio data as input to run a generating AI model, which analyzes the audio characteristics in real time. Data calculations such as volume, sound quality, and frequency spectrum are performed, and parts that require noise reduction or volume adjustment are identified. The analysis results output the points that need adjustment and the adjusted audio characteristics.
[0623] Step 3:
[0624] The server works in conjunction with a generative AI model to activate the emotion engine. It analyzes changes in intonation, tempo, and volume from the speech to estimate the user's emotional state. The analyzed emotional state is detected as input, and the emotion is output as a state such as stress or relaxation. This information is used for subsequent speech adjustments.
[0625] Step 4:
[0626] The server activates an audio adjustment device based on the analyzed voice characteristics and emotional state. Specifically, it adjusts the voice to make it easier to hear by applying noise cancellation and volume adjustments. For example, if stress is detected, it processes the voice to calm it down. This process generates a clear voice, which is then prepared for transmission.
[0627] Step 5:
[0628] The server sends the adjusted audio data back to the terminal. The terminal provides the received audio to the user through an audio output device. Here, it is output as the audio the user actually hears. The server also transcribes the audio data into text, extracts important information, and supplies it to the user interface display as feedback information. This allows the user to access information including emotional feedback.
[0629] Step 6:
[0630] Users refer to text information and sentiment feedback displayed on the device's screen. This allows them to understand their own expressions, obtain the information necessary for communication, and use it for the next step. By reviewing the displayed feedback information, they can gain insights to communicate effectively.
[0631] (Application Example 2)
[0632] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0633] Improving the quality of voice communication is crucial for the elderly and those with limited communication abilities. However, in addition to ensuring clear voice quality, there is a lack of systems that can accurately grasp the emotional state of the other party and respond appropriately as needed. Furthermore, real-time emotion recognition is required to detect and respond to potential security risks arising from emotional shifts at an early stage. Solving these challenges is essential.
[0634] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0635] In this invention, the server includes an acoustic input means, an acoustic analysis means, an adjustment means, an emotion recognition means, and an alarm means. This makes it easy for the elderly to use and enables real-time detection and response to security risks based on emotions.
[0636] An "acoustic input device" is a device that acquires voice signals emitted by elderly people and converts them into voice information.
[0637] An "acoustic analysis device" is a device that analyzes received audio information in real time using a generation model to understand acoustic characteristics and emotional states.
[0638] A "adjustment device" is a device that adjusts the volume and sound quality based on the results of acoustic analysis, converting the sound into one that is easy for elderly people to hear.
[0639] "Audio output means" refers to a device for providing adjusted sound to the user.
[0640] A "language analysis device" is a device that has the function of extracting important words from audio information and displaying that information on the user's screen.
[0641] A "display device" is a device that visually shows analyzed emotional information and important words on the user's screen.
[0642] An "emotion recognition device" is a device that analyzes the characteristics of speech from speech information obtained by an acoustic analysis device to estimate an emotional state.
[0643] An "alarm device" is a device that has the function of detecting potential security risks and providing necessary notifications based on an estimated emotional state.
[0644] The system of this invention supports voice communication for the elderly and provides emotion-based adaptive feedback and security features. A detailed embodiment of this system is shown below.
[0645] First, the terminal uses an acoustic input device to acquire an audio signal from the user and converts it into digital data as audio information. This audio information is then sent to the server.
[0646] The server uses acoustic analysis tools and generative AI models to analyze speech information in real time. Specifically, the acoustic analysis engine analyzes changes in speech intonation, tempo, and volume, and estimates the user's emotional state through emotion recognition tools. As generative AI models, for example, Google Cloud's Speech-to-Text API and Emotion Recognition API can be used.
[0647] Next, the adjustment means adjusts the volume and sound quality based on the analyzed information to create a sound that is easy for the user to hear. This adjusted sound is then provided to the user by the sound output means.
[0648] Furthermore, the language analysis tool extracts important words from the audio information and displays them visually on the user's screen via the display tool. Simultaneously, emotional information is also displayed on the screen, allowing the user to understand their own emotional state.
[0649] The alarm system activates when the estimated emotional state indicates a certain level of security risk. Specifically, when the emotional state is determined to be stressful, it sends a notification to registered contacts using communication methods such as Twilio or SendGrid.
[0650] This system allows elderly individuals to safely enjoy communication not only through voice but also through emotional feedback. For example, when an elderly person feels lonely, their family is notified, prompting them to take care of them.
[0651] Examples of prompt messages include the following:
[0652] "Analyze the current emotional state and determine whether the elderly person's emotions are calm."
[0653] "Use the user's voice data to detect their stress levels and send alerts if necessary."
[0654] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0655] Step 1:
[0656] The terminal acquires the user's voice using an acoustic input device. The input is the user's voice signal, which the terminal converts into digital data. The converted digital audio data is then sent to the server.
[0657] Step 2:
[0658] The server uses speech analysis tools to analyze digital speech data through a generative AI model. The input is digital speech data, and the server analyzes changes in intonation, tempo, and volume. As a result of the analysis, acoustic characteristics and emotional states are extracted.
[0659] Step 3:
[0660] The server adjusts the volume and sound quality based on the analysis results using adjustment mechanisms. The input is acoustic characteristics, and the server adjusts the sound to be easily audible to the user, generating clear, adjusted audio data. The adjusted audio data is then transmitted to the terminal.
[0661] Step 4:
[0662] The terminal outputs adjusted audio to the user using an audio output means. The input is adjusted audio data, allowing the user to hear clear audio.
[0663] Step 5:
[0664] The server extracts important words from audio information using language analysis tools. The input is audio information, and the server performs language analysis to extract important words. Once important words are extracted, the data is transmitted to the terminal via a display device.
[0665] Step 6:
[0666] The terminal uses a display mechanism to show important words and estimated sentiment information on the user's screen. The input consists of extracted important words and sentiment information, allowing the user to visually confirm their emotional state.
[0667] Step 7:
[0668] The server utilizes emotion recognition mechanisms to detect potential security risks based on estimated emotional states. If the emotion indicates stress or danger, it uses that state as input to send a notification command to the alarm system.
[0669] Step 8:
[0670] The server uses alarm mechanisms to notify pre-configured emergency contacts via communication means as needed. The input is information about detected security risks, a notification is output, and relevant parties are prompted to take prompt action.
[0671] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0672] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0673] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0674] [Fourth Embodiment]
[0675] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0676] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0677] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0678] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0679] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0680] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0681] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0682] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0683] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0684] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0685] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0686] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0687] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0688] This invention is a system that supports voice communication for the elderly and consists of multiple components. Specifically, it comprises a terminal held by the user, a server that processes the voice, and a network connecting the two.
[0689] The user initiates a call using the device. The device has a voice input mechanism and captures the user's voice as soon as they begin speaking, converting it into digital data. The converted voice data is then sent to the server using a protocol optimized to maintain low latency.
[0690] Upon receiving this audio data, the server first analyzes the data in real time using an audio analysis tool. A generative AI model is utilized to extract information on important frequency bands and background noise from the audio. Based on the analysis results, the server adjusts the volume and sound quality to make it easier for elderly people to hear. This adjusted audio is then returned to the user's terminal through an audio output tool and played back as clear audio.
[0691] In addition, the server uses text analysis to convert the audio data into text. From the transcribed data, important keywords are identified and displayed on the terminal's display through the user interface. This allows for visual supplementation of parts that were difficult to understand through auditory means.
[0692] As a concrete example, consider a case where user A uses the device to make a call with family members. The family members' voices are captured by the device's microphone and immediately sent to the server. The server analyzes the audio in real time, adjusts the sound quality to match user A's hearing characteristics, and then returns the adjusted audio to the device. Simultaneously, the conversation is transcribed into text, and particularly important keywords are displayed on the screen, allowing user A to confirm the conversation content not only aurally but also visually.
[0693] This system makes telephone communication more enjoyable and stress-free for users with hearing impairments. This invention could be a significant support for elderly people in maintaining independent living.
[0694] The following describes the processing flow.
[0695] Step 1:
[0696] The device captures the user's voice conversation using a microphone. The audio signal is converted from analog to digital and processed as audio data.
[0697] Step 2:
[0698] The terminal transmits the converted audio data to the server over the network. The audio data is transferred using an efficient protocol to maintain low latency.
[0699] Step 3:
[0700] The server analyzes the received audio data. It applies a generative AI model to analyze the frequency characteristics and noise components of the audio in real time.
[0701] Step 4:
[0702] Based on the results of the audio analysis, the server optimizes the volume and sound quality to make it easier for the user to hear. It emphasizes specific frequency bands and reduces background noise.
[0703] Step 5:
[0704] The server re-encodes the optimized audio and sends it to the terminal. The adjusted audio is then supplied to the user via the audio output device, resulting in clear playback.
[0705] Step 6:
[0706] The server simultaneously converts the audio data into text. It uses speech recognition technology to identify important keywords and extract them as text.
[0707] Step 7:
[0708] The terminal receives text data sent from the server and displays it on the screen. Important keywords are presented in a visually recognizable format within the user interface.
[0709] Step 8:
[0710] Users can review the audio-text displayed on their device, supplementing any parts of the conversation that were difficult to understand. This process makes it easier to accurately understand the content of the conversation.
[0711] (Example 1)
[0712] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0713] Providing voice communication tailored to the hearing characteristics of the elderly is challenging, and conventional communication devices may not allow for sufficient comprehension of conversations. Furthermore, in environments with significant background noise, hearing difficulties increase, potentially leading to the loss of important information. This can result in users experiencing stress and discomfort during communication.
[0714] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0715] In this invention, the server includes an input device that collects voice information and converts it into digital data, an analysis device that uses an artificial intelligence model to perform real-time analysis, and an adjustment device that adjusts the volume and sound quality to convert it into easily audible voice. This enables clear voice communication tailored to individual hearing characteristics.
[0716] "Audio information" refers to human speech and sounds, which are signals collected for digital processing.
[0717] "Digital data" refers to information obtained by converting analog signals into a format that can be processed by computers and digital devices.
[0718] An "input device" is a hardware or software component that collects audio information and converts it into digital data.
[0719] An "artificial intelligence model" is an algorithm that learns patterns and features from large amounts of data and performs analysis and prediction.
[0720] "Real-time analysis" is a process that processes input data immediately and outputs results quickly.
[0721] An "analysis device" is a component used to process received digital data using an artificial intelligence model.
[0722] "Volume and sound quality adjustment" is the process of adapting the volume and tone of received audio to the individual characteristics of each device.
[0723] A "regulating device" is a device that has the function of optimizing volume and sound quality.
[0724] "Easy-to-listen audio" refers to audio signals that are easy to understand and less burdensome to the listener, adjusted according to the user's auditory characteristics.
[0725] An "output device" is a device or interface for providing a pre-tuned audio signal to the user.
[0726] "Key terms" are words or phrases in audio data that contain particularly important information.
[0727] An "analysis device" is a device or software that has the function of extracting key words from audio data.
[0728] A "display function" refers to an interface or device that visually shows the extracted main terms to the user.
[0729] This invention is a system that supports voice communication for the elderly, and its main components include an input device, an analysis device, an adjustment device, an output device, an analytical device, and a display function.
[0730] Input device:
[0731] The user uses a device to acquire voice information and converts it into digital data. This process utilizes a microphone and a digital signal processing chip.
[0732] Analysis equipment:
[0733] The server analyzes the received digital data in real time using an artificial intelligence model. This AI model performs noise reduction and frequency enhancement. For example, it is possible to set prompts such as "reduce ambient noise and highlight important conversations."
[0734] Adjustment device:
[0735] The server adjusts the volume and sound quality based on the analysis results. This adjustment process uses equalization and sound pressure optimization techniques to adapt to the user's auditory characteristics.
[0736] Output device:
[0737] The device plays pre-tuned audio transmitted from the server, ensuring the user receives clear sound. Playback utilizes the device's speakers or headphones.
[0738] Analytical device and display function:
[0739] The server converts the audio data into text and extracts key phrases. This process identifies important information and displays it on the terminal's display. This allows the user to visually confirm the content of the conversation.
[0740] As a concrete example, when user A uses the device to make a call with family members, the family members' voices are captured by the device's microphone and immediately transmitted to the server as digital data. The server analyzes the audio in real time, removes noise, and sends the newly adjusted audio back to the device. At the same time, important keywords are displayed as text on the device's screen to complement the conversation. This allows user A to communicate smoothly within the family using both their ears and eyes.
[0741] This system provides hearing aids tailored to the hearing characteristics of elderly individuals, supporting verbal communication in daily life. This allows users to enjoy a more secure and independent life.
[0742] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0743] Step 1:
[0744] The terminal acquires the user's voice information and converts it from analog to digital. The input is the user's voice, captured by the terminal's microphone. This voice signal is converted into digital data by a digital signal processing chip. The output is digital voice data, which is then used for further processing.
[0745] Step 2:
[0746] The terminal sends the converted digital audio data to the server using an optimized low-latency protocol. The input is the digital audio data from step 1. Based on the transmission protocol (e.g., WebSocket or UDP), the data is transported to the server, minimizing network latency. The output is the completion of the transfer of the audio data to the server.
[0747] Step 3:
[0748] The server analyzes received digital audio data in real time using a generating AI model. The input is digital audio data sent from the terminal. Based on the prompt "remove noise and emphasize important frequencies," the AI model analyzes the audio data, extracts important frequency bands, and reduces noise. The output is the analyzed audio information.
[0749] Step 4:
[0750] The server adjusts the volume and sound quality based on the analysis results. This process uses analyzed audio information from an AI model as input. Equalization and compression are used for adjustment to optimize the audio to suit the user's auditory characteristics. The output is the adjusted audio data.
[0751] Step 5:
[0752] The server sends the adjusted audio data to the terminal. The input is the adjusted audio data obtained in step 4. The data is sent to the terminal again using the optimized protocol, maintaining low latency. The output is the reception of the adjusted audio data on the terminal.
[0753] Step 6:
[0754] The terminal plays the pre-tuned audio received from the server. The input is the pre-tuned audio data sent from the server. Clear audio is provided to the user using the terminal's speaker or headphones. The output is audio playback in an audible format.
[0755] Step 7:
[0756] The server transcribes the audio data into text and extracts key keywords through analysis. The input is the audio data obtained in step 3. Using speech recognition technology, it generates transcribed data and extracts key phrases from the conversation. The output is the transcribed data and key phrases.
[0757] Step 8:
[0758] The terminal displays the extracted keywords on its screen. Input consists of transcribed data and key phrases from the server. The user interface complements auditory input by visually displaying the main points of the conversation. Output is a display of keywords as visual information on the screen.
[0759] (Application Example 1)
[0760] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0761] For the elderly, the insufficient accuracy of voice recognition when giving voice instructions for electronic transactions, and the inability to visually confirm the content of the voice instructions, are major obstacles. Furthermore, sound quality issues caused by background noise and other noises affect the accuracy of voice instructions, highlighting the need for effective solutions.
[0762] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0763] In this invention, the server includes voice acquisition means, voice analysis means, voice adjustment means, and electronic transaction support means. This makes it possible to instantly adjust the sound quality in response to voice instructions and provide voice feedback that is easy for the elderly to hear. Furthermore, by visually presenting important transaction details, it is possible to prevent misunderstandings of transactions based solely on voice instructions and to support the elderly in confidently conducting electronic transactions by voice.
[0764] A "voice acquisition means" is a device that captures a user's speech as an audio signal and converts it into audio information.
[0765] A "voice analysis device" is a device that has the function of analyzing acquired voice information in real time using a generation AI model and extracting important content.
[0766] A "sound adjustment means" is a device that, based on analyzed sound information, adjusts the volume and sound quality to match the user's auditory characteristics, converting it into easily understandable sound.
[0767] "Audio playback means" refers to a device for outputting adjusted audio to the user.
[0768] A "text analysis device" is a device that extracts important words and phrases from audio information and processes them as text information.
[0769] An "information presentation means" is a device that displays important keywords extracted by a text analysis means on a user interface.
[0770] An "electronic transaction support device" is a device that receives voice instructions from a user and performs the necessary information processing to facilitate electronic transactions.
[0771] A "display support device" is a device that has the function of displaying information so that transaction details can be visually confirmed.
[0772] This invention is a system that supports electronic transactions via voice commands for the elderly. The system consists of a user terminal, a server, and a network connecting them. When a user gives voice commands for an electronic transaction, the voice is captured by a microphone on the terminal and transmitted to the server as voice data.
[0773] The server uses the Google Cloud Speech-to-Text API to perform speech analysis, converting the audio data into text. Furthermore, it extracts important keywords from the analyzed text using a natural language processing (NLP) library and presents them visually to the user.
[0774] To reduce audio noise and adjust sound quality to suit the hearing characteristics of elderly individuals, a generative AI model using TensorFlow further processes the audio data. The processed audio is returned to the device and played back in a format that is easy for the user to understand.
[0775] The terminal's display shows important information about electronic transactions conducted via voice commands, such as the trading partner and the amount. This display helps the user confirm the transaction.
[0776] For example, if a user gives a voice command such as, "I want to buy a 500 yen item at the nearby supermarket," the command is instantly converted into text, and the necessary information is displayed. Furthermore, the AI model analyzes the voice command and returns feedback with appropriately adjusted volume and sound quality. A confirmation prompt appears asking, "Do you want to proceed with purchasing the 500 yen item at the supermarket based on this voice command?"
[0777] An example of a prompt is, "If a user indicates their intention to make an electronic transaction via voice, explain in detail how to process this intention." This prompt is useful for confirming and understanding the steps involved in voice analysis and transaction processing using a generative AI model.
[0778] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0779] Step 1:
[0780] The user issues voice commands to the terminal. The terminal captures the voice signal using its built-in microphone and converts it into digital audio data. This data is then sent to the server.
[0781] Step 2:
[0782] The server converts the received digital audio data into text using the Google Cloud Speech-to-Text API. The input is audio data, and the output is text data. The server extracts the converted text data and passes it on to the next processing step.
[0783] Step 3:
[0784] The server analyzes text data using generative AI models and natural language processing (NLP) techniques. The analysis extracts and identifies key keywords. The input is text data, and the output is a list of key keywords. This makes important instructions regarding transactions easier to understand.
[0785] Step 4:
[0786] The server organizes the extracted keywords and constructs the information necessary to visualize the details of electronic transactions. This information is sent to the user's terminal and displayed on the screen. The input is a list of important keywords, and the output is transaction information for display.
[0787] Step 5:
[0788] The server analyzes the audio data using an AI model and adjusts the sound quality to suit the hearing characteristics of elderly individuals. This adjusted audio data is then sent to the terminal and played back. The input is the original audio data, and the output is optimized audio data. This allows the user to receive clear audio feedback.
[0789] Step 6:
[0790] The terminal simultaneously presents the user with transaction information for display and pre-recorded audio. The user can verify transaction details both visually and aurally. Input consists of transaction information for display and pre-recorded audio data, while the visual display and audio output constitute the output.
[0791] Step 7:
[0792] Once the user confirms the details and indicates their intention to execute the transaction, the details are sent back from the terminal to the server, and the electronic transaction is completed. The input is the confirmed transaction information, and the output is a transaction success message. This ensures that electronic transactions are conducted safely and smoothly.
[0793] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0794] This invention provides a voice communication support system for the elderly that incorporates an emotion engine that recognizes the user's emotions in addition to voice analysis. This system consists of a user terminal and a server, which are connected via a network.
[0795] The user makes a call using their device, and the audio is captured by a voice input device. The captured audio data is digitized and sent to a server. The server analyzes the audio data using a voice analysis device and understands the characteristics of the audio using a generative AI model. Based on this, the server adjusts the volume and sound quality to make it easy for the user to hear.
[0796] This system is equipped with an emotion engine, and the server analyzes the user's emotional state from their voice in real time. The emotion engine analyzes changes in intonation, tempo, and volume of the voice to infer the user's emotions. The resulting emotional state is then used for further optimization in voice adjustment.
[0797] For example, if the emotion engine determines that the user is stressed, the server adjusts the voice to a calmer tone and further reduces background noise. The adjusted voice is then sent back from the server to the terminal and delivered to the user as clear audio.
[0798] Furthermore, along with the audio data transcribed into text by the text analysis method, emotional information analyzed by the emotion engine is displayed on the user interface. Specifically, simple feedback corresponding to the emotional state is displayed on the screen, allowing the user to obtain information based on their own emotional state. For example, if the user's emotions are analyzed as being calm, a message such as "Relaxed" will be displayed.
[0799] This system will allow elderly individuals to enjoy more fulfilling communication not only through voice but also through emotional feedback. The introduction of an emotion engine will enable individualized adaptation that takes into account the user's emotions, which is expected to further improve the quality of communication.
[0800] The following describes the processing flow.
[0801] Step 1:
[0802] The device captures the user's call audio using a microphone and converts the analog signal into digital audio data. The audio data is then encoded into a format that allows for real-time processing.
[0803] Step 2:
[0804] The terminal sends the converted audio data to the server. The transmission is performed using a secure protocol to minimize latency.
[0805] Step 3:
[0806] When the server receives audio data, it first uses an audio analysis tool to analyze the audio frequency and volume. This analysis provides the basic data needed to adjust the sound quality to make it easier for the user to hear.
[0807] Step 4:
[0808] The server uses an emotion engine to analyze the emotional nuances contained in the speech. It analyzes the intonation, tempo, and volume fluctuations of the speech to infer the user's emotional state.
[0809] Step 5:
[0810] The server adjusts the volume and frequency of the audio based on the results of voice analysis and emotion analysis. The adjustments are adapted to the user's auditory characteristics and estimated emotional state.
[0811] Step 6:
[0812] The server sends back audio, adjusted to a sound quality that allows the user to relax, to the terminal via an audio output device. The terminal then plays the adjusted audio clearly.
[0813] Step 7:
[0814] The server converts the audio data into text, identifying and extracting important keywords. The extracted text information, along with emotional feedback, is sent to the terminal.
[0815] Step 8:
[0816] The device displays received text and sentiment information on its screen. Users can visually confirm key points of the conversation and receive feedback on their own emotional state.
[0817] Step 9:
[0818] By reviewing the displayed information and understanding parts that were difficult to hear or their own emotional state, users can facilitate smoother communication.
[0819] (Example 2)
[0820] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0821] When elderly people communicate with others via voice, the sound quality and volume may be insufficient, or the user's emotional state may not be understood, leading to communication that proceeds without proper understanding. This can hinder smooth communication and result in an unsatisfactory experience. Therefore, there is a need for technology that improves sound quality and provides feedback that responds to the user's emotions, thereby enriching the communication experience.
[0822] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0823] In this invention, the server includes means for receiving digitized speech and analyzing speech characteristics in real time using a generation AI model; a speech adjustment device that adjusts the volume and sound quality based on the analyzed speech characteristics and emotional state to convert it into an easily audible sound; and a text processing device that generates text from the speech and extracts important information. As a result, users can receive clearer and easier-to-understand speech and also receive feedback tailored to their emotional state.
[0824] "Elderly people" generally refers to people who are older, and in this invention, they are a group that requires special consideration when communication is primarily conducted via voice.
[0825] "Digitalization" is the process of sampling analog audio signals and converting them into digital signals, and it is a fundamental step that enables data processing for the entire system.
[0826] A "generative AI model" is a type of artificial intelligence used to analyze speech characteristics, designed to perform intelligible speech conversion and emotion analysis in real time.
[0827] "Sound characteristics" refer to the physical and sensory characteristics of an audio signal, and specifically include volume, sound quality, frequency spectrum, intonation, tempo, etc.
[0828] A "voice adjustment device" is a device that adjusts voice to an optimal state based on analyzed voice characteristics and emotional state, thereby allowing users to receive voices that are easier to hear.
[0829] A "text processing device" is a device that converts audio signals into text and extracts important information, making it useful for providing information to users.
[0830] "Feedback" is information provided based on the user's voice and emotional state, and is intended to help users understand their own expressions and facilitate smooth communication.
[0831] This system is designed to help elderly people communicate smoothly and comfortably using voice. The system mainly consists of terminals, servers, and various software components.
[0832] The user initiates a voice call using their device. The device has a built-in microphone that captures the user's voice as an analog signal. The captured audio is digitized by the device, sampled, and converted into a digital audio format. This digitized audio data is then transmitted to the server via the network.
[0833] The server processes the received audio data using a generative AI model. Specifically, it analyzes the characteristics of the audio in real time and performs adjustments such as noise reduction and volume adjustment. Furthermore, the server analyzes the intonation and tempo of the audio using its built-in emotion engine and infers the user's emotional state. The results of this emotion analysis are also fed back into the audio adjustment process, generating the optimal audio for the user.
[0834] The server also converts the audio data into text and extracts important information. This text, along with feedback on the emotional state, is displayed in the user interface, allowing the user to understand their own emotional state while communicating.
[0835] As a concrete example, here is an example of a prompt sentence to give to a generative AI model: "Based on the audio data the user is speaking, use the emotion engine to analyze the emotions and provide appropriate feedback."
[0836] This system is expected to improve the quality of communication among the elderly by enabling higher-quality voice communication and providing more comprehensive feedback based on emotional understanding.
[0837] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0838] Step 1:
[0839] The user initiates a call using the device. The device's microphone captures the user's speech and recognizes it as an analog audio signal. The device then samples and quantizes this input audio, converting it into a digital audio format. In this process, the analog signal is transformed into digital data. Digital audio data is generated and prepared for the next processing step.
[0840] Step 2:
[0841] The terminal transmits digitized audio data to the server via the network. The server uses the received audio data as input to run a generating AI model, which analyzes the audio characteristics in real time. Data calculations such as volume, sound quality, and frequency spectrum are performed, and parts that require noise reduction or volume adjustment are identified. The analysis results output the points that need adjustment and the adjusted audio characteristics.
[0842] Step 3:
[0843] The server works in conjunction with a generative AI model to activate the emotion engine. It analyzes changes in intonation, tempo, and volume from the speech to estimate the user's emotional state. The analyzed emotional state is detected as input, and the emotion is output as a state such as stress or relaxation. This information is used for subsequent speech adjustments.
[0844] Step 4:
[0845] The server activates an audio adjustment device based on the analyzed voice characteristics and emotional state. Specifically, it adjusts the voice to make it easier to hear by applying noise cancellation and volume adjustments. For example, if stress is detected, it processes the voice to calm it down. This process generates a clear voice, which is then prepared for transmission.
[0846] Step 5:
[0847] The server sends the adjusted audio data back to the terminal. The terminal provides the received audio to the user through an audio output device. Here, it is output as the audio the user actually hears. The server also transcribes the audio data into text, extracts important information, and supplies it to the user interface display as feedback information. This allows the user to access information including emotional feedback.
[0848] Step 6:
[0849] Users refer to text information and sentiment feedback displayed on the device's screen. This allows them to understand their own expressions, obtain the information necessary for communication, and use it for the next step. By reviewing the displayed feedback information, they can gain insights to communicate effectively.
[0850] (Application Example 2)
[0851] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0852] Improving the quality of voice communication is crucial for the elderly and those with limited communication abilities. However, in addition to ensuring clear voice quality, there is a lack of systems that can accurately grasp the emotional state of the other party and respond appropriately as needed. Furthermore, real-time emotion recognition is required to detect and respond to potential security risks arising from emotional shifts at an early stage. Solving these challenges is essential.
[0853] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0854] In this invention, the server includes an acoustic input means, an acoustic analysis means, an adjustment means, an emotion recognition means, and an alarm means. This makes it easy for the elderly to use and enables real-time detection and response to security risks based on emotions.
[0855] An "acoustic input device" is a device that acquires voice signals emitted by elderly people and converts them into voice information.
[0856] An "acoustic analysis device" is a device that analyzes received audio information in real time using a generation model to understand acoustic characteristics and emotional states.
[0857] A "adjustment device" is a device that adjusts the volume and sound quality based on the results of acoustic analysis, converting the sound into one that is easy for elderly people to hear.
[0858] "Audio output means" refers to a device for providing adjusted sound to the user.
[0859] A "language analysis device" is a device that has the function of extracting important words from audio information and displaying that information on the user's screen.
[0860] A "display device" is a device that visually shows analyzed emotional information and important words on the user's screen.
[0861] An "emotion recognition device" is a device that analyzes the characteristics of speech from speech information obtained by an acoustic analysis device to estimate an emotional state.
[0862] An "alarm device" is a device that has the function of detecting potential security risks and providing necessary notifications based on an estimated emotional state.
[0863] The system of this invention supports voice communication for the elderly and provides emotion-based adaptive feedback and security features. A detailed embodiment of this system is shown below.
[0864] First, the terminal uses an acoustic input device to acquire an audio signal from the user and converts it into digital data as audio information. This audio information is then sent to the server.
[0865] The server uses acoustic analysis tools and generative AI models to analyze speech information in real time. Specifically, the acoustic analysis engine analyzes changes in speech intonation, tempo, and volume, and estimates the user's emotional state through emotion recognition tools. As generative AI models, for example, Google Cloud's Speech-to-Text API and Emotion Recognition API can be used.
[0866] Next, the adjustment means adjusts the volume and sound quality based on the analyzed information to create a sound that is easy for the user to hear. This adjusted sound is then provided to the user by the sound output means.
[0867] Furthermore, the language analysis tool extracts important words from the audio information and displays them visually on the user's screen via the display tool. Simultaneously, emotional information is also displayed on the screen, allowing the user to understand their own emotional state.
[0868] The alarm system activates when the estimated emotional state indicates a certain level of security risk. Specifically, when the emotional state is determined to be stressful, it sends a notification to registered contacts using communication methods such as Twilio or SendGrid.
[0869] This system allows elderly individuals to safely enjoy communication not only through voice but also through emotional feedback. For example, when an elderly person feels lonely, their family is notified, prompting them to take care of them.
[0870] Examples of prompt messages include the following:
[0871] "Analyze the current emotional state and determine whether the elderly person's emotions are calm."
[0872] "Use the user's voice data to detect their stress levels and send alerts if necessary."
[0873] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0874] Step 1:
[0875] The terminal acquires the user's voice using an acoustic input device. The input is the user's voice signal, which the terminal converts into digital data. The converted digital audio data is then sent to the server.
[0876] Step 2:
[0877] The server uses speech analysis tools to analyze digital speech data through a generative AI model. The input is digital speech data, and the server analyzes changes in intonation, tempo, and volume. As a result of the analysis, acoustic characteristics and emotional states are extracted.
[0878] Step 3:
[0879] The server adjusts the volume and sound quality based on the analysis results using adjustment mechanisms. The input is acoustic characteristics, and the server adjusts the sound to be easily audible to the user, generating clear, adjusted audio data. The adjusted audio data is then transmitted to the terminal.
[0880] Step 4:
[0881] The terminal outputs adjusted audio to the user using an audio output means. The input is adjusted audio data, allowing the user to hear clear audio.
[0882] Step 5:
[0883] The server extracts important words from audio information using language analysis tools. The input is audio information, and the server performs language analysis to extract important words. Once important words are extracted, the data is transmitted to the terminal via a display device.
[0884] Step 6:
[0885] The terminal uses a display mechanism to show important words and estimated sentiment information on the user's screen. The input consists of extracted important words and sentiment information, allowing the user to visually confirm their emotional state.
[0886] Step 7:
[0887] The server utilizes emotion recognition mechanisms to detect potential security risks based on estimated emotional states. If the emotion indicates stress or danger, it uses that state as input to send a notification command to the alarm system.
[0888] Step 8:
[0889] The server uses alarm mechanisms to notify pre-configured emergency contacts via communication means as needed. The input is information about detected security risks, a notification is output, and relevant parties are prompted to take prompt action.
[0890] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0891] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0892] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0893] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0894] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0895] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0896] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0897] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0898] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0899] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0900] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0901] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0902] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0903] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0904] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0905] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0906] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0907] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0908] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0909] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0910] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0911] The following is further disclosed regarding the embodiments described above.
[0912] (Claim 1)
[0913] A voice input means that captures the voice signals when elderly people converse and converts them into voice data,
[0914] A voice analysis means that receives the aforementioned voice data and performs voice analysis in real time using an artificial intelligence model,
[0915] An adjustment means that adjusts the volume and sound quality based on the analysis results and converts it into an easy-to-hear voice,
[0916] Audio output means for outputting the adjusted audio to the user,
[0917] A text analysis means for extracting important words from the aforementioned audio data,
[0918] A display means for displaying the aforementioned important terms on the user interface,
[0919] A system that includes this.
[0920] (Claim 2)
[0921] The system according to claim 1, wherein the voice analysis means comprises a voice filtering function for reducing background noise.
[0922] (Claim 3)
[0923] The system according to claim 1, wherein the adjustment means has a function to learn the user's auditory characteristics and individually optimize the adjustment of the sound.
[0924] "Example 1"
[0925] (Claim 1)
[0926] An input device that collects audio information and converts it into digital data,
[0927] An analysis device that receives the aforementioned digital data and analyzes it in real time using an artificial intelligence model,
[0928] An adjustment device that adjusts the volume and sound quality based on the analysis results and converts it into an easy-to-listen-to sound,
[0929] An output device that provides the aforementioned adjusted audio to an individual,
[0930] An analysis device for extracting key terms from the aforementioned digital data,
[0931] A display function that shows the aforementioned main terms on a display device,
[0932] A system that includes this.
[0933] (Claim 2)
[0934] The system according to claim 1, wherein the analysis device has an audio processing function that reduces ambient noise.
[0935] (Claim 3)
[0936] The system according to claim 1, wherein the adjustment device has a function to learn the user's auditory characteristics and individually optimize the adjustment of the sound.
[0937] "Application Example 1"
[0938] (Claim 1)
[0939] A voice acquisition method that captures the voice signals when elderly people converse and converts them into voice information,
[0940] A voice analysis means that receives the aforementioned voice information and performs voice analysis in real time using a generation AI model,
[0941] A voice adjustment means that adjusts the volume and sound quality based on the analysis results and converts it into an easy-to-listen voice,
[0942] Audio playback means for outputting the adjusted audio to the user,
[0943] A text analysis means for extracting important words from the aforementioned audio information,
[0944] Information presentation means for displaying the aforementioned important terms on the user interface,
[0945] An electronic transaction support system that receives voice commands and processes payment information,
[0946] A display support means that converts the aforementioned payment information into text and displays important content,
[0947] A system that includes this.
[0948] (Claim 2)
[0949] The system according to claim 1, wherein the voice analysis means comprises a voice processing function for reducing background noise.
[0950] (Claim 3)
[0951] The system according to claim 1, wherein the voice adjustment means has a function to learn the user's auditory characteristics and individually optimize the voice adjustment.
[0952] "Example 2 of combining an emotion engine"
[0953] (Claim 1)
[0954] A device that captures and digitizes the voices of elderly people when they speak,
[0955] A processing device that receives the digitized audio and analyzes the audio characteristics in real time using a generation AI model,
[0956] A sound adjustment device that adjusts the volume and sound quality based on the analyzed voice characteristics and emotional state, and converts the sound into a sound that is easy to hear,
[0957] An audio output device that outputs the adjusted audio to the user,
[0958] A text processing device that generates text from the aforementioned audio and extracts important information,
[0959] A device that displays the aforementioned text and sentiment analysis results on a user display device,
[0960] A system that includes this.
[0961] (Claim 2)
[0962] The system according to claim 1, wherein the processing device comprises an emotion analysis function for analyzing the emotional state of speech.
[0963] (Claim 3)
[0964] The system according to claim 1, wherein the voice adjustment device has a function to dynamically optimize voice adjustment based on the analyzed emotional state.
[0965] "Application example 2 when combining with an emotional engine"
[0966] (Claim 1)
[0967] An acoustic input means that acquires voice signals when elderly people converse and converts them into voice information,
[0968] Acoustic analysis means that receives the aforementioned audio information and performs audio analysis in real time using a generation model,
[0969] An adjustment means that adjusts the volume and sound quality based on the analysis results and converts it into a sound that is easy to hear,
[0970] The aforementioned adjusted sound is provided to the user via an acoustic output means,
[0971] A language analysis method for extracting important words from audio information,
[0972] A display means for displaying the aforementioned important words and analyzed sentiment information on the user's screen,
[0973] The aforementioned acoustic analysis means includes an emotion recognition means that estimates an emotional state from speech using an emotion engine,
[0974] An alarm means that detects and notifies of potential security risks based on the aforementioned emotional state,
[0975] A system that includes this.
[0976] (Claim 2)
[0977] The system according to claim 1, wherein the acoustic analysis means comprises an acoustic filtering function for reducing background noise.
[0978] (Claim 3)
[0979] The system according to claim 1, wherein the adjustment means has a function to learn the user's auditory characteristics and individually optimize the sound adjustment. [Explanation of Symbols]
[0980] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A voice input means that captures the voice signals when elderly people converse and converts them into voice data, A voice analysis means that receives the aforementioned voice data and performs voice analysis in real time using an artificial intelligence model, An adjustment means that adjusts the volume and sound quality based on the analysis results and converts it into an easy-to-hear voice, Audio output means for outputting the adjusted audio to the user, A text analysis means for extracting important words from the aforementioned audio data, A display means for displaying the aforementioned important terms on the user interface, A system that includes this.
2. The system according to claim 1, wherein the voice analysis means is equipped with a voice filtering function for reducing background noise.
3. The system according to claim 1, wherein the adjustment means has a function to learn the user's auditory characteristics and individually optimize the adjustment of the sound.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A