system
The communication system addresses hearing loss and speed differences by converting audio to text and adjusting audio tone and speed, improving communication quality for the elderly.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-15
- Publication Date
- 2026-04-27
AI Technical Summary
Communication between the elderly and their families is hindered by age-related hearing loss and differences in conversation speed, leading to one-sided conversations and misunderstandings, with conventional hearing aids and speech recognition systems failing to provide real-time voice speed adjustment and visual support.
A communication system that includes a display device for converting audio to text, a processing device for adjusting audio speed and tone, and a communication device for wireless transmission to hearing aids, enabling real-time audio adjustment and visual confirmation of text.
Enhances communication quality by allowing the elderly to hear adjusted audio in real time and visually confirm the text, making conversations more effective and comfortable.
Smart Images

Figure 2026070214000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the communication between the elderly and their families, there is a problem that smooth conversation becomes difficult due to age-related hearing loss and differences in conversation speed. As a result, one-sided conversations and misunderstandings are likely to occur, increasing the number of stressful situations for both parties. In addition, conventional hearing aids and speech recognition systems do not sufficiently provide real-time voice speed adjustment and visual support for text, resulting in a problem of deteriorating communication quality.
Means for Solving the Problems
[0005] This invention provides a communication system equipped with a display device that acquires audio data, converts it into text data, and then displays it. Furthermore, it also includes a communication device that adjusts the speed and tone of the audio data to make it easier for the user to hear, and transmits it to a hearing aid via wireless communication. As a result, the user can hear the adjusted audio in real time and visually confirm it as text. The combination of adjusted audio and text display solves conventional problems and enables more effective and comfortable communication.
[0006] "Audio data" refers to data that records the voice spoken by a user in digital format.
[0007] A "device" is a hardware device used to acquire audio data.
[0008] "Text data" refers to character information obtained by converting audio data using speech recognition technology.
[0009] A "processing device" is a computer unit that has the function of converting audio data into text data.
[0010] A "display device" is a device such as a screen or monitor that visually displays converted text data.
[0011] "Speed" refers to the temporal speed at which audio data is played back.
[0012] "Tone" refers to the element that controls the pitch and quality of sound in audio data.
[0013] A "modification device" is a device that has the function of changing the speed and tone of audio data.
[0014] A "communication device" is hardware or software used to transmit adjusted audio data to other devices or hearing aids.
[0015] A "system" is a series of elements configured to achieve a certain purpose by the cooperation of multiple devices and functions.
[0016] "Wireless communication" is a means of transmitting information via radio waves without using wires such as cables.
[0017] A "hearing aid" is a portable device that amplifies sound and provides it to a user to assist with hearing impairments.
Brief Description of the Drawings
[0018] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12]It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.
Mode for Carrying Out the Invention
[0019] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0020] First, the language used in the following description will be explained.
[0021] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0022] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0023] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0024] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0026] [First Embodiment]
[0027] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0028] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0031] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0034] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0038] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0039] This invention is a system for facilitating smooth communication between the elderly and their families. The invention mainly consists of five main components: a voice input device, a processing device, a display device, a regulating device, and a communication device.
[0040] First, when a user speaks into a voice input device, the device acquires voice data. This voice data is converted into a digital format and sent to a processing unit. The processing unit uses a voice recognition module located in the cloud or within the device to convert the voice data into text data. This text data is sent to a display device and presented visually according to user settings such as font size and color.
[0041] Simultaneously, the server optimizes the speed and tone of the audio data to individual needs using an adjustment device. Specifically, it slows down the speed to make the audio easier to understand and adjusts the tone of the voice to maintain a natural sound. The new audio data generated by this adjustment process is transmitted wirelessly to the user's hearing aids via a communication device. This allows the user to hear the adjusted audio in real time.
[0042] For example, when a user says "Let's have lunch together" into a voice input device, the device receives the voice, and the processing unit converts it into the text "Let's have lunch together." This text is then displayed on the display device with an adjusted font size. Meanwhile, the server slows down the voice speed by 30%, and this slowed-down voice is sent to the hearing aid, making it easier for elderly people to hear.
[0043] This invention makes everyday conversations with elderly people who have difficulty with voice communication more comfortable and enables communication that is easier for both parties to understand.
[0044] The following describes the processing flow.
[0045] Step 1:
[0046] The user begins speaking into the device's microphone. The device captures this voice and temporarily stores it as audio data in its internal memory.
[0047] Step 2:
[0048] The device sends the stored audio data to a server in the cloud for processing. The data is transmitted securely using a communication protocol.
[0049] Step 3:
[0050] The server inputs the received audio data into a speech recognition algorithm, converting the audio into text data. This process is performed using a speech recognition API.
[0051] Step 4:
[0052] The converted text data is sent from the server to the terminal. The terminal adjusts the font size and color of the received text according to the user's settings and displays it on the screen.
[0053] Step 5:
[0054] The server adjusts the speed and tone of the original audio data to convert it into easily understandable speech. For this purpose, it uses an audio processing algorithm.
[0055] Step 6:
[0056] The adjusted audio data is sent from the server to the terminal, which then streams it to the user's hearing aids via Bluetooth.
[0057] Step 7:
[0058] Users receive calibrated audio through their hearing aids and understand conversations in real time. Along with displayed text, their listening comprehension is improved.
[0059] (Example 1)
[0060] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0061] In recent years, there has been an increasing number of cases where communication between the elderly and their families is difficult due to hearing and visual limitations. Many current speech recognition technologies and communication systems have limitations in supporting real-time, two-way communication, and a particular lack of systems that provide coordinated audio and visual displays simultaneously has been pointed out. There is a need to solve these problems and provide systems that enable smooth communication for everyone.
[0062] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0063] In this invention, the server includes means for acquiring audio information, means for converting the audio information into text information, and means for visually displaying the text information. This makes it possible to provide auditory support and visual information to make it easier for elderly people to understand audio in real time.
[0064] "Means for acquiring voice information" refers to a device that has the function of capturing voice signals emitted by a user as digital data.
[0065] "Device for converting audio information into text information" refers to a device that performs the process of converting acquired audio information into a string of characters using an algorithm.
[0066] "Means for visually displaying textual information" refers to a device that enables the visual presentation of converted textual information using a display device such as a screen.
[0067] "Devices and means for adjusting the speed and pitch of audio information" refers to devices that have the function of adjusting the playback speed and pitch of audio in order to facilitate the understanding of the audio.
[0068] "Communication means for transmitting adjusted audio information" refers to a device that has a communication function for transmitting adjusted audio data to other devices, particularly hearing assistance devices.
[0069] "Hearing assistance devices" refer to assistive devices that amplify or adjust sound to provide audio information to users with hearing impairments.
[0070] "Display font size and color according to user settings" refers to a function that adjusts the font size and color tone according to the user's preferences.
[0071] This invention is a system for facilitating smooth voice communication between the elderly and their families. The system mainly consists of a device for acquiring voice information, a device for converting voice information into text information, a device for visually displaying the text information, a device for adjusting the speed and tone of the voice information, and a communication device for transmitting the adjusted voice information.
[0072] The user sends a message using a voice input device. The terminal receives this audio and converts it into a digital format. This process uses hardware such as a microphone and an A / D converter. The audio data is sent from the terminal to the processing unit, which uses speech recognition software to convert the audio into text. Specifically, a cloud platform that provides speech recognition services may be used as the software.
[0073] Next, the generated text information is sent to a display device within the terminal and visualized using large fonts and specified colors. This makes it easy for elderly people with visual impairments to recognize the content.
[0074] Simultaneously, the server adjusts the speed and tone of the audio data via a tuning device to suit individual needs. The adjusted audio is played back at the optimal speed and tone to make it easier for elderly people to understand.
[0075] The server then transmits the adjusted audio information to the user's hearing aid using wireless communication. This uses communication technologies such as Bluetooth or Wi-Fi.
[0076] As a concrete example, suppose a user says "Let's have lunch together" into a voice input device. The device receives the voice and a processing unit converts it into text information: "Let's have lunch together." This text information is displayed in a large font on the display device, while the server slows down the voice speed by 30% to produce easier-to-understand audio. This adjusted audio is sent to a hearing aid, allowing elderly people to hear the adjusted audio in real time.
[0077] In the generative AI model, the prompt phrase "How to use voice input devices to improve communication with the elderly" is used. This prompt serves as a guide to improve the performance of speech recognition and speech synthesis by the generative AI model.
[0078] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0079] Step 1:
[0080] When a user speaks into a voice input device, voice data is input. The terminal receives this voice data and performs a process to convert the analog voice into a digital format. An analog-to-digital converter is used for this conversion. The resulting digital voice data is then output.
[0081] Step 2:
[0082] The terminal transmits digital audio data to the processing unit. The processing unit performs speech recognition using a generative AI model and generates text data from the audio data. Specifically, it takes audio waveform data as input, analyzes it at the phoneme level, and outputs text data by performing the optimal text conversion.
[0083] Step 3:
[0084] A terminal that receives character data from a processing unit transmits that data to a display device. The display device adjusts the font size and color according to user settings and displays it as visual information. Based on the input character data, character information processed into a format that is easy for the user to read is output.
[0085] Step 4:
[0086] The server receives the digital audio data and adjusts the speed and tone of the audio. This process applies algorithms that stretch the time axis and audio filters to generate easy-to-understand audio data. The adjusted audio data is then output.
[0087] Step 5:
[0088] The server wirelessly transmits the adjusted audio data to the user's hearing aid via a communication device. Wireless technologies such as the Bluetooth protocol are used for communication. Finally, the outputted audio is played back in real time from the hearing aid, enabling elderly people to communicate easily.
[0089] (Application Example 1)
[0090] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0091] This solution addresses the problem of elderly and hearing-impaired individuals experiencing difficulties in shopping and receiving services due to insufficient information, hindering smooth communication. Furthermore, it is necessary to prevent confusion and decreased customer satisfaction caused by inadequate and timely provision of store information.
[0092] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0093] In this invention, the server includes means for acquiring audio data, means for converting the audio data into text data, and means for providing the text data and the adjusted audio data to the guidance device. This makes it possible for elderly customers and customers with hearing impairments to receive guidance information in a more easily understandable format in physical stores.
[0094] "Means for acquiring audio data" refers to a device that collects audio using an input device and processes it as a digital signal.
[0095] "Means of converting to text data" refers to a device that performs the process of analyzing audio data and converting it into textual information.
[0096] "Means of display" refers to a display device for visually presenting the converted text data.
[0097] "Means for adjusting the speed and tone of audio data" refers to a device that performs the process of adjusting the playback speed and tone of audio to suit individual needs.
[0098] "Means for transmitting adjusted audio data" refers to a device for transmitting adjusted audio data wirelessly or via a wired connection to an auxiliary device or receiving terminal.
[0099] "Means of providing information to the guidance device" refers to a communication device for providing necessary information to guidance equipment within the store.
[0100] This system enables smooth communication for the elderly and those with hearing impairments in physical stores by using voice input devices, processing devices, display devices, adjustment devices, and communication devices. The system is centered around a server and is configured as follows:
[0101] The server acquires voice data from customers in the store via voice input devices. The acquired voice data is converted into text data using speech recognition software such as Google® Cloud Speech-to-Text. The converted text data is presented as visual information using a display device. The display device uses e-ink displays or LCD panels, and is displayed in a size and color that takes user visibility into consideration.
[0102] Furthermore, the adjustment device uses audio processing software such as Adobe Audition and Librosa to adjust the speed and tone of the audio data, providing audio optimized for hearing. This adjusted audio data is transmitted to the assistive device using Bluetooth or Wi-Fi. This creates an environment where elderly or hearing-impaired customers can clearly understand the information.
[0103] As a concrete example, if an elderly person in a store uses voice input to ask, "Where is this item located?", the terminal captures this voice, and the server immediately processes the data to provide both visual and auditory guidance. Customers can receive guidance both displayed on a large screen and audio guidance through assistive devices.
[0104] An example of an input prompt for a generated AI model is "a method to display in real time the content of a question asked by an elderly person in a store and send an audio guide to an assistive device." This is expected to improve the accuracy of guidance and usability.
[0105] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0106] Step 1:
[0107] The server acquires voice data from the user through the voice input device. The input voice is converted into digital data using a microphone. The voice input device then prepares to capture the user's voice and send it to the server.
[0108] Step 2:
[0109] The server sends the acquired audio data to the processing unit for speech-to-text conversion. Here, a speech recognition service such as Google Cloud Speech-to-Text is called to convert the audio data into text data. The input is audio data, and the output is text data. At this stage, the audio input is processed to accurately convert it into text.
[0110] Step 3:
[0111] The server sends the converted text data to a display device, where it is displayed as visual information. The display is configured to show the received text using a highly legible font and color. The input is text data, and the output is visual information that the user can confirm. This allows the user to verify the information provided.
[0112] Step 4:
[0113] The server sends audio data to an adjustment device to adjust the speed and timbre of the audio. This adjustment uses software such as Librosa or Adobe Audition, and aims to clarify the sound through data processing. The input is the original audio data, and the output is the adjusted audio data. At this stage, audio optimized for auditory perception is generated.
[0114] Step 5:
[0115] The server transmits the adjusted audio data to the assistive device via a communication device. Wireless communication using Bluetooth or Wi-Fi is then used to deliver the audio to the user's assistive device. The input is the adjusted audio data, and the output is the audio heard by the user. The user can auditorily understand the guidance content by receiving the adjusted audio guidance.
[0116] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0117] This invention relates to a communication system that not only acquires voice data and converts it into text data, but also recognizes the user's emotions. The system includes a voice input device, a processing device, a display device, a regulating device, a communication device, and an emotion engine.
[0118] When a user speaks into a voice input device, the device collects voice data. This voice data is sent to a server for processing and converted into text data by a speech recognition algorithm. Similarly, an emotion engine extracts emotional information from the user's voice. This emotional information influences the display of the converted text. For example, if the user is excited, visual feedback such as a change in text color can be provided.
[0119] The server uses the emotion engine's output to automatically adjust the speed and tone of the audio data, and then transmits this adjusted audio data to the user's hearing aids via a communication device. This process results in more natural and easily understandable communication for the listener.
[0120] For example, if a user smiles and says, "I'm looking forward to tomorrow's presentation," the device will capture the audio, and the processing unit will transcribe it into text, "I'm looking forward to tomorrow's presentation," while also recognizing the emotion of "happiness." Based on this information, the display device will show the text in a brighter color, and the server will adjust the tone of the audio slightly higher before sending it to the hearing aid, thereby conveying the user's emotions more richly.
[0121] This invention provides a more personalized experience by comprehensively capturing the user's emotions and adjusting the entire communication accordingly. This enables smoother dialogue between the elderly and their families, overcoming the limitations of conventional voice communication.
[0122] The following describes the processing flow.
[0123] Step 1:
[0124] The user speaks into the device's microphone. The device captures this voice and temporarily stores it as digital audio data.
[0125] Step 2:
[0126] The terminal sends the acquired audio data to the server. On the server side, the audio data is processed and converted into text data using speech recognition technology.
[0127] Step 3:
[0128] The server inputs audio data into the emotion engine before and after speech recognition, analyzing the user's emotions from their voice. This emotion data is then used for subsequent processing.
[0129] Step 4:
[0130] The server sends the converted text data back to the terminal, simultaneously adding information based on the user's emotions. The terminal receives this text data and displays it on the screen, visually adjusting the font size and color according to the emotions.
[0131] Step 5:
[0132] The server adjusts the speed and tone of the voice data based on the output of the emotion engine. Emotion-appropriate voice emphasis is applied to create a more natural conversation.
[0133] Step 6:
[0134] The adjusted audio data is sent from the server to the terminal, which then streams this audio to the user's hearing aids via Bluetooth.
[0135] Step 7:
[0136] Through the hearing aid, users receive both calibrated audio and visually calibrated text data, enabling them to understand the content and emotions of conversations in real time.
[0137] (Example 2)
[0138] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0139] With the increasing diversity of modern audio content, users are seeking richer communication that includes not only the content of the audio data but also the speaker's emotions and tone. However, conventional audio data conversion technologies and display methods have been unable to adequately reflect these emotions and tones, resulting in limited information transmission to users. Against this backdrop, there is a need to develop a new system that not only converts audio data to text but also adjusts the audio while considering emotional information.
[0140] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0141] In this invention, the server includes means for acquiring voice information, processing means for converting the voice information into a data format, and means for extracting emotional information from the voice information. This makes it possible to understand the emotions and tone of the voice spoken by the user and to provide voice and display adjusted based on that information.
[0142] "Audio information" refers to data that is created by converting the waveform of sound emitted by a user into an electrical signal and representing it in digital format.
[0143] "Data format" refers to the state in which audio information has been converted into text or other processable digital information.
[0144] "Processing means" means a device or process that acquires audio information, converts it into a data format, and performs other necessary processing.
[0145] "Emotional information" refers to data that analyzes audio information and indicates the speaker's psychological and emotional state.
[0146] "Display means" refers to a device or method for visually displaying digital information.
[0147] "Adjustment means" refers to a process or device that modifies the speed and timbre of audio information to enable more appropriate transmission.
[0148] "Communication means" means technology or equipment for transmitting coordinated voice information to other devices or equipment.
[0149] A "hearing assistance device" refers to a device that outputs audio data in a format suitable for hearing, helping people with hearing impairments to hear sounds.
[0150] "User" refers to an individual or group that uses the system.
[0151] "Display characteristics" refer to elements such as font size and color that visually represent data formats.
[0152] This invention relates to a system that uses a specific voice processing algorithm to effectively analyze a user's voice information and obtain desired information. This system converts the voice information into text format and further extracts emotional information, enabling display and sound adjustment.
[0153] Users can speak to the system through a voice input device. This voice input device consists of a terminal including a standard microphone. The terminal sends the acquired voice information to a server. The server uses speech recognition software to convert the voice information into a data format. Tools such as the Google Cloud Speech-to-Text API are used for this process. Furthermore, the server uses an emotion recognition system to extract emotion information from the voice information. A general-purpose emotion analysis engine is a possible technology used here.
[0154] Based on the extracted emotional information, the display device on the terminal shows text data to the user, visually representing emotions with appropriate colors and font sizes. The server also processes the speed and tone of the audio information and transmits the adjusted audio information to the hearing assistance device via wireless communication. This makes the audio clearer and easier to understand.
[0155] For example, if a user says, "I'm looking forward to tomorrow's presentation," the server transcribes this phrase into text and further recognizes the positive emotion. The display device shows this text in a brighter color, and the server adjusts the tone of voice slightly higher before transmitting it.
[0156] An example of a prompt message is: "Based on the voice information obtained from the user, please extract emotional information and provide it as text data. Also, please suggest adjustments to the text display color and voice tone." This enables emotionally rich and intuitive communication for the user.
[0157] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0158] Step 1:
[0159] The user speaks through a voice input device. The terminal uses a microphone to acquire the user's voice information. The input is the user's voice waveform, and the output is digital voice data. This voice data is then temporarily stored for subsequent processing.
[0160] Step 2:
[0161] The terminal sends the acquired audio data to the server. Here, the terminal uses an internet connection to transfer data to the server. The input is audio data in digital format, and the output is the data sent to the server. This process involves specific actions using protocols for secure and high-speed data transfer.
[0162] Step 3:
[0163] The server uses speech recognition software to convert audio data into text format. The input is digital audio data received by the server, and the output is text data. In this process, the server uses the Google Cloud Speech-to-Text API and other tools to analyze the characteristics of the audio waveform and perform the specific operations of converting it into text.
[0164] Step 4:
[0165] The server uses an emotion recognition engine to extract emotional information from audio data. The input is digital audio data received by the server, and the output is an emotion label (e.g., "happy," "sad"). The server performs specific actions to identify the speaker's emotions by analyzing the tone, pitch, and speed of the speech.
[0166] Step 5:
[0167] The terminal processes data to change how text data is displayed based on emotional information. Input consists of text data and emotional information, while output is a visually altered text display. The display device performs specific actions, such as changing the color of the text, according to the emotional information.
[0168] Step 6:
[0169] The server adjusts the speed and tone of the audio data based on emotional information. The input is the original audio data and emotional information, and the output is the adjusted audio data. The server uses audio editing software to perform the specific actions required to adjust the tone of the audio.
[0170] Step 7:
[0171] The server transmits the adjusted audio data to the user's hearing aid using a communication method. The input is the adjusted audio data, and the output is the audio received by the user's hearing aid. The server performs the specific operation of transmitting the audio using a wireless communication protocol.
[0172] (Application Example 2)
[0173] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0174] In face-to-face communication at physical stores, there is a problem in immediately grasping customer emotions and reflecting them in service. With conventional methods, voice data is limited to simple text conversion, and there are limitations in grasping customer emotions and nuances. As a result, it is difficult to provide rapid feedback for improving service quality and customer satisfaction.
[0175] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0176] In this invention, the server includes means for acquiring an acoustic signal, processing means for converting the acoustic signal into symbolic data, and emotion recognition means for analyzing the data to recognize emotions. This enables the server to instantly grasp the emotional state of the customer and provide feedback to improve the quality of service.
[0177] An "acoustic signal" is a physical representation of sound information collected from the environment or people.
[0178] "Symbolic data" refers to digital data that includes character information converted from acoustic signals.
[0179] "Processing means" refers to a device that has software or hardware functionality for converting acoustic signals into symbolic data.
[0180] "Display means" refers to displays or other display devices that visually present symbolic data.
[0181] An "emotion recognition system" is a system that has the function of analyzing and identifying emotional states from acoustic signals or symbolic data.
[0182] "Adjustment means" refers to a device or function for modifying acoustic signals or display content based on recognized emotions.
[0183] "Communication means" refers to networks and devices used to transfer processed information to external devices.
[0184] An "external device" is another electronic device that can receive and process acoustic signals and data.
[0185] "Visual information" refers to visible data such as characters, colors, and shapes displayed by a system.
[0186] A "user" is a person or organization that uses the system.
[0187] The system for carrying out this invention requires several main components. The server uses a smartphone or other voice input device to acquire acoustic signals. This allows the system to collect voices spoken by the user from the environment. The voice signals are converted into symbolic data using the Google Cloud Speech-to-Text API. This converted data is visualized by a display means. The display means can be a general display or a mobile screen.
[0188] Furthermore, the server uses the Emotion Recognition SDK to recognize emotions from symbolic data. This emotion recognition mechanism analyzes the tone and speed of the voice to identify emotions such as positive, negative, and neutral. Based on this information, the server uses an adjustment mechanism to change the display content and color tone of the symbolic data, providing visual feedback that fits the customer's emotions.
[0189] The communication method transmits processed data to external devices, such as smartphones or tablets, and provides feedback to service providers and in-store staff through these devices. In this way, customer needs can be met in real time, and the quality of service can be improved.
[0190] For example, when a user looks at a product in a store and makes a comment, that comment is captured on their smartphone. The system then performs emotion recognition and notifies the store staff that "the customer is showing interest in the product."
[0191] Examples of prompts for using generative AI models are as follows:
[0192] "The customer is speaking into their smartphone. We analyze their tone and speed of voice to identify their emotions and display a notification in the app."
[0193] This invention provides a useful means to optimize the customer experience in physical stores and to provide services efficiently and effectively.
[0194] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0195] Step 1:
[0196] The terminal acquires the user's speech as an acoustic signal from the microphone. The input is ambient sound, and the output is digitized acoustic data. This data is transmitted to the processing unit.
[0197] Step 2:
[0198] The server converts the received audio signal into symbolic data using the Google Cloud Speech-to-Text API. The input is digital audio data, and the output is symbolic data in text format. A speech recognition algorithm is used for this conversion.
[0199] Step 3:
[0200] The server uses a generative AI model to perform emotion recognition from the content of symbolic data. The input is the symbolic data obtained in step 2, and the output is emotion information (e.g., positive, negative, neutral). The emotion engine achieves this by analyzing the tone and content of the speech.
[0201] Step 4:
[0202] The server dynamically adjusts the display format of symbolic data using emotional information. Input consists of symbolic data and emotional information, while output is display data as visual information for the user. Specifically, this includes actions such as changing text color and font size.
[0203] Step 5:
[0204] The communication system transmits pre-tuned acoustic signals and symbolic data to an external device. The input is pre-tuned data, and the output is information displayed on a receiving device such as a smartphone or tablet. This allows in-store staff to respond to customers in a way that is tailored to their emotions.
[0205] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0206] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0207] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0208] [Second Embodiment]
[0209] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0210] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0211] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0212] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0213] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0214] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0215] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0216] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0217] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0218] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0219] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0220] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0221] This invention is a system for facilitating smooth communication between the elderly and their families. The invention mainly consists of five main components: a voice input device, a processing device, a display device, a regulating device, and a communication device.
[0222] First, when a user speaks into a voice input device, the device acquires voice data. This voice data is converted into a digital format and sent to a processing unit. The processing unit uses a voice recognition module located in the cloud or within the device to convert the voice data into text data. This text data is sent to a display device and presented visually according to user settings such as font size and color.
[0223] Simultaneously, the server optimizes the speed and tone of the audio data to individual needs using an adjustment device. Specifically, it slows down the speed to make the audio easier to understand and adjusts the tone of the voice to maintain a natural sound. The new audio data generated by this adjustment process is transmitted wirelessly to the user's hearing aids via a communication device. This allows the user to hear the adjusted audio in real time.
[0224] For example, when a user says "Let's have lunch together" into a voice input device, the device receives the voice, and the processing unit converts it into the text "Let's have lunch together." This text is then displayed on the display device with an adjusted font size. Meanwhile, the server slows down the voice speed by 30%, and this slowed-down voice is sent to the hearing aid, making it easier for elderly people to hear.
[0225] This invention makes everyday conversations with elderly people who have difficulty with voice communication more comfortable and enables communication that is easier for both parties to understand.
[0226] The following describes the processing flow.
[0227] Step 1:
[0228] The user begins speaking into the device's microphone. The device captures this voice and temporarily stores it as audio data in its internal memory.
[0229] Step 2:
[0230] The device sends the stored audio data to a server in the cloud for processing. The data is transmitted securely using a communication protocol.
[0231] Step 3:
[0232] The server inputs the received audio data into a speech recognition algorithm, converting the audio into text data. This process is performed using a speech recognition API.
[0233] Step 4:
[0234] The converted text data is sent from the server to the terminal. The terminal adjusts the font size and color of the received text according to the user's settings and displays it on the screen.
[0235] Step 5:
[0236] The server adjusts the speed and tone of the original audio data to convert it into easily understandable speech. For this purpose, it uses an audio processing algorithm.
[0237] Step 6:
[0238] The adjusted audio data is sent from the server to the terminal, which then streams it to the user's hearing aids via Bluetooth.
[0239] Step 7:
[0240] Users receive pre-calibrated audio through their hearing aids and understand conversations in real time. Along with displayed text, their listening comprehension is improved.
[0241] (Example 1)
[0242] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0243] In recent years, there has been an increasing number of cases where communication between the elderly and their families is difficult due to hearing and visual limitations. Many current speech recognition technologies and communication systems have limitations in supporting real-time, two-way communication, and a particular lack of systems that provide coordinated audio and visual displays simultaneously has been pointed out. There is a need to solve these problems and provide systems that enable smooth communication for everyone.
[0244] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0245] In this invention, the server includes means for acquiring audio information, means for converting the audio information into text information, and means for visually displaying the text information. This makes it possible to provide auditory support and visual information to make it easier for elderly people to understand audio in real time.
[0246] "Means for acquiring voice information" refers to a device that has the function of capturing voice signals emitted by a user as digital data.
[0247] "Device for converting audio information into text information" refers to a device that performs the process of converting acquired audio information into a string of characters using an algorithm.
[0248] "Means for visually displaying textual information" refers to a device that enables the visual presentation of converted textual information using a display device such as a screen.
[0249] "Devices and means for adjusting the speed and pitch of audio information" refers to devices that have the function of adjusting the playback speed and pitch of audio in order to facilitate the understanding of the audio.
[0250] "Communication means for transmitting adjusted audio information" refers to a device that has a communication function for transmitting adjusted audio data to other devices, particularly hearing assistance devices.
[0251] "Hearing assistance devices" refer to assistive devices that amplify or adjust sound to provide audio information to users with hearing impairments.
[0252] "Display font size and color according to user settings" refers to a function that adjusts the font size and color tone according to the user's preferences.
[0253] This invention is a system for facilitating smooth voice communication between the elderly and their families. The system mainly consists of a device for acquiring voice information, a device for converting voice information into text information, a device for visually displaying the text information, a device for adjusting the speed and tone of the voice information, and a communication device for transmitting the adjusted voice information.
[0254] The user sends a message using a voice input device. The terminal receives this audio and converts it into a digital format. This process uses hardware such as a microphone and an A / D converter. The audio data is sent from the terminal to the processing unit, which uses speech recognition software to convert the audio into text. Specifically, a cloud platform that provides speech recognition services may be used as the software.
[0255] Next, the generated text information is sent to a display device within the terminal and visualized using large fonts and specified colors. This makes it easy for elderly people with visual impairments to recognize the content.
[0256] Simultaneously, the server adjusts the speed and tone of the audio data to individual needs via an adjustment device. The adjusted audio is played back at the optimal speed and tone to make it easier for elderly people to understand.
[0257] The server then transmits the adjusted audio information to the user's hearing aid using wireless communication. This uses communication technologies such as Bluetooth or Wi-Fi.
[0258] As a concrete example, suppose a user says "Let's have lunch together" into a voice input device. The device receives the voice and a processing unit converts it into text information: "Let's have lunch together." This text information is displayed in a large font on the display device, while the server slows down the voice speed by 30% to produce easier-to-understand audio. This adjusted audio is sent to a hearing aid, allowing elderly people to hear the adjusted audio in real time.
[0259] In the generative AI model, the prompt phrase "How to use voice input devices to improve communication with the elderly" is used. This prompt serves as a guide to improve the performance of speech recognition and speech synthesis by the generative AI model.
[0260] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0261] Step 1:
[0262] When a user speaks into a voice input device, voice data is input. The terminal receives this voice data and performs a process to convert the analog voice into a digital format. An analog-to-digital converter is used for this conversion. The resulting digital voice data is then output.
[0263] Step 2:
[0264] The terminal transmits digital audio data to the processing unit. The processing unit performs speech recognition using a generative AI model and generates text data from the audio data. Specifically, it takes audio waveform data as input, analyzes it at the phoneme level, and outputs text data by performing the optimal text conversion.
[0265] Step 3:
[0266] A terminal that receives character data from a processing unit transmits that data to a display device. The display device adjusts the font size and color according to user settings and displays it as visual information. Based on the input character data, character information processed into a format that is easy for the user to read is output.
[0267] Step 4:
[0268] The server receives the digital audio data and adjusts the speed and tone of the audio. This process applies algorithms that stretch the time axis and audio filters to generate easy-to-understand audio data. The adjusted audio data is then output.
[0269] Step 5:
[0270] The server wirelessly transmits the adjusted audio data to the user's hearing aid via a communication device. Wireless technologies such as the Bluetooth protocol are used for communication. Finally, the outputted audio is played back in real time from the hearing aid, enabling elderly people to communicate easily.
[0271] (Application Example 1)
[0272] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0273] This solution addresses the problem of elderly and hearing-impaired individuals experiencing difficulties in shopping and receiving services due to insufficient information, hindering smooth communication. Furthermore, it is necessary to prevent confusion and decreased customer satisfaction caused by inadequate and timely provision of store information.
[0274] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0275] In this invention, the server includes means for acquiring audio data, means for converting the audio data into text data, and means for providing the text data and the adjusted audio data to the guidance device. This makes it possible for elderly customers and customers with hearing impairments to receive guidance information in a more easily understandable format in physical stores.
[0276] "Means for acquiring audio data" refers to a device that collects audio using an input device and processes it as a digital signal.
[0277] "Means of converting to text data" refers to a device that performs the process of analyzing audio data and converting it into textual information.
[0278] "Means of display" refers to a display device for visually presenting the converted text data.
[0279] "Means for adjusting the speed and tone of audio data" refers to a device that performs the process of adjusting the playback speed and tone of audio to suit individual needs.
[0280] "Means for transmitting adjusted audio data" refers to a device for transmitting adjusted audio data wirelessly or via a wired connection to an auxiliary device or receiving terminal.
[0281] "Means of providing information to the guidance device" refers to a communication device for providing necessary information to guidance equipment within the store.
[0282] This system enables smooth communication for the elderly and those with hearing impairments in physical stores by using voice input devices, processing devices, display devices, adjustment devices, and communication devices. The system is centered around a server and is configured as follows:
[0283] The server acquires voice data from customers in the store through a voice input device. The acquired voice data is converted into text data using voice recognition software such as Google Cloud Speech-to-Text. The converted text data is presented as visual information using a display device. The display device uses an e-ink display, an LCD panel, etc., and is displayed in a size and color considering the visibility of the user.
[0284] Furthermore, the adjustment device uses voice processing software such as Adobe Audition or Librosa to adjust the speed and tone of the voice data, providing voice optimized for hearing. This adjusted voice data is transmitted to the auxiliary device using Bluetooth or Wi-Fi. This creates an environment where elderly customers or customers with hearing impairments can clearly understand the information.
[0285] As a specific example, when an elderly person in the store makes a voice input such as "Where is the location of this item?", the terminal captures this voice, and the server immediately processes the data to provide visual display and auditory guidance. The customer can receive both the guidance displayed on the large display and the voice guidance through the auxiliary device.
[0286] As an example of the input prompt sentence for the generated AI model, "A method of displaying in real time the content of questions asked by elderly people in the store by voice and transmitting voice guidance to the auxiliary device" can be cited. This is expected to improve the guidance accuracy and user usability.
[0287] The flow of the specific process in Application Example 1 will be described using FIG. 12.
[0288] Step 1:
[0289] The server acquires voice data from the user through the voice input device. The input voice is converted into digital data using a microphone. The voice input device then prepares to capture the user's voice and send it to the server.
[0290] Step 2:
[0291] The server sends the acquired audio data to the processing unit for speech-to-text conversion. Here, a speech recognition service such as Google Cloud Speech-to-Text is called to convert the audio data into text data. The input is audio data, and the output is text data. At this stage, the audio input is processed to accurately convert it into text.
[0292] Step 3:
[0293] The server sends the converted text data to a display device, where it is displayed as visual information. The display is configured to show the received text using a highly legible font and color. The input is text data, and the output is visual information that the user can confirm. This allows the user to verify the information provided.
[0294] Step 4:
[0295] The server sends audio data to an adjustment device to adjust the speed and timbre of the audio. This adjustment uses software such as Librosa or Adobe Audition, and aims to clarify the sound through data processing. The input is the original audio data, and the output is the adjusted audio data. At this stage, audio optimized for auditory perception is generated.
[0296] Step 5:
[0297] The server transmits the adjusted audio data to the assistive device via a communication device. Wireless communication using Bluetooth or Wi-Fi is then used to deliver the audio to the user's assistive device. The input is the adjusted audio data, and the output is the audio heard by the user. The user can auditorily understand the guidance content by receiving the adjusted audio guidance.
[0298] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0299] This invention relates to a communication system that not only acquires voice data and converts it into text data, but also recognizes the user's emotions. The system includes a voice input device, a processing device, a display device, a regulating device, a communication device, and an emotion engine.
[0300] When a user speaks into a voice input device, the device collects voice data. This voice data is sent to a server for processing and converted into text data by a speech recognition algorithm. Similarly, an emotion engine extracts emotional information from the user's voice. This emotional information influences the display of the converted text. For example, if the user is excited, visual feedback such as a change in text color can be provided.
[0301] The server uses the output of the emotion engine to automatically adjust the speed and tone of the audio data, and then transmits this adjusted audio data to the user's hearing aids via a communication device. This process results in more natural and easily understandable communication for the listener.
[0302] As a specific example, when a user says with a smile "I'm looking forward to tomorrow's presentation", the terminal acquires the voice, and the processing device recognizes the emotion of "happy" along with the text conversion of "I'm looking forward to tomorrow's presentation". Based on this information, the display device displays the text in bright colors, and the server adjusts the tone of the voice slightly higher and sends it to the hearing aid, making it possible to convey the user's emotion more richly. <00009The server sends the converted text data back to the terminal, simultaneously adding information based on the user's emotions. The terminal receives this text data and displays it on the screen, visually adjusting the font size and color according to the emotions.
[0313] Step 5:
[0314] The server adjusts the speed and tone of the voice data based on the output of the emotion engine. Emotion-appropriate voice emphasis is applied to create a more natural conversation.
[0315] Step 6:
[0316] The adjusted audio data is sent from the server to the terminal, which then streams this audio to the user's hearing aids via Bluetooth.
[0317] Step 7:
[0318] Through the hearing aid, users receive both calibrated audio and visually calibrated text data, enabling them to understand the content and emotions of conversations in real time.
[0319] (Example 2)
[0320] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0321] With the increasing diversity of modern audio content, users are seeking richer communication that includes not only the content of the audio data but also the speaker's emotions and tone. However, conventional audio data conversion technologies and display methods have been unable to adequately reflect these emotions and tones, resulting in limited information transmission to users. Against this backdrop, there is a need to develop a new system that not only converts audio data to text but also adjusts the audio while considering emotional information.
[0322] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0323] In this invention, the server includes means for acquiring voice information, processing means for converting the voice information into a data format, and means for extracting emotional information from the voice information. This makes it possible to understand the emotions and tone of the voice spoken by the user and to provide voice and display adjusted based on that information.
[0324] "Audio information" refers to data that is created by converting the waveform of sound emitted by a user into an electrical signal and representing it in digital format.
[0325] "Data format" refers to the state in which audio information has been converted into text or other processable digital information.
[0326] "Processing means" means a device or process that acquires audio information, converts it into a data format, and performs other necessary processing.
[0327] "Emotional information" refers to data that analyzes audio information and indicates the speaker's psychological and emotional state.
[0328] "Display means" refers to a device or method for visually displaying digital information.
[0329] "Adjustment means" refers to a process or device that modifies the speed and timbre of audio information to enable more appropriate transmission.
[0330] "Communication means" means technology or equipment for transmitting coordinated voice information to other devices or equipment.
[0331] A "hearing assistance device" refers to a device that outputs audio data in a format suitable for hearing, helping people with hearing impairments to hear sounds.
[0332] "User" refers to an individual or group that uses the system.
[0333] "Display characteristics" refer to elements such as font size and color that visually represent data formats.
[0334] This invention relates to a system that uses a specific voice processing algorithm to effectively analyze a user's voice information and obtain desired information. This system converts the voice information into text format and further extracts emotional information, enabling display and sound adjustment.
[0335] Users can speak to the system through a voice input device. This voice input device consists of a terminal including a standard microphone. The terminal sends the acquired voice information to a server. The server uses speech recognition software to convert the voice information into a data format. Tools such as the Google Cloud Speech-to-Text API are used for this process. Furthermore, the server uses an emotion recognition system to extract emotion information from the voice information. A general-purpose emotion analysis engine is a possible technology used here.
[0336] Based on the extracted emotional information, the display device on the terminal shows text data to the user, visually representing emotions with appropriate colors and font sizes. The server also processes the speed and tone of the audio information and transmits the adjusted audio information to the hearing assistance device via wireless communication. This makes the audio clearer and easier to understand.
[0337] For example, if a user says, "I'm looking forward to tomorrow's presentation," the server transcribes this phrase into text and further recognizes the positive emotion. The display device shows this text in a brighter color, and the server adjusts the tone of voice slightly higher before transmitting it.
[0338] An example of a prompt message is: "Based on the voice information obtained from the user, please extract emotional information and provide it as text data. Also, please suggest adjustments to the text display color and voice tone." This enables emotionally rich and intuitive communication for the user.
[0339] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0340] Step 1:
[0341] The user speaks through a voice input device. The terminal uses a microphone to acquire the user's voice information. The input is the user's voice waveform, and the output is digital voice data. The terminal then performs a specific action to temporarily store this voice data in preparation for subsequent processing.
[0342] Step 2:
[0343] The terminal sends the acquired audio data to the server. Here, the terminal uses an internet connection to transfer data to the server. The input is audio data in digital format, and the output is the data sent to the server. This process involves specific actions using protocols for secure and high-speed data transfer.
[0344] Step 3:
[0345] The server uses speech recognition software to convert audio data into text format. The input is digital audio data received by the server, and the output is text data. In this process, the server uses the Google Cloud Speech-to-Text API and other tools to analyze the characteristics of the audio waveform and perform the specific operations of converting it into text.
[0346] Step 4:
[0347] The server uses an emotion recognition engine to extract emotional information from audio data. The input is digital audio data received by the server, and the output is an emotion label (e.g., "happy," "sad"). The server performs specific actions to identify the speaker's emotions by analyzing the tone, pitch, and speed of the speech.
[0348] Step 5:
[0349] The terminal processes data to change how text data is displayed based on emotional information. Input consists of text data and emotional information, while output is a visually altered text display. The display device performs specific actions, such as changing the color of the text, according to the emotional information.
[0350] Step 6:
[0351] The server adjusts the speed and tone of the audio data based on emotional information. The input is the original audio data and emotional information, and the output is the adjusted audio data. The server uses audio editing software to perform the specific actions required to adjust the tone of the audio.
[0352] Step 7:
[0353] The server transmits the adjusted audio data to the user's hearing aid using a communication method. The input is the adjusted audio data, and the output is the audio received by the user's hearing aid. The server performs the specific operation of transmitting the audio using a wireless communication protocol.
[0354] (Application Example 2)
[0355] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0356] In face-to-face communication at physical stores, there is a problem in immediately grasping customer emotions and reflecting them in service. With conventional methods, voice data is limited to simple text conversion, and there are limitations in grasping customer emotions and nuances. As a result, it is difficult to provide rapid feedback for improving service quality and customer satisfaction.
[0357] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0358] In this invention, the server includes means for acquiring an acoustic signal, processing means for converting the acoustic signal into symbolic data, and emotion recognition means for analyzing the data to recognize emotions. This enables the server to instantly grasp the emotional state of the customer and provide feedback to improve the quality of service.
[0359] An "acoustic signal" is a physical representation of sound information collected from the environment or people.
[0360] "Symbolic data" refers to digital data that includes character information converted from acoustic signals.
[0361] "Processing means" refers to a device that has software or hardware functionality for converting acoustic signals into symbolic data.
[0362] "Display means" refers to displays and other display devices that visually present symbolic data.
[0363] An "emotion recognition system" is a system that has the function of analyzing and identifying emotional states from acoustic signals or symbolic data.
[0364] "Adjustment means" refers to a device or function for modifying acoustic signals or display content based on recognized emotions.
[0365] "Communication means" refers to networks and devices used to transfer processed information to external devices.
[0366] An "external device" is another electronic device that can receive and process acoustic signals and data.
[0367] "Visual information" refers to visible data such as characters, colors, and shapes displayed by a system.
[0368] A "user" is a person or organization that uses the system.
[0369] The system for carrying out this invention requires several main components. The server uses a smartphone or other voice input device to acquire acoustic signals. This allows the system to collect voices spoken by the user from the environment. The voice signals are converted into symbolic data using the Google Cloud Speech-to-Text API. This converted data is visualized by a display means. The display means can be a general display or a mobile screen.
[0370] Furthermore, the server uses the Emotion Recognition SDK to recognize emotions from symbolic data. This emotion recognition mechanism analyzes the tone and speed of the voice to identify emotions such as positive, negative, and neutral. Based on this information, the server uses an adjustment mechanism to change the display content and color tone of the symbolic data, providing visual feedback that fits the customer's emotions.
[0371] The communication method transmits processed data to external devices, such as smartphones or tablets, and provides feedback to service providers and in-store staff through these devices. In this way, customer needs can be met in real time, and the quality of service can be improved.
[0372] For example, when a user looks at a product in a store and makes a comment, that comment is captured on their smartphone. The system then performs emotion recognition and notifies the store staff that "the customer is showing interest in the product."
[0373] Examples of prompts for using generative AI models are as follows:
[0374] "The customer is speaking into their smartphone. We analyze their tone and speed of voice to identify their emotions and display a notification in the app."
[0375] This invention provides a useful means to optimize the customer experience in physical stores and to provide services efficiently and effectively.
[0376] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0377] Step 1:
[0378] The terminal acquires the user's speech as an acoustic signal from the microphone. The input is ambient sound, and the output is digitized acoustic data. This data is transmitted to the processing unit.
[0379] Step 2:
[0380] The server converts the received audio signal into symbolic data using the Google Cloud Speech-to-Text API. The input is digital audio data, and the output is symbolic data in text format. A speech recognition algorithm is used for this conversion.
[0381] Step 3:
[0382] The server uses a generative AI model to perform emotion recognition from the content of symbolic data. The input is the symbolic data obtained in step 2, and the output is emotion information (e.g., positive, negative, neutral). The emotion engine achieves this by analyzing the tone and content of the speech.
[0383] Step 4:
[0384] The server dynamically adjusts the display format of symbolic data using emotional information. Input consists of symbolic data and emotional information, while output is display data as visual information for the user. Specifically, this includes actions such as changing text color and font size.
[0385] Step 5:
[0386] The communication system transmits pre-tuned acoustic signals and symbolic data to an external device. The input is pre-tuned data, and the output is information displayed on a receiving device such as a smartphone or tablet. This allows in-store staff to respond to customers in a way that is tailored to their emotions.
[0387] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0388] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0389] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0390] [Third Embodiment]
[0391] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0392] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0393] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0394] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0395] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0396] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0397] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0398] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0399] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0400] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0401] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0402] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0403] This invention is a system for facilitating smooth communication between the elderly and their families. The invention mainly consists of five main components: a voice input device, a processing device, a display device, a regulating device, and a communication device.
[0404] First, when a user speaks into a voice input device, the device acquires voice data. This voice data is converted into a digital format and sent to a processing unit. The processing unit uses a voice recognition module located in the cloud or within the device to convert the voice data into text data. This text data is sent to a display device and presented visually according to user settings such as font size and color.
[0405] Simultaneously, the server optimizes the speed and tone of the audio data to individual needs using an adjustment device. Specifically, it slows down the speed to make the audio easier to understand and adjusts the tone of the voice to maintain a natural sound. The new audio data generated by this adjustment process is transmitted wirelessly to the user's hearing aids via a communication device. This allows the user to hear the adjusted audio in real time.
[0406] For example, when a user says "Let's have lunch together" into a voice input device, the device receives the voice, and the processing unit converts it into the text "Let's have lunch together." This text is then displayed on the display device with an adjusted font size. Meanwhile, the server slows down the voice speed by 30%, and this slowed-down voice is sent to the hearing aid, making it easier for elderly people to hear.
[0407] This invention makes everyday conversations with elderly people who have difficulty with voice communication more comfortable and enables communication that is easier for both parties to understand.
[0408] The following describes the processing flow.
[0409] Step 1:
[0410] The user begins speaking into the device's microphone. The device captures this voice and temporarily stores it as audio data in its internal memory.
[0411] Step 2:
[0412] The device sends the stored audio data to a server in the cloud for processing. The data is transmitted securely using a communication protocol.
[0413] Step 3:
[0414] The server inputs the received audio data into a speech recognition algorithm, converting the audio into text data. This process is performed using a speech recognition API.
[0415] Step 4:
[0416] The converted text data is sent from the server to the terminal. The terminal adjusts the font size and color of the received text according to the user's settings and displays it on the screen.
[0417] Step 5:
[0418] The server adjusts the speed and tone of the original audio data to convert it into easily understandable speech. For this purpose, it uses an audio processing algorithm.
[0419] Step 6:
[0420] The adjusted audio data is sent from the server to the terminal, which then streams it to the user's hearing aids via Bluetooth.
[0421] Step 7:
[0422] Users receive pre-calibrated audio through their hearing aids and understand conversations in real time. Along with displayed text, their listening comprehension is improved.
[0423] (Example 1)
[0424] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0425] In recent years, there has been an increasing number of cases where communication between the elderly and their families is difficult due to hearing and visual limitations. Many current speech recognition technologies and communication systems have limitations in supporting real-time, two-way communication, and a particular lack of systems that provide coordinated audio and visual displays simultaneously has been pointed out. There is a need to solve these problems and provide systems that enable smooth communication for everyone.
[0426] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0427] In this invention, the server includes means for acquiring audio information, means for converting the audio information into text information, and means for visually displaying the text information. This makes it possible to provide auditory support and visual information to make it easier for elderly people to understand audio in real time.
[0428] "Means for acquiring voice information" refers to a device that has the function of capturing voice signals emitted by a user as digital data.
[0429] "Device for converting audio information into text information" refers to a device that performs the process of converting acquired audio information into a string of characters using an algorithm.
[0430] "Means for visually displaying textual information" refers to a device that enables the visual presentation of converted textual information using a display device such as a screen.
[0431] "Devices and means for adjusting the speed and pitch of audio information" refers to devices that have the function of adjusting the playback speed and pitch of audio in order to facilitate the understanding of the audio.
[0432] "Communication means for transmitting adjusted audio information" refers to a device that has a communication function for transmitting adjusted audio data to other devices, particularly hearing assistance devices.
[0433] "Hearing assistance devices" refer to assistive devices that amplify or adjust sound to provide audio information to users with hearing impairments.
[0434] "Display font size and color according to user settings" refers to a function that adjusts the font size and color tone according to the user's preferences.
[0435] This invention is a system for facilitating smooth voice communication between the elderly and their families. The system mainly consists of a device for acquiring voice information, a device for converting voice information into text information, a device for visually displaying the text information, a device for adjusting the speed and tone of the voice information, and a communication device for transmitting the adjusted voice information.
[0436] The user sends a message using a voice input device. The terminal receives this audio and converts it into a digital format. This process uses hardware such as a microphone and an A / D converter. The audio data is sent from the terminal to the processing unit, which uses speech recognition software to convert the audio into text. Specifically, a cloud platform that provides speech recognition services may be used as the software.
[0437] Next, the generated text information is sent to a display device within the terminal and visualized using large fonts and specified colors. This makes it easy for elderly people with visual impairments to recognize the content.
[0438] Simultaneously, the server adjusts the speed and tone of the audio data via a tuning device to suit individual needs. The adjusted audio is played back at the optimal speed and tone to make it easier for elderly people to understand.
[0439] The server then transmits the adjusted audio information to the user's hearing aid using wireless communication. This uses communication technologies such as Bluetooth or Wi-Fi.
[0440] As a concrete example, suppose a user says "Let's have lunch together" into a voice input device. The device receives the voice and a processing unit converts it into text information: "Let's have lunch together." This text information is displayed in a large font on the display device, while the server slows down the voice speed by 30% to produce easier-to-understand audio. This adjusted audio is sent to a hearing aid, allowing elderly people to hear the adjusted audio in real time.
[0441] In the generative AI model, the prompt phrase "How to use voice input devices to improve communication with the elderly" is used. This prompt serves as a guide to improve the performance of speech recognition and speech synthesis by the generative AI model.
[0442] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0443] Step 1:
[0444] When a user speaks into a voice input device, voice data is input. The terminal receives this voice data and performs a process to convert the analog voice into a digital format. An analog-to-digital converter is used for this conversion. The resulting digital voice data is then output.
[0445] Step 2:
[0446] The terminal transmits digital audio data to the processing unit. The processing unit performs speech recognition using a generative AI model and generates text data from the audio data. Specifically, it takes audio waveform data as input, analyzes it at the phoneme level, and outputs text data by performing the optimal text conversion.
[0447] Step 3:
[0448] A terminal that receives character data from a processing unit transmits that data to a display device. The display device adjusts the font size and color according to user settings and displays it as visual information. Based on the input character data, character information processed into a format that is easy for the user to read is output.
[0449] Step 4:
[0450] The server receives the digital audio data and adjusts the speed and tone of the audio. This process applies algorithms that stretch the time axis and audio filters to generate easy-to-understand audio data. The adjusted audio data is then output.
[0451] Step 5:
[0452] The server wirelessly transmits the adjusted audio data to the user's hearing aid via a communication device. Wireless technologies such as the Bluetooth protocol are used for communication. Finally, the outputted audio is played back in real time from the hearing aid, enabling elderly people to communicate easily.
[0453] (Application Example 1)
[0454] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0455] This solution addresses the problem of elderly and hearing-impaired individuals experiencing difficulties in shopping and receiving services due to insufficient information, hindering smooth communication. Furthermore, it is necessary to prevent confusion and decreased customer satisfaction caused by inadequate and timely provision of store information.
[0456] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0457] In this invention, the server includes means for acquiring audio data, means for converting the audio data into text data, and means for providing the text data and the adjusted audio data to the guidance device. This makes it possible for elderly customers and customers with hearing impairments to receive guidance information in a more easily understandable format in physical stores.
[0458] "Means for acquiring audio data" refers to a device that collects audio using an input device and processes it as a digital signal.
[0459] "Means of converting to text data" refers to a device that performs the process of analyzing audio data and converting it into textual information.
[0460] "Means of display" refers to a display device for visually presenting the converted text data.
[0461] "Means for adjusting the speed and tone of audio data" refers to a device that performs the process of adjusting the playback speed and tone of audio to suit individual needs.
[0462] "Means for transmitting adjusted audio data" refers to a device for transmitting adjusted audio data wirelessly or via a wired connection to an auxiliary device or receiving terminal.
[0463] "Means of providing information to the guidance device" refers to a communication device for providing necessary information to guidance equipment within the store.
[0464] This system enables smooth communication for the elderly and those with hearing impairments in physical stores by using voice input devices, processing devices, display devices, adjustment devices, and communication devices. The system is centered around a server and is configured as follows:
[0465] The server acquires voice data from customers in the store via voice input devices. The acquired voice data is converted into text data using speech recognition software such as Google Cloud Speech-to-Text. The converted text data is presented as visual information using a display device. The display device uses e-ink displays or LCD panels, and is displayed in a size and color that takes user visibility into consideration.
[0466] Furthermore, the adjustment device uses audio processing software such as Adobe Audition and Librosa to adjust the speed and tone of the audio data, providing audio optimized for hearing. This adjusted audio data is transmitted to the assistive device using Bluetooth or Wi-Fi. This creates an environment where elderly or hearing-impaired customers can clearly understand the information.
[0467] As a concrete example, if an elderly person in a store uses voice input to ask, "Where is this item located?", the terminal captures this voice, and the server immediately processes the data to provide both visual and auditory guidance. Customers can receive guidance both displayed on a large screen and audio guidance through assistive devices.
[0468] An example of an input prompt for a generated AI model is "a method to display in real time the content of a question asked by an elderly person in a store and send an audio guide to an assistive device." This is expected to improve the accuracy of guidance and usability.
[0469] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0470] Step 1:
[0471] The server acquires voice data from the user through the voice input device. The input voice is converted into digital data using a microphone. The voice input device then prepares to capture the user's voice and send it to the server.
[0472] Step 2:
[0473] The server sends the acquired audio data to the processing unit for speech-to-text conversion. Here, a speech recognition service such as Google Cloud Speech-to-Text is called to convert the audio data into text data. The input is audio data, and the output is text data. At this stage, the audio input is processed to accurately convert it into text.
[0474] Step 3:
[0475] The server sends the converted text data to a display device, where it is displayed as visual information. The display is configured to show the received text using a highly legible font and color. The input is text data, and the output is visual information that the user can confirm. This allows the user to verify the information provided.
[0476] Step 4:
[0477] The server sends audio data to an adjustment device to adjust the speed and timbre of the audio. This adjustment uses software such as Librosa or Adobe Audition, and aims to clarify the sound through data processing. The input is the original audio data, and the output is the adjusted audio data. At this stage, audio optimized for auditory perception is generated.
[0478] Step 5:
[0479] The server transmits the adjusted audio data to the assistive device via a communication device. Wireless communication using Bluetooth or Wi-Fi is then used to deliver the audio to the user's assistive device. The input is the adjusted audio data, and the output is the audio heard by the user. The user can auditorily understand the guidance content by receiving the adjusted audio guidance.
[0480] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0481] This invention relates to a communication system that not only acquires voice data and converts it into text data, but also recognizes the user's emotions. The system includes a voice input device, a processing device, a display device, a regulating device, a communication device, and an emotion engine.
[0482] When a user speaks into a voice input device, the device collects voice data. This voice data is sent to a server for processing and converted into text data by a speech recognition algorithm. Similarly, an emotion engine extracts emotional information from the user's voice. This emotional information influences the display of the converted text. For example, if the user is excited, visual feedback such as a change in text color can be provided.
[0483] The server uses the output of the emotion engine to automatically adjust the speed and tone of the audio data, and then transmits this adjusted audio data to the user's hearing aids via a communication device. This process results in more natural and easily understandable communication for the listener.
[0484] For example, if a user smiles and says, "I'm looking forward to tomorrow's presentation," the device captures the audio, and the processing unit recognizes the emotion of "happiness" along with converting it to text. Based on this information, the display device displays the text in a brighter color, and the server adjusts the tone of the audio slightly higher before sending it to the hearing aid, thereby conveying the user's emotions more richly.
[0485] This invention provides a more personalized experience by comprehensively capturing the user's emotions and adjusting the entire communication accordingly. This enables smoother dialogue between the elderly and their families, overcoming the limitations of conventional voice communication.
[0486] The following describes the processing flow.
[0487] Step 1:
[0488] The user speaks into the device's microphone. The device captures this voice and temporarily stores it as digital audio data.
[0489] Step 2:
[0490] The terminal sends the acquired audio data to the server. On the server side, the audio data is processed and converted into text data using speech recognition technology.
[0491] Step 3:
[0492] The server inputs audio data into the emotion engine before and after speech recognition, analyzing the user's emotions from their voice. This emotion data is then used for subsequent processing.
[0493] Step 4:
[0494] The server sends the converted text data back to the terminal, simultaneously adding information based on the user's emotions. The terminal receives this text data and displays it on the screen, visually adjusting the font size and color according to the emotions.
[0495] Step 5:
[0496] The server adjusts the speed and tone of the voice data based on the output of the emotion engine. Emotion-appropriate voice emphasis is applied to create a more natural conversation.
[0497] Step 6:
[0498] The adjusted audio data is sent from the server to the terminal, which then streams this audio to the user's hearing aids via Bluetooth.
[0499] Step 7:
[0500] Through the hearing aid, users receive both calibrated audio and visually calibrated text data, enabling them to understand the content and emotions of conversations in real time.
[0501] (Example 2)
[0502] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0503] With the increasing diversity of modern audio content, users are seeking richer communication that includes not only the content of the audio data but also the speaker's emotions and tone. However, conventional audio data conversion technologies and display methods have been unable to adequately reflect these emotions and tones, resulting in limited information transmission to users. Against this backdrop, there is a need to develop a new system that not only converts audio data to text but also adjusts the audio while considering emotional information.
[0504] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0505] In this invention, the server includes means for acquiring voice information, processing means for converting the voice information into a data format, and means for extracting emotional information from the voice information. This makes it possible to understand the emotions and tone of the voice spoken by the user and to provide voice and display adjusted based on that information.
[0506] "Audio information" refers to data that is created by converting the waveform of sound emitted by a user into an electrical signal and representing it in digital format.
[0507] "Data format" refers to the state in which audio information has been converted into text or other processable digital information.
[0508] "Processing means" means a device or process that acquires audio information, converts it into a data format, and performs other necessary processing.
[0509] "Emotional information" refers to data that analyzes audio information and indicates the speaker's psychological and emotional state.
[0510] "Display means" refers to a device or method for visually displaying digital information.
[0511] "Adjustment means" refers to a process or device that modifies the speed and timbre of audio information to enable more appropriate transmission.
[0512] "Communication means" means technology or equipment for transmitting coordinated voice information to other devices or equipment.
[0513] A "hearing assistance device" refers to a device that outputs audio data in a format suitable for hearing, helping people with hearing impairments to hear sounds.
[0514] "User" refers to an individual or group that uses the system.
[0515] "Display characteristics" refer to elements such as font size and color that visually represent data formats.
[0516] This invention relates to a system that uses a specific voice processing algorithm to effectively analyze a user's voice information and obtain desired information. This system converts the voice information into text format and further extracts emotional information, enabling display and sound adjustment.
[0517] Users can speak to the system through a voice input device. This voice input device consists of a terminal including a standard microphone. The terminal sends the acquired voice information to a server. The server uses speech recognition software to convert the voice information into a data format. Tools such as the Google Cloud Speech-to-Text API are used for this process. Furthermore, the server uses an emotion recognition system to extract emotion information from the voice information. A general-purpose emotion analysis engine is a possible technology used here.
[0518] Based on the extracted emotional information, the display device on the terminal shows text data to the user, visually representing emotions with appropriate colors and font sizes. The server also processes the speed and tone of the audio information and transmits the adjusted audio information to the hearing assistance device via wireless communication. This makes the audio clearer and easier to understand.
[0519] For example, if a user says, "I'm looking forward to tomorrow's presentation," the server transcribes this phrase into text and further recognizes the positive emotion. The display device shows this text in a brighter color, and the server adjusts the tone of voice slightly higher before transmitting it.
[0520] An example of a prompt message is: "Based on the voice information obtained from the user, please extract emotional information and provide it as text data. Also, please suggest adjustments to the text display color and voice tone." This enables emotionally rich and intuitive communication for the user.
[0521] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0522] Step 1:
[0523] The user speaks through a voice input device. The terminal uses a microphone to acquire the user's voice information. The input is the user's voice waveform, and the output is digital voice data. The terminal then performs a specific action to temporarily store this voice data in preparation for subsequent processing.
[0524] Step 2:
[0525] The terminal sends the acquired audio data to the server. Here, the terminal uses an internet connection to transfer data to the server. The input is audio data in digital format, and the output is the data sent to the server. This process involves specific actions using protocols for secure and high-speed data transfer.
[0526] Step 3:
[0527] The server uses speech recognition software to convert audio data into text format. The input is digital audio data received by the server, and the output is text data. In this process, the server uses the Google Cloud Speech-to-Text API and other tools to analyze the characteristics of the audio waveform and perform the specific operations of converting it into text.
[0528] Step 4:
[0529] The server uses an emotion recognition engine to extract emotional information from audio data. The input is digital audio data received by the server, and the output is an emotion label (e.g., "happy," "sad"). The server performs specific actions to identify the speaker's emotions by analyzing the tone, pitch, and speed of the speech.
[0530] Step 5:
[0531] The terminal processes data to change how text data is displayed based on emotional information. Input consists of text data and emotional information, while output is a visually altered text display. The display device performs specific actions, such as changing the color of the text, according to the emotional information.
[0532] Step 6:
[0533] The server adjusts the speed and tone of the audio data based on emotional information. The input is the original audio data and emotional information, and the output is the adjusted audio data. The server uses audio editing software to perform the specific actions required to adjust the tone of the audio.
[0534] Step 7:
[0535] The server transmits the adjusted audio data to the user's hearing aid using a communication method. The input is the adjusted audio data, and the output is the audio received by the user's hearing aid. The server performs the specific operation of transmitting the audio using a wireless communication protocol.
[0536] (Application Example 2)
[0537] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0538] In face-to-face communication at physical stores, there is a problem in immediately grasping customer emotions and reflecting them in service. With conventional methods, voice data is limited to simple text conversion, and there are limitations in grasping customer emotions and nuances. As a result, it is difficult to provide rapid feedback for improving service quality and customer satisfaction.
[0539] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0540] In this invention, the server includes means for acquiring an acoustic signal, processing means for converting the acoustic signal into symbolic data, and emotion recognition means for analyzing the data to recognize emotions. This enables the server to instantly grasp the emotional state of the customer and provide feedback to improve the quality of service.
[0541] An "acoustic signal" is a physical representation of sound information collected from the environment or people.
[0542] "Symbolic data" refers to digital data that includes character information converted from acoustic signals.
[0543] "Processing means" refers to a device that has software or hardware functionality for converting acoustic signals into symbolic data.
[0544] "Display means" refers to displays and other display devices that visually present symbolic data.
[0545] An "emotion recognition system" is a system that has the function of analyzing and identifying emotional states from acoustic signals or symbolic data.
[0546] "Adjustment means" refers to a device or function for modifying acoustic signals or display content based on recognized emotions.
[0547] "Communication means" refers to networks and devices used to transfer processed information to external devices.
[0548] An "external device" is another electronic device that can receive and process acoustic signals and data.
[0549] "Visual information" refers to visible data such as characters, colors, and shapes displayed by a system.
[0550] A "user" is a person or organization that uses the system.
[0551] The system for carrying out this invention requires several main components. The server uses a smartphone or other voice input device to acquire acoustic signals. This allows the system to collect voices spoken by the user from the environment. The voice signals are converted into symbolic data using the Google Cloud Speech-to-Text API. This converted data is visualized by a display means. The display means can be a general display or a mobile screen.
[0552] Furthermore, the server uses the Emotion Recognition SDK to recognize emotions from symbolic data. This emotion recognition mechanism analyzes the tone and speed of the voice to identify emotions such as positive, negative, and neutral. Based on this information, the server uses an adjustment mechanism to change the display content and color tone of the symbolic data, providing visual feedback that fits the customer's emotions.
[0553] The communication method transmits processed data to external devices, such as smartphones or tablets, and provides feedback to service providers and in-store staff through these devices. In this way, customer needs can be met in real time, and the quality of service can be improved.
[0554] For example, when a user looks at a product in a store and makes a comment, that comment is captured on their smartphone. The system then performs emotion recognition and notifies the store staff that "the customer is showing interest in the product."
[0555] Examples of prompts for using generative AI models are as follows:
[0556] "The customer is speaking into their smartphone. We analyze their tone and speed of voice to identify their emotions and display a notification in the app."
[0557] This invention provides a useful means to optimize the customer experience in physical stores and to provide services efficiently and effectively.
[0558] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0559] Step 1:
[0560] The terminal acquires the user's speech as an acoustic signal from the microphone. The input is ambient sound, and the output is digitized acoustic data. This data is transmitted to the processing unit.
[0561] Step 2:
[0562] The server converts the received audio signal into symbolic data using the Google Cloud Speech-to-Text API. The input is digital audio data, and the output is symbolic data in text format. A speech recognition algorithm is used for this conversion.
[0563] Step 3:
[0564] The server uses a generative AI model to perform emotion recognition from the content of symbolic data. The input is the symbolic data obtained in step 2, and the output is emotion information (e.g., positive, negative, neutral). The emotion engine achieves this by analyzing the tone and content of the speech.
[0565] Step 4:
[0566] The server dynamically adjusts the display format of symbolic data using emotional information. Input consists of symbolic data and emotional information, while output is display data as visual information for the user. Specifically, this includes actions such as changing text color and font size.
[0567] Step 5:
[0568] The communication system transmits pre-tuned acoustic signals and symbolic data to an external device. The input is pre-tuned data, and the output is information displayed on a receiving device such as a smartphone or tablet. This allows in-store staff to respond to customers in a way that is tailored to their emotions.
[0569] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0570] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0571] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0572] [Fourth Embodiment]
[0573] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0574] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0575] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0576] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0577] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0578] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0579] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0580] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0581] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0582] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0583] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0584] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0585] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0586] This invention is a system for facilitating smooth communication between the elderly and their families. The invention mainly consists of five main components: a voice input device, a processing device, a display device, a regulating device, and a communication device.
[0587] First, when a user speaks into a voice input device, the device acquires voice data. This voice data is converted into a digital format and sent to a processing unit. The processing unit uses a voice recognition module located in the cloud or within the device to convert the voice data into text data. This text data is sent to a display device and presented visually according to user settings such as font size and color.
[0588] Simultaneously, the server optimizes the speed and tone of the audio data to individual needs using an adjustment device. Specifically, it slows down the speed to make the audio easier to understand and adjusts the tone of the voice to maintain a natural sound. The new audio data generated by this adjustment process is transmitted wirelessly to the user's hearing aids via a communication device. This allows the user to hear the adjusted audio in real time.
[0589] For example, when a user says "Let's have lunch together" into a voice input device, the device receives the voice, and the processing unit converts it into the text "Let's have lunch together." This text is then displayed on the display device with an adjusted font size. Meanwhile, the server slows down the voice speed by 30%, and this slowed-down voice is sent to the hearing aid, making it easier for elderly people to hear.
[0590] This invention makes everyday conversations with elderly people who have difficulty with voice communication more comfortable and enables communication that is easier for both parties to understand.
[0591] The following describes the processing flow.
[0592] Step 1:
[0593] The user begins speaking into the device's microphone. The device captures this voice and temporarily stores it as audio data in its internal memory.
[0594] Step 2:
[0595] The device sends the stored audio data to a server in the cloud for processing. The data is transmitted securely using a communication protocol.
[0596] Step 3:
[0597] The server inputs the received audio data into a speech recognition algorithm, converting the audio into text data. This process is performed using a speech recognition API.
[0598] Step 4:
[0599] The converted text data is sent from the server to the terminal. The terminal adjusts the font size and color of the received text according to the user's settings and displays it on the screen.
[0600] Step 5:
[0601] The server adjusts the speed and tone of the original audio data to convert it into easily understandable speech. For this purpose, it uses an audio processing algorithm.
[0602] Step 6:
[0603] The adjusted audio data is sent from the server to the terminal, which then streams it to the user's hearing aids via Bluetooth.
[0604] Step 7:
[0605] Users receive pre-calibrated audio through their hearing aids and understand conversations in real time. Along with displayed text, their listening comprehension is improved.
[0606] (Example 1)
[0607] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0608] In recent years, there has been an increasing number of cases where communication between the elderly and their families is difficult due to hearing and visual limitations. Many current speech recognition technologies and communication systems have limitations in supporting real-time, two-way communication, and a particular lack of systems that provide coordinated audio and visual displays simultaneously has been pointed out. There is a need to solve these problems and provide systems that enable smooth communication for everyone.
[0609] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0610] In this invention, the server includes means for acquiring audio information, means for converting the audio information into text information, and means for visually displaying the text information. This makes it possible to provide auditory support and visual information to make it easier for elderly people to understand audio in real time.
[0611] "Means for acquiring voice information" refers to a device that has the function of capturing voice signals emitted by a user as digital data.
[0612] "Device for converting audio information into text information" refers to a device that performs the process of converting acquired audio information into a string of characters using an algorithm.
[0613] "Means for visually displaying textual information" refers to a device that enables the visual presentation of converted textual information using a display device such as a screen.
[0614] "Devices and means for adjusting the speed and pitch of audio information" refers to devices that have the function of adjusting the playback speed and pitch of audio in order to facilitate the understanding of the audio.
[0615] "Communication means for transmitting adjusted audio information" refers to a device that has a communication function for transmitting adjusted audio data to other devices, particularly hearing assistance devices.
[0616] "Hearing assistance devices" refer to assistive devices that amplify or adjust sound to provide audio information to users with hearing impairments.
[0617] "Display font size and color according to user settings" refers to a function that adjusts the font size and color tone according to the user's preferences.
[0618] This invention is a system for facilitating smooth voice communication between the elderly and their families. The system mainly consists of a device for acquiring voice information, a device for converting voice information into text information, a device for visually displaying the text information, a device for adjusting the speed and tone of the voice information, and a communication device for transmitting the adjusted voice information.
[0619] The user sends a message using a voice input device. The terminal receives this audio and converts it into a digital format. This process uses hardware such as a microphone and an A / D converter. The audio data is sent from the terminal to the processing unit, which uses speech recognition software to convert the audio into text. Specifically, a cloud platform that provides speech recognition services may be used as the software.
[0620] Next, the generated text information is sent to a display device within the terminal and visualized using large fonts and specified colors. This makes it easy for elderly people with visual impairments to recognize the content.
[0621] Simultaneously, the server adjusts the speed and tone of the audio data to individual needs via an adjustment device. The adjusted audio is played back at the optimal speed and tone to make it easier for elderly people to understand.
[0622] The server then transmits the adjusted audio information to the user's hearing aid using wireless communication. This uses communication technologies such as Bluetooth or Wi-Fi.
[0623] As a concrete example, suppose a user says "Let's have lunch together" into a voice input device. The device receives the voice and a processing unit converts it into text information: "Let's have lunch together." This text information is displayed in a large font on the display device, while the server slows down the voice speed by 30% to produce easier-to-understand audio. This adjusted audio is sent to a hearing aid, allowing elderly people to hear the adjusted audio in real time.
[0624] In the generative AI model, the prompt phrase "How to use voice input devices to improve communication with the elderly" is used. This prompt serves as a guide to improve the performance of speech recognition and speech synthesis by the generative AI model.
[0625] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0626] Step 1:
[0627] When a user speaks into a voice input device, voice data is input. The terminal receives this voice data and performs a process to convert the analog voice into a digital format. An analog-to-digital converter is used for this conversion. The resulting digital voice data is then output.
[0628] Step 2:
[0629] The terminal transmits digital audio data to the processing unit. The processing unit performs speech recognition using a generative AI model and generates text data from the audio data. Specifically, it takes audio waveform data as input, analyzes it at the phoneme level, and outputs text data by performing the optimal text conversion.
[0630] Step 3:
[0631] A terminal that receives character data from a processing unit transmits that data to a display device. The display device adjusts the font size and color according to user settings and displays it as visual information. Based on the input character data, character information processed into a format that is easy for the user to read is output.
[0632] Step 4:
[0633] The server receives the digital audio data and adjusts the speed and tone of the audio. This process applies algorithms that stretch the time axis and audio filters to generate easy-to-understand audio data. The adjusted audio data is then output.
[0634] Step 5:
[0635] The server wirelessly transmits the adjusted audio data to the user's hearing aid via a communication device. Wireless technologies such as the Bluetooth protocol are used for communication. Finally, the outputted audio is played back in real time from the hearing aid, enabling elderly people to communicate easily.
[0636] (Application Example 1)
[0637] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0638] This solution addresses the problem of elderly and hearing-impaired individuals experiencing difficulties in shopping and receiving services due to insufficient information, hindering smooth communication. Furthermore, it is necessary to prevent confusion and decreased customer satisfaction caused by inadequate and timely provision of store information.
[0639] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0640] In this invention, the server includes means for acquiring audio data, means for converting the audio data into text data, and means for providing the text data and the adjusted audio data to the guidance device. This makes it possible for elderly customers and customers with hearing impairments to receive guidance information in a more easily understandable format in physical stores.
[0641] "Means for acquiring audio data" refers to a device that collects audio using an input device and processes it as a digital signal.
[0642] "Means of converting to text data" refers to a device that performs the process of analyzing audio data and converting it into textual information.
[0643] "Means of display" refers to a display device for visually presenting the converted text data.
[0644] "Means for adjusting the speed and tone of audio data" refers to a device that performs the process of adjusting the playback speed and tone of audio to suit individual needs.
[0645] "Means for transmitting adjusted audio data" refers to a device for transmitting adjusted audio data wirelessly or via a wired connection to an auxiliary device or receiving terminal.
[0646] "Means of providing information to the guidance device" refers to a communication device for providing necessary information to guidance equipment within the store.
[0647] This system enables smooth communication for the elderly and those with hearing impairments in physical stores by using voice input devices, processing devices, display devices, adjustment devices, and communication devices. The system is centered around a server and is configured as follows:
[0648] The server acquires voice data from customers in the store via voice input devices. The acquired voice data is converted into text data using speech recognition software such as Google Cloud Speech-to-Text. The converted text data is presented as visual information using a display device. The display device uses e-ink displays or LCD panels, and is displayed in a size and color that takes user visibility into consideration.
[0649] Furthermore, the adjustment device uses audio processing software such as Adobe Audition and Librosa to adjust the speed and tone of the audio data, providing audio optimized for hearing. This adjusted audio data is transmitted to the assistive device using Bluetooth or Wi-Fi. This creates an environment where elderly or hearing-impaired customers can clearly understand the information.
[0650] As a concrete example, if an elderly person in a store uses voice input to ask, "Where is this item located?", the terminal captures this voice, and the server immediately processes the data to provide both visual and auditory guidance. Customers can receive guidance both displayed on a large screen and audio guidance through assistive devices.
[0651] An example of an input prompt for a generated AI model is "a method to display in real time the content of a question asked by an elderly person in a store and send an audio guide to an assistive device." This is expected to improve the accuracy of guidance and usability.
[0652] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0653] Step 1:
[0654] The server acquires voice data from the user through the voice input device. The input voice is converted into digital data using a microphone. The voice input device then prepares to capture the user's voice and send it to the server.
[0655] Step 2:
[0656] The server sends the acquired audio data to the processing unit for speech-to-text conversion. Here, a speech recognition service such as Google Cloud Speech-to-Text is called to convert the audio data into text data. The input is audio data, and the output is text data. At this stage, the audio input is processed to accurately convert it into text.
[0657] Step 3:
[0658] The server sends the converted text data to a display device, where it is displayed as visual information. The display is configured to show the received text using a highly legible font and color. The input is text data, and the output is visual information that the user can confirm. This allows the user to verify the information provided.
[0659] Step 4:
[0660] The server sends audio data to an adjustment device to adjust the speed and timbre of the audio. This adjustment uses software such as Librosa or Adobe Audition, and aims to clarify the sound through data processing. The input is the original audio data, and the output is the adjusted audio data. At this stage, audio optimized for auditory perception is generated.
[0661] Step 5:
[0662] The server transmits the adjusted audio data to the assistive device via a communication device. Wireless communication using Bluetooth or Wi-Fi is then used to deliver the audio to the user's assistive device. The input is the adjusted audio data, and the output is the audio heard by the user. The user can auditorily understand the guidance content by receiving the adjusted audio guidance.
[0663] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0664] This invention relates to a communication system that not only acquires voice data and converts it into text data, but also recognizes the user's emotions. The system includes a voice input device, a processing device, a display device, a regulating device, a communication device, and an emotion engine.
[0665] When a user speaks into a voice input device, the device collects voice data. This voice data is sent to a server for processing and converted into text data by a speech recognition algorithm. Similarly, an emotion engine extracts emotional information from the user's voice. This emotional information influences the display of the converted text. For example, if the user is excited, visual feedback such as a change in text color can be provided.
[0666] The server uses the output of the emotion engine to automatically adjust the speed and tone of the audio data, and then transmits this adjusted audio data to the user's hearing aids via a communication device. This process results in more natural and easily understandable communication for the listener.
[0667] For example, if a user smiles and says, "I'm looking forward to tomorrow's presentation," the device captures the audio, and the processing unit recognizes the emotion of "happiness" along with converting it to text. Based on this information, the display device displays the text in a brighter color, and the server adjusts the tone of the audio slightly higher before sending it to the hearing aid, thereby conveying the user's emotions more richly.
[0668] This invention provides a more personalized experience by comprehensively capturing the user's emotions and adjusting the entire communication accordingly. This enables smoother dialogue between the elderly and their families, overcoming the limitations of conventional voice communication.
[0669] The following describes the processing flow.
[0670] Step 1:
[0671] The user speaks into the device's microphone. The device captures this voice and temporarily stores it as digital audio data.
[0672] Step 2:
[0673] The terminal sends the acquired audio data to the server. On the server side, the audio data is processed and converted into text data using speech recognition technology.
[0674] Step 3:
[0675] The server inputs audio data into the emotion engine before and after speech recognition, analyzing the user's emotions from their voice. This emotion data is then used for subsequent processing.
[0676] Step 4:
[0677] The server sends the converted text data back to the terminal, simultaneously adding information based on the user's emotions. The terminal receives this text data and displays it on the screen, visually adjusting the font size and color according to the emotions.
[0678] Step 5:
[0679] The server adjusts the speed and tone of the voice data based on the output of the emotion engine. Emotion-appropriate voice emphasis is applied to create a more natural conversation.
[0680] Step 6:
[0681] The adjusted audio data is sent from the server to the terminal, which then streams this audio to the user's hearing aids via Bluetooth.
[0682] Step 7:
[0683] Through the hearing aid, users receive both calibrated audio and visually calibrated text data, enabling them to understand the content and emotions of conversations in real time.
[0684] (Example 2)
[0685] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0686] With the increasing diversity of modern audio content, users are seeking richer communication that includes not only the content of the audio data but also the speaker's emotions and tone. However, conventional audio data conversion technologies and display methods have been unable to adequately reflect these emotions and tones, resulting in limited information transmission to users. Against this backdrop, there is a need to develop a new system that not only converts audio data to text but also adjusts the audio while considering emotional information.
[0687] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0688] In this invention, the server includes means for acquiring voice information, processing means for converting the voice information into a data format, and means for extracting emotional information from the voice information. This makes it possible to understand the emotions and tone of the voice spoken by the user and to provide voice and display adjusted based on that information.
[0689] "Audio information" refers to data that is created by converting the waveform of sound emitted by a user into an electrical signal and representing it in digital format.
[0690] "Data format" refers to the state in which audio information has been converted into text or other processable digital information.
[0691] "Processing means" means a device or process that acquires audio information, converts it into a data format, and performs other necessary processing.
[0692] "Emotional information" refers to data that analyzes audio information and indicates the speaker's psychological and emotional state.
[0693] "Display means" refers to a device or method for visually displaying digital information.
[0694] "Adjustment means" refers to a process or device that modifies the speed and timbre of audio information to enable more appropriate transmission.
[0695] "Communication means" means technology or equipment for transmitting coordinated voice information to other devices or equipment.
[0696] A "hearing assistance device" refers to a device that outputs audio data in a format suitable for hearing, helping people with hearing impairments to hear sounds.
[0697] "User" refers to an individual or group that uses the system.
[0698] "Display characteristics" refer to elements such as font size and color that visually represent data formats.
[0699] This invention relates to a system that uses a specific voice processing algorithm to effectively analyze a user's voice information and obtain desired information. This system converts the voice information into text format and further extracts emotional information, enabling display and sound adjustment.
[0700] Users can speak to the system through a voice input device. This voice input device consists of a terminal including a standard microphone. The terminal sends the acquired voice information to a server. The server uses speech recognition software to convert the voice information into a data format. Tools such as the Google Cloud Speech-to-Text API are used for this process. Furthermore, the server uses an emotion recognition system to extract emotion information from the voice information. A general-purpose emotion analysis engine is a possible technology used here.
[0701] Based on the extracted emotional information, the display device on the terminal shows text data to the user, visually representing emotions with appropriate colors and font sizes. The server also processes the speed and tone of the audio information and transmits the adjusted audio information to the hearing assistance device via wireless communication. This makes the audio clearer and easier to understand.
[0702] For example, if a user says, "I'm looking forward to tomorrow's presentation," the server transcribes this phrase into text and further recognizes the positive emotion. The display device shows this text in a brighter color, and the server adjusts the tone of voice slightly higher before transmitting it.
[0703] An example of a prompt message is: "Based on the voice information obtained from the user, please extract emotional information and provide it as text data. Also, please suggest adjustments to the text display color and voice tone." This enables emotionally rich and intuitive communication for the user.
[0704] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0705] Step 1:
[0706] The user speaks through a voice input device. The terminal uses a microphone to acquire the user's voice information. The input is the user's voice waveform, and the output is digital voice data. The terminal then performs a specific action to temporarily store this voice data in preparation for subsequent processing.
[0707] Step 2:
[0708] The terminal sends the acquired audio data to the server. Here, the terminal uses an internet connection to transfer data to the server. The input is audio data in digital format, and the output is the data sent to the server. This process involves specific actions using protocols for secure and high-speed data transfer.
[0709] Step 3:
[0710] The server uses speech recognition software to convert audio data into text format. The input is digital audio data received by the server, and the output is text data. In this process, the server uses the Google Cloud Speech-to-Text API and other tools to analyze the characteristics of the audio waveform and perform the specific operations of converting it into text.
[0711] Step 4:
[0712] The server uses an emotion recognition engine to extract emotional information from audio data. The input is digital audio data received by the server, and the output is an emotion label (e.g., "happy," "sad"). The server performs specific actions to identify the speaker's emotions by analyzing the tone, pitch, and speed of the speech.
[0713] Step 5:
[0714] The terminal processes data to change how text data is displayed based on emotional information. Input consists of text data and emotional information, while output is a visually altered text display. The display device performs specific actions, such as changing the color of the text, according to the emotional information.
[0715] Step 6:
[0716] The server adjusts the speed and tone of the audio data based on emotional information. The input is the original audio data and emotional information, and the output is the adjusted audio data. The server uses audio editing software to perform the specific actions required to adjust the tone of the audio.
[0717] Step 7:
[0718] The server transmits the adjusted audio data to the user's hearing aid using a communication method. The input is the adjusted audio data, and the output is the audio received by the user's hearing aid. The server performs the specific operation of transmitting the audio using a wireless communication protocol.
[0719] (Application Example 2)
[0720] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0721] In face-to-face communication at physical stores, there is a problem in immediately grasping customer emotions and reflecting them in service. With conventional methods, voice data is limited to simple text conversion, and there are limitations in grasping customer emotions and nuances. As a result, it is difficult to provide rapid feedback for improving service quality and customer satisfaction.
[0722] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0723] In this invention, the server includes means for acquiring an acoustic signal, processing means for converting the acoustic signal into symbolic data, and emotion recognition means for analyzing the data to recognize emotions. This enables the server to instantly grasp the emotional state of the customer and provide feedback to improve the quality of service.
[0724] An "acoustic signal" is a physical representation of sound information collected from the environment or people.
[0725] "Symbolic data" refers to digital data that includes character information converted from acoustic signals.
[0726] "Processing means" refers to a device that has software or hardware functionality for converting acoustic signals into symbolic data.
[0727] "Display means" refers to displays and other display devices that visually present symbolic data.
[0728] An "emotion recognition system" is a system that has the function of analyzing and identifying emotional states from acoustic signals or symbolic data.
[0729] "Adjustment means" refers to a device or function for modifying acoustic signals or display content based on recognized emotions.
[0730] "Communication means" refers to networks and devices used to transfer processed information to external devices.
[0731] An "external device" is another electronic device that can receive and process acoustic signals and data.
[0732] "Visual information" refers to visible data such as characters, colors, and shapes displayed by a system.
[0733] A "user" is a person or organization that uses the system.
[0734] The system for carrying out this invention requires several main components. The server uses a smartphone or other voice input device to acquire acoustic signals. This allows the system to collect voices spoken by the user from the environment. The voice signals are converted into symbolic data using the Google Cloud Speech-to-Text API. This converted data is visualized by a display means. The display means can be a general display or a mobile screen.
[0735] Furthermore, the server uses the Emotion Recognition SDK to recognize emotions from symbolic data. This emotion recognition mechanism analyzes the tone and speed of the voice to identify emotions such as positive, negative, and neutral. Based on this information, the server uses an adjustment mechanism to change the display content and color tone of the symbolic data, providing visual feedback that fits the customer's emotions.
[0736] The communication method transmits processed data to external devices, such as smartphones or tablets, and provides feedback to service providers and in-store staff through these devices. In this way, customer needs can be met in real time, and the quality of service can be improved.
[0737] For example, when a user looks at a product in a store and makes a comment, that comment is captured on their smartphone. The system then performs emotion recognition and notifies the store staff that "the customer is showing interest in the product."
[0738] Examples of prompts for using generative AI models are as follows:
[0739] "The customer is speaking into their smartphone. We analyze their tone and speed of voice to identify their emotions and display a notification in the app."
[0740] This invention provides a useful means to optimize the customer experience in physical stores and to provide services efficiently and effectively.
[0741] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0742] Step 1:
[0743] The terminal acquires the user's speech as an acoustic signal from the microphone. The input is ambient sound, and the output is digitized acoustic data. This data is transmitted to the processing unit.
[0744] Step 2:
[0745] The server converts the received audio signal into symbolic data using the Google Cloud Speech-to-Text API. The input is digital audio data, and the output is symbolic data in text format. A speech recognition algorithm is used for this conversion.
[0746] Step 3:
[0747] The server uses a generative AI model to perform emotion recognition from the content of symbolic data. The input is the symbolic data obtained in step 2, and the output is emotion information (e.g., positive, negative, neutral). The emotion engine achieves this by analyzing the tone and content of the speech.
[0748] Step 4:
[0749] The server dynamically adjusts the display format of symbolic data using emotional information. Input consists of symbolic data and emotional information, while output is display data as visual information for the user. Specifically, this includes actions such as changing text color and font size.
[0750] Step 5:
[0751] The communication system transmits pre-tuned acoustic signals and symbolic data to an external device. The input is pre-tuned data, and the output is information displayed on a receiving device such as a smartphone or tablet. This allows in-store staff to respond to customers in a way that is tailored to their emotions.
[0752] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0753] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0754] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0755] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0756] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0757] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0758] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0759] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0760] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0761] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0762] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0763] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0764] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0765] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0766] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0767] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0768] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0769] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0770] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0771] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0772] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0773] The following is further disclosed regarding the embodiments described above.
[0774] (Claim 1)
[0775] A device that acquires audio data,
[0776] A processing device that converts the aforementioned audio data into text data,
[0777] A display device that displays the aforementioned text data,
[0778] An adjustment device for adjusting the speed and tone of the aforementioned audio data,
[0779] A communication device that transmits the adjusted voice data,
[0780] A communication system that includes this.
[0781] (Claim 2)
[0782] The system according to claim 1, characterized in that the communication device transmits the adjusted audio data to the hearing aid via wireless communication.
[0783] (Claim 3)
[0784] The system according to claim 1, characterized in that the display device displays the text data with the font size and color changed according to the user's settings.
[0785] "Example 1"
[0786] (Claim 1)
[0787] Means for acquiring audio information,
[0788] A device means for converting the aforementioned audio information into text information,
[0789] Means for visually displaying the aforementioned textual information,
[0790] A device means for adjusting the speed and tone of the aforementioned audio information,
[0791] A communication means for transmitting the adjusted voice information,
[0792] A system that includes this.
[0793] (Claim 2)
[0794] The system according to claim 1, characterized in that the communication means transmits the adjusted voice information to the hearing assistance device via wireless communication.
[0795] (Claim 3)
[0796] The system according to claim 1, characterized in that the visual display means displays the character information with a modified font size and color according to the user's settings.
[0797] "Application Example 1"
[0798] (Claim 1)
[0799] Means for acquiring audio data,
[0800] Means for converting the aforementioned audio data into text data,
[0801] means for displaying the aforementioned text data,
[0802] Means for adjusting the speed and tone of the aforementioned audio data,
[0803] means for transmitting the adjusted audio data,
[0804] Means for providing the aforementioned text data and adjusted audio data to the guidance device,
[0805] A system that includes this.
[0806] (Claim 2)
[0807] The system according to claim 1, characterized in that the transmitting means transmits the adjusted audio data to an auxiliary device via wireless communication.
[0808] (Claim 3)
[0809] The system according to claim 1, characterized in that the display means displays the text data in a format that changes according to the recipient's settings.
[0810] "Example 2 of combining an emotion engine"
[0811] (Claim 1)
[0812] Means for acquiring audio information,
[0813] Processing means for converting the aforementioned audio information into a data format,
[0814] Means for displaying the aforementioned data format,
[0815] A means for extracting emotional information from the aforementioned audio information,
[0816] Means for changing the display method of the data format based on the aforementioned emotional information,
[0817] Means for adjusting the speed and tone of the aforementioned audio information,
[0818] A communication means for transmitting the adjusted voice information,
[0819] A system that includes this.
[0820] (Claim 2)
[0821] The system according to claim 1, characterized in that the communication means transmits the adjusted voice information to the hearing assistance device via wireless communication.
[0822] (Claim 3)
[0823] The system according to claim 1, characterized in that the display means displays the data format by changing the display characteristics according to the user's settings.
[0824] "Application example 2 when combining with an emotional engine"
[0825] (Claim 1)
[0826] Means for acquiring acoustic signals,
[0827] Processing means for converting the aforementioned acoustic signal into symbolic data,
[0828] A display means for displaying the aforementioned symbolic data,
[0829] An emotion recognition means that analyzes the aforementioned data to recognize emotions,
[0830] An adjustment means for adjusting display data and sound signals based on the aforementioned emotions,
[0831] Communication means for transmitting the adjusted acoustic signal to an external device,
[0832] A system that includes this.
[0833] (Claim 2)
[0834] The system according to claim 1, characterized in that the communication means transmits the adjusted acoustic signal to the hearing aid via wireless communication.
[0835] (Claim 3)
[0836] The system according to claim 1, characterized in that the display means displays the symbol data by changing the size and color of the visual information according to the user's settings. [Explanation of symbols]
[0837] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A device that acquires audio data, A processing device that converts the aforementioned audio data into text data, A display device that displays the aforementioned text data, An adjustment device for adjusting the speed and tone of the aforementioned audio data, A communication device that transmits the adjusted voice data, A communication system that includes this.
2. The system according to claim 1, characterized in that the communication device transmits the adjusted audio data to the hearing aid via wireless communication.
3. The system according to claim 1, characterized in that the display device displays the text data with the font size and color changed according to the user's settings.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A