System

The system addresses communication challenges for the hearing-impaired by analyzing sound environments, optimizing hearing aids, converting audio to text, and providing sign language commentary, enhancing daily life experiences.

JP2026014974APending Publication Date: 2026-01-29SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024116448
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Hearing-impaired and hard-of-hearing individuals face challenges in supplementing auditory information with vision alone, leading to communication difficulties, increased psychological and cognitive risks, and inconvenience in daily life, with existing methods failing to adapt appropriately to individual situations.

Method used

A system that analyzes sound environments in real time to recommend hearing aid settings, converts audio data into text, provides video content with sign language commentary, and detects specific sounds or words, using a combination of wearable devices and a central server to enhance communication and information access.

Benefits of technology

Improves the quality of life for hearing-impaired individuals by enabling efficient communication and information acquisition through real-time sound analysis, hearing aid optimization, and provision of barrier-free information, reducing risks and enhancing safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026014974000001_ABST
    Figure 2026014974000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for analyzing a sound environment in real-time and recommending appropriate hearing aid settings; means for converting audio data to text; means for providing video content with sign language commentary; and means for detecting and notifying specific sounds and words.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] The purpose of this invention is to compensate for communication loss for the hearing-impaired and hard of hearing, to avoid inconvenience and danger in daily life, and to mitigate the increased psychological and cognitive risks caused by reduced communication with those around them. Currently, it is difficult to supplement auditory information with vision alone, and it is difficult to always choose an appropriate method according to each person's situation and ability. We provide a system that solves these problems and improves the quality of life of individuals. [Means for solving the problem]

[0005] The present invention solves the above-mentioned problems with a system that includes the following means: a means for analyzing the sound environment in real time and recommending appropriate hearing aid settings, a means for converting audio data into text, a means for providing video content with sign language commentary, and a means for detecting and notifying specific sounds and words. The system also includes means for transmitting sound environment data to a server and receiving hearing aid settings from the server, and means for transmitting collected audio data to a speech recognition engine and returning it to a terminal as text data. This configuration helps the hearing-impaired and hard-of-hearing individuals communicate and acquire information smoothly in their daily lives, improving their quality of life.

[0006] The "sound environment" refers to the entire auditory information, including surrounding environmental sounds, conversations, noises, alarm sounds, etc.

[0007] "Analyzing in real time" means instantly collecting data on the sound environment and analyzing it without delay.

[0008] "Hearing aid settings" refers to the volume and frequency adjustment parameters that allow a hearing aid to function optimally in a particular sound environment.

[0009] "Audio data" refers to sound information collected by a microphone or recording device expressed as digital data.

[0010] "Converting to text" refers to analyzing audio data and expressing its contents as text information.

[0011] "Video content with sign language commentary" refers to video information that includes audio information and supplementary explanations in sign language.

[0012] The "specific sounds and words" refer to sounds and words that have specific meanings preset by the user and that express situations that require attention.

[0013] "Notifying" refers to notifying the user of a specific event or sound by visual, vibration, audio, or other means when that event or sound occurs.

[0014] "Server" refers to a computer system placed on a network for the purpose of analyzing data and providing information.

[0015] "Terminal" refers to a device worn or used by a user that has functions such as collecting, notifying, and displaying audio.

[0016] A "voice recognition engine" refers to software or hardware that analyzes speech and converts it into text data. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11]FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0025] [First embodiment]

[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0038] The present invention is a system for solving various problems that hearing-impaired and hard-of-hearing people encounter in their daily lives, and specifically includes a means for analyzing the sound environment in real time and recommending appropriate hearing aid settings, a means for converting audio data into text, a means for providing video content with sign language commentary, and a means for detecting and notifying specific sounds and words. Each function of this system is described in detail below.

[0039] Automatic hearing aid optimization

[0040] When a user puts on the wearable device and turns it on, the device automatically collects the surrounding sound environment using a microphone. The sound environment data is analyzed by the device to identify sound types such as speech, noise, and alarms. The analyzed data is sent to a server, which calculates the most suitable hearing aid settings. The calculated settings are sent back to the device and suggested to the user. If the user approves, the device automatically adjusts the hearing aid settings.

[0041] Utilizing voice recognition technology

[0042] When a user wears the device and activates the voice recognition function, the device continuously collects surrounding sounds. The collected voice data is sent to a server in real time, where a voice recognition engine analyzes the voice data and converts it into text. The converted text data is displayed on the device's display, allowing the user to visually check the content of conversations around them.

[0043] Barrier-free information provided

[0044] When a user uses a device in a station or public facility, the device collects public information such as announcements and guidance announcements. The collected information is sent to a server, which searches a database of video content with sign language commentary. When relevant video content is found, the server returns a link to the device and notifies the user. When the user selects the link, the video with sign language commentary is played on the device.

[0045] Personal Settings Function

[0046] The user can set specific sounds or words (such as their name or "danger") in advance, and the device will monitor the surrounding sounds in real time. When the set sound or word is detected, the corresponding data is sent to the server. After confirmation, the server sends a notification instruction to the device. The device will then alert the user using a visual alert or vibration.

[0047] Specific examples

[0048] 1. Examples of automatic hearing aid optimization:

[0049] When a user is talking with a friend at a coffee shop, the device recognizes the surrounding noise and the conversation and sends the data to the server. The server calculates the settings to reduce noise and emphasize the conversation and sends them back to the device. The device then displays a message to the user asking, "Would you like to change the settings to reduce noise and emphasize the conversation?" and the settings are applied if the user approves.

[0050] 2. Examples of voice recognition technology:

[0051] When a user is participating in a meeting at work, the device sends the contents of the meeting to a speech recognition engine, which converts the conversation into text and displays "About the next presentation" on the device display, allowing the user to easily understand the content.

[0052] 3. Examples of providing barrier-free information:

[0053] When the user is at a station, the device collects the station's departure announcements, and the server finds and sends a video with sign language explanations to the device. The user views the video with sign language explanations, which include the message "The next train departs at 10:25," and accurately understands the information.

[0054] 4. Examples of personalization features:

[0055] When a user is relaxing at home, a specific word, "danger," is set, and if that word is heard from outside, the device will vibrate to notify the user and display "Danger detected" on the screen to warn them.

[0056] In this way, the present invention provides an effective means for improving the quality of life of the deaf and hard of hearing.

[0057] The processing flow will be explained below.

[0058] Automatic hearing aid optimization

[0059] Step 1:

[0060] The device detects that the user is wearing a wearable device and has it turned on.

[0061] Step 2:

[0062] The device collects the surrounding sound environment in real time using its built-in microphone and generates sound environment data.

[0063] Step 3:

[0064] The device analyzes the sound environment data collected and identifies conversation sounds, noise, alarm sounds, etc.

[0065] Step 4:

[0066] The device sends the analysis results to the server.

[0067] Step 5:

[0068] The server analyzes the received sound environment data and calculates the optimal hearing aid settings.

[0069] Step 6:

[0070] The server returns the calculation results to the terminal.

[0071] Step 7:

[0072] The device will prompt the user, "There are new environment-based hearing aid settings. Would you like to apply them?"

[0073] Step 8:

[0074] The user selects "Yes" or "No."

[0075] Step 9:

[0076] If the user selects "Yes," the device will automatically change the hearing aid settings.

[0077] Step 10:

[0078] The device will notify the user that the new settings have been applied.

[0079] Utilizing voice recognition technology

[0080] Step 1:

[0081] The user puts on the device and enables the speech recognition feature.

[0082] Step 2:

[0083] The device continuously collects surrounding sounds using the built-in microphone.

[0084] Step 3:

[0085] The device sends the collected voice data to the server.

[0086] Step 4:

[0087] The server analyzes the received voice data and converts it into text using a voice recognition engine.

[0088] Step 5:

[0089] The server returns the converted text data to the terminal.

[0090] Step 6:

[0091] The device displays the received text on the display.

[0092] Step 7:

[0093] The device updates the text display in real time as new audio data is collected.

[0094] Barrier-free information provided

[0095] Step 1:

[0096] The user wears the device and enables the barrier-free information provision function.

[0097] Step 2:

[0098] The device uses a built-in microphone to collect public announcements and guidance broadcasts.

[0099] Step 3:

[0100] The device sends the collected voice data to the server.

[0101] Step 4:

[0102] The server analyzes the audio data and searches a database for relevant video content with sign language commentary.

[0103] Step 5:

[0104] The server sends a link to the appropriate video content back to the device.

[0105] Step 6:

[0106] The device notifies the user, "There is a related sign language explanation video. Would you like to play it?"

[0107] Step 7:

[0108] The user selects "Yes" or "No."

[0109] Step 8:

[0110] If the user selects "Yes," the device will open the video link and play the sign language explanation video.

[0111] Personal Settings Function

[0112] Step 1:

[0113] The user puts on the device and accesses the personalization screen.

[0114] Step 2:

[0115] The user sets and saves a specific sound or word (such as a name or "danger").

[0116] Step 3:

[0117] The device collects surrounding sounds in real time using a built-in microphone.

[0118] Step 4:

[0119] The device continuously compares the sounds it collects with preset sounds and words.

[0120] Step 5:

[0121] When the device detects a specific sound or word, it sends that data to the server.

[0122] Step 6:

[0123] The server analyzes the received data and confirms the detection results.

[0124] Step 7:

[0125] The server sends a notification instruction to the terminal.

[0126] Step 8:

[0127] The device will notify the user with a visual alert and / or vibration.

[0128] Step 9:

[0129] The device will display "A specific sound has been detected. Please be careful."

[0130] Example 1

[0131] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0132] This invention aims to solve various problems that the hearing impaired and hard of hearing face in their daily lives. Specifically, it aims to improve the quality of life for the hearing impaired by providing technology that efficiently analyzes surrounding sounds and automatically configures optimal hearing aids, as well as technology that converts speech to text, provides video content that incorporates sign language, and detects and notifies users of specific sounds and words.

[0133] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0134] In this invention, the server includes a means for analyzing the sound environment in real time and recommending appropriate hearing aid settings, a means for converting recorded audio data into text, and a means for providing video content with sign language commentary, thereby enabling hearing-impaired or hard-of-hearing people to efficiently grasp the surrounding sound environment and receive appropriate hearing assistance.

[0135] "Sound environment" refers to the state and situation of sound waves such as surrounding voices and noise.

[0136] "Analyzing in real time" means analyzing data immediately after it is collected, without any time delay.

[0137] A "hearing assistive device" is a device that assists the hearing of people with hearing loss or hearing impairments, and has the function of amplifying sound or emphasizing specific sounds.

[0138] "Audio recording data" refers to data that has been digitized and saved from audio collected by a microphone or other device.

[0139] "Converting to text" refers to the process of converting data such as audio and video into text information.

[0140] "Video content with sign language commentary" refers to videos or video materials that include visual sign language commentary.

[0141] "Detecting specific sounds or words" refers to the recognition system identifying and detecting predefined sounds or words.

[0142] "Notify" refers to informing a user of specific information or events.

[0143] "Device" refers to electronic equipment or devices, including hearing aids and wearable devices.

[0144] "Central server" refers to the main computer system that receives data from multiple devices, processes and analyzes it, and returns the results.

[0145] "Collect" refers to gathering data or information.

[0146] "Transmit" refers to sending data or information to another device or system.

[0147] "Calculating" refers to the process of deriving a specific result based on analytical results or algorithms.

[0148] "Automatically adjust" means that the device is automatically set to the optimum state without requiring any user operation.

[0149] "Public facilities" refers to buildings and places for use by the general public, including stations and libraries.

[0150] "Linking" means associating and making accessible particular information or content.

[0151] "Monitor" refers to watching for a particular condition or state.

[0152] This system analyzes the sound environment, converts speech into text, provides video content with sign language commentary, and detects and notifies specific sounds and words to solve various problems faced by the hearing impaired and hard of hearing in their daily lives. The system of this invention functions mainly through cooperation between the device (wearable terminal) and a central server.

[0153] Hardware and software used

[0154] 1. Device (wearable device)

[0155] microphone

[0156] Built-in processor

[0157] Communication modules (Wi-Fi, 5G, etc.)

[0158] display

[0159] Vibration Motor

[0160] battery

[0161] 2. Central Server

[0162] Generative AI Models

[0163] Speech Recognition Engine

[0164] Database (including video content with sign language commentary)

[0165] Communication Interface

[0166] Program processing

[0167] Analysis of sound environments and optimization of hearing aids

[0168] When a user puts on the device and turns it on, the device automatically collects the surrounding sound environment using a microphone. The sound environment data is analyzed by the device to identify sound types such as speech, noise, and alarms. This data is sent to a central server, where a generative AI model calculates optimal hearing aid settings. The calculation results are sent back to the device and recommended to the user. If the user approves, the device automatically adjusts the hearing aid settings.

[0169] Examples:

[0170] When a user is talking with a friend at a coffee shop, the device recognizes the surrounding noise and the conversation and sends the data to a central server. The server calculates the settings to reduce noise and enhance the conversation and sends them back to the device. The device then displays a message asking, "Do you want to change the settings to reduce noise and enhance the conversation?" and the settings are applied if the user confirms.

[0171] Utilizing voice recognition technology

[0172] When a user enables the device's voice recognition function, the device will continuously collect surrounding sounds. The collected voice data is sent to a central server in real time, where the voice recognition engine converts the speech into text. The converted text data is then displayed on the device's display, allowing the user to visually check the content of the conversations around them.

[0173] Examples:

[0174] When a user is participating in a meeting at work, the device sends the contents of the meeting to a speech recognition engine, and the converted text is displayed on the device's display as "About the next presentation," allowing the user to easily understand the content.

[0175] Providing video content with sign language commentary

[0176] When a user uses a device in a station or public facility, the device collects public information such as announcements and guidance announcements. The collected information is sent to a central server, which searches a database of videos with sign language descriptions. When relevant video content is found, the central server sends a link back to the device and notifies the user. When the user selects the link, the video with sign language descriptions is played on the device.

[0177] Examples:

[0178] When a user is at a station, the terminal collects the station's departure announcements, and the central server finds the video with sign language explanations and sends it to the terminal. The user views the video with sign language explanations, which include the message "The next train departs at 10:25," and accurately understands the information.

[0179] Detecting and announcing specific sounds and words

[0180] Users can pre-set specific sounds or words (e.g., "danger"), and the device will monitor surrounding sounds in real time. When the device detects a specific sound or word, the corresponding data is sent to a central server, which then confirms the detection and sends a notification instruction to the device. The device will then alert the user with a visual alert or vibration.

[0181] Examples:

[0182] If the user is relaxing at home and the word "danger" is set and hears that word from outside, the device will vibrate to notify the user and display "Danger detected" on the screen to warn them.

[0183] As described above, the present invention integrates multiple technical means to provide an effective method for the hearing impaired and hard of hearing to efficiently obtain information necessary in daily life.

[0184] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0185] Analysis of sound environments and optimization of hearing aids

[0186] Step 1:

[0187] When a user wears the device and turns it on, the device automatically collects the surrounding sound environment with its microphone. The input here is the surrounding sound environment, and the output is digitized sound environment data. The device's microphone converts analog audio into digital data.

[0188] Step 2:

[0189] The device analyzes the sound environment data collected in real time to identify the type of sound, such as speech, noise, or alarm sounds. The input is digitized sound environment data, and the output is information about the identified sound type (voice pattern data). The device's built-in processor runs an algorithm to analyze the data.

[0190] Step 3:

[0191] The device sends the analyzed voice pattern data to a central server. The input is the voice pattern data, and the output is the transmitted data. The data is sent to a cloud server using the device's communication module (Wi-Fi, 5G, etc.).

[0192] Step 4:

[0193] Based on the voice pattern data received by the server, a generative AI model calculates the optimal hearing aid settings. The input is voice pattern data, and the output is optimal setting data. The generative AI model analyzes the voice data and derives the appropriate parameters.

[0194] Step 5:

[0195] The server returns the optimal setting data to the terminal. The input is the optimal setting data, and the output is the setting data sent to the terminal. The data is sent using the server's communication interface.

[0196] Step 6:

[0197] The device receives the optimal setting data and displays the suggestion to the user. The input is the received setting data, and the output is the display of the suggestion. The device display shows "Would you like to change the settings to reduce noise and enhance speech?"

[0198] Step 7:

[0199] When the user presses the approval button, the device automatically adjusts the settings of the hearing aid. The input is the user's approval, and the output is the adjusted settings of the hearing aid. The device provides feedback to the user by vibration or sound.

[0200] Utilizing voice recognition technology

[0201] Step 1:

[0202] When a user enables the voice recognition function of the device, the device continuously collects surrounding sounds. The input is the surrounding sounds, and the output is digitized voice data. The microphone collects the voice and converts it into digital data.

[0203] Step 2:

[0204] The terminal transmits the collected voice data to a central server in real time. The input is the digitized voice data, and the output is the transmitted voice data. The data is transmitted using a communication module.

[0205] Step 3:

[0206] The server analyzes the voice data using a voice recognition engine and converts it into text data. The input is voice data and the output is text data. The voice recognition engine analyzes the voice pattern and generates the corresponding text.

[0207] Step 4:

[0208] Text data is sent from the central server to the terminal. The input is the text data, and the output is the text data sent to the terminal. The data is sent using the server's communication interface.

[0209] Step 5:

[0210] Text data is displayed on the terminal display. The input is the received text data, and the output is the text content displayed on the display. The user can check the content of the surrounding conversation by looking at the displayed text.

[0211] Providing video content with sign language commentary

[0212] Step 1:

[0213] When a user uses a device in a station or public facility, the terminal collects announcements and guidance announcements. The input is the public facility announcements, and the output is digitized voice data. The microphone collects the voice and converts it into digital data.

[0214] Step 2:

[0215] The collected voice data is transmitted to a central server. The input is the digitized voice data and the output is the transmitted voice data. A communication module is used to transmit the data to the server.

[0216] Step 3:

[0217] The server analyzes the audio data and searches a database of video content with sign language commentary. The input is the audio data, and the output is a link to related video content. The server's generative AI model analyzes the audio and finds related videos from the database.

[0218] Step 4:

[0219] The server returns a link to the associated video content to the terminal. The input is the link to the video content, and the output is the link sent to the terminal. The link is sent using a communication interface.

[0220] Step 5:

[0221] When the user selects a link, the device plays a video with sign language descriptions. The input is the link selected by the user and the output is the video that is played. The video with sign language descriptions appears on the device's display.

[0222] Detecting and announcing specific sounds and words

[0223] Step 1:

[0224] The user sets specific sounds or words (e.g., "danger") in advance on the device. The input is the user's settings for the sounds or words, and the output is the saved setting data. The sounds or words are entered using the device's setting screen and saved in a database.

[0225] Step 2:

[0226] The device monitors the surrounding sound in real time. The input is the surrounding sound, and the output is digitized audio data. The microphone collects the audio and converts it into digital data.

[0227] Step 3:

[0228] When a preset sound or word is detected, the device sends data to a central server. The input is the detected voice data, and the output is the transmitted data. The data is sent to the server using a communication module.

[0229] Step 4:

[0230] The server checks the received data and sends a notification instruction to the device. The input is the received voice data, and the output is the notification instruction. The server's generative AI model analyzes the data and determines the appropriate response.

[0231] Step 5:

[0232] The device alerts the user with a visual alert or vibration. The input is a notification instruction, and the output is a visual alert or vibration notification. The device activates the vibration motor and displays "Danger detected" on the display.

[0233] (Application example 1)

[0234] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0235] The present invention aims to solve the problem that hearing-impaired and hard-of-hearing people have difficulty recognizing voices and grasping important information in daily life and in certain situations (e.g., when using self-driving vehicles). In particular, there is a problem that emergency situations and important announcements are difficult to convey to hearing-impaired and hard-of-hearing people. This may result in situations where the safety and comfort of users are lacking.

[0236] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0237] In this invention, the server includes a means for analyzing the sound environment in real time and recommending appropriate hearing aid settings, a means for converting voice data into text, and a means for providing video content with sign language commentary. This allows hearing-impaired and hard-of-hearing individuals to properly understand the surrounding sound environment and automatically optimize their hearing aid settings. Furthermore, by providing a means for collecting voice data in real time and notifying the user based on specific keywords, it becomes possible to reliably communicate emergencies and important announcements. Furthermore, by including a means for displaying voice recognition results on an in-vehicle display, users can visually confirm surrounding voice information, thereby improving safety and comfort when using autonomous vehicles.

[0238] "Analyzing the sound environment in real time and recommending appropriate hearing aid settings" means continuously monitoring the sound environment using sensors and microphones, classifying the type and level of sound using an analysis device, and providing optimal hearing aid settings based on that.

[0239] "Converting voice data to text" refers to the process of analyzing collected voice data using speech recognition technology and converting it into a corresponding text format.

[0240] "Providing video content with sign language commentary" is a service that allows deaf and hard of hearing people to obtain information visually by providing video information with sign language interpretation.

[0241] "Detecting and notifying specific sounds and words" refers to a mechanism that detects specific pre-set sounds or utterances and notifies the user.

[0242] "Collecting voice data in real time and notifying the user based on specific keywords" means accumulating voice data in real time, identifying specific keywords in the data, and providing the user with appropriate alerts.

[0243] "Displaying voice recognition results on an in-vehicle display" means displaying text data obtained through voice recognition technology on a display inside an autonomous vehicle to provide information visually to the user.

[0244] The present invention is a system for solving various problems that hearing-impaired and hard-of-hearing people encounter in their daily lives and when using self-driving vehicles. This system includes a means for analyzing the sound environment in real time and recommending appropriate hearing aid settings, a means for converting voice data into text, a means for providing video content with sign language commentary, and a means for detecting and notifying specific sounds and words.

[0245] System configuration

[0246] The system consists of the following main components:

[0247] Hardware

[0248] 1. Microphone: A device that collects ambient sounds, allowing for real-time monitoring of the sound environment.

[0249] 2. Server: Analyzes the sound data and provides instructions for optimal settings for the hearing aid and text conversion.

[0250] 3. Autonomous vehicle display: A device for displaying voice recognition results and system notifications.

[0251] 4. Vibration motor: To notify the user when a specific sound or keyword is detected.

[0252] software

[0253] 1. Speech recognition engine (e.g., Google Speech Recognition API): Collects speech in real time and converts it into text.

[0254] 2. Analysis algorithm: Analyzes the sound environment and recommends appropriate hearing aid settings.

[0255] 3. Notification system: Detects specific keywords or voice anomalies and notifies the user.

[0256] System Operation

[0257] 1. Real-time analysis of sound environments

[0258] When the user operates the system, the microphone begins to collect the surrounding sound environment. The collected sound environment data is sent to the server, where an analysis algorithm identifies the type of sound, such as speech, noise, or alarm. Based on the analyzed data, the server calculates the optimal hearing aid settings and sends them back to the device. If the user approves, the device automatically adjusts the hearing aid settings.

[0259] 2. Use of voice recognition technology

[0260] When the user wears the device and activates the voice recognition function, the device continuously collects surrounding sounds. The collected voice data is sent to a server in real time, where a voice recognition engine analyzes the voice data and converts it into text. The converted text data is displayed on the vehicle's display, allowing the user to visually check the content of surrounding conversations and announcements.

[0261] 3. Providing barrier-free information

[0262] When a user is using an autonomous vehicle, the device collects information such as in-car announcements and emergency alerts. The collected information is sent to a server, which searches a database of video content with sign language commentary. When relevant video content is found, the server returns a link to the device and notifies the user. When the user selects the link, the video with sign language commentary is played on the in-car display.

[0263] 4. Personal settings function

[0264] The user can set specific sounds or words (such as their name or "danger") in advance, and the device will monitor the surrounding sounds in real time. When the set sound or word is detected, the corresponding data is sent to the server. After confirmation, the server sends a notification instruction to the device. The device will then alert the user using a visual alert or vibration.

[0265] Specific examples

[0266] 1. A concrete example of automatic hearing aid optimization

[0267] When a user is talking with a friend at a coffee shop, the device recognizes the surrounding noise and the conversation and sends the data to the server. The server calculates the settings to reduce noise and emphasize the conversation and sends them back to the device. The device then displays a message to the user asking, "Would you like to change the settings to reduce noise and emphasize the conversation?" and the settings are applied if the user approves.

[0268] 2. Specific examples of voice recognition technology

[0269] When a user is participating in a meeting at work, the device sends the contents of the meeting to a speech recognition engine, which converts the conversation into text and displays "About the next presentation" on the device display, allowing the user to easily understand the content.

[0270] 3. Examples of providing barrier-free information

[0271] When the user is at a station, the device collects the station's departure announcements, and the server finds and sends a video with sign language explanations to the device. The user views the video with sign language explanations, which include the message "The next train departs at 10:25," and accurately understands the information.

[0272] 4. Examples of personal settings functions

[0273] When a user is relaxing at home, a specific word, "danger," is set, and if that word is heard from outside, the device will vibrate to notify the user and display "Danger detected" on the screen to warn them.

[0274] Example prompts for generative AI models

[0275] "You are designing a self-driving car system for the hearing impaired. The system uses speech recognition to convert emergency situations and announcements into text and notify the user. Extract the emergency content from the following audio data and display it as text."

[0276] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0277] Step 1:

[0278] The user operates the system. The user gets into the autonomous vehicle and a microphone connected to the terminal collects surrounding sounds in real time. The collected sound data is input into the terminal.

[0279] Step 2:

[0280] The device analyzes the collected sound data and classifies the sound environment. The device's internal analysis algorithm classifies the sound data into categories such as speech, noise, and alarm sounds. This analyzed data is sent to the server.

[0281] Step 3:

[0282] The server receives the sound environment data and calculates the optimal hearing aid settings. Based on the analyzed sound data, calculations are performed to automatically optimize the hearing aid settings. These optimized settings are then sent from the server to the device.

[0283] Step 4:

[0284] The device presents the hearing aid settings to the user. The device asks the user whether to apply the new settings and obtains their approval. When the user presses the approval button, the device changes the hearing aid settings.

[0285] Step 5:

[0286] The device continuously collects voice data and sends it to a voice recognition engine. Based on the voice data input, real-time voice recognition processing is performed and the data is converted into text. This text data is then sent back to the device.

[0287] Step 6:

[0288] The device displays the voice recognition results on the vehicle display, and text data is output to the display so that the user can visually check the content of surrounding conversations and announcements.

[0289] Step 7:

[0290] The device monitors specific keywords and sounds to detect abnormalities. When a pre-defined keyword (e.g., "danger" or "emergency") is detected, the data is sent to the server. After the server confirms the detection, it sends a notification instruction to the device.

[0291] Step 8:

[0292] The device will alert the user. If an abnormal sound or specific keyword is detected, the device will alert the user with a visual alert or vibration notification.

[0293] The above process will create a system that improves the safety and convenience of people who are deaf or hard of hearing when using self-driving vehicles.

[0294] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0295] The present invention is a system for solving various problems encountered in daily life by people with hearing impairments or hard of hearing. Specifically, it includes a means for analyzing the sound environment in real time and recommending appropriate hearing aid settings, a means for converting audio data into text, a means for providing video content with sign language commentary, and a means for detecting and notifying specific sounds or words. Furthermore, by combining it with an emotion engine that recognizes and analyzes the user's emotions, further user assistance is provided. Each function of this system is described in detail below.

[0296] Automatic hearing aid optimization

[0297] When a user puts on the wearable device and turns it on, the device automatically collects the surrounding sound environment using a microphone. The sound environment data is analyzed by the device to identify sound types such as speech, noise, and alarms. The analyzed data is sent to a server, which calculates the most suitable hearing aid settings. The calculated settings are sent back to the device and suggested to the user. If the user approves, the device automatically adjusts the hearing aid settings.

[0298] Utilizing voice recognition technology

[0299] When a user wears the device and activates the voice recognition function, the device continuously collects surrounding sounds. The collected voice data is sent to a server in real time, where a voice recognition engine analyzes the voice data and converts it into text. The converted text data is displayed on the device's display, allowing the user to visually check the content of conversations around them.

[0300] Barrier-free information provided

[0301] When a user uses a device in a station or public facility, the device collects public information such as announcements and guidance announcements. The collected information is sent to a server, which searches a database of video content with sign language commentary. When relevant video content is found, the server returns a link to the device and notifies the user. When the user selects the link, the video with sign language commentary is played on the device.

[0302] Personal Settings Function

[0303] The user can set specific sounds or words (such as their name or "danger") in advance, and the device will monitor the surrounding sounds in real time. When the set sound or word is detected, the corresponding data is sent to the server. After confirmation, the server sends a notification instruction to the device. The device will then alert the user using a visual alert or vibration.

[0304] Utilizing the Emotion Engine

[0305] The emotion engine monitors the user's emotional state in real time while the user is wearing the device. The emotion engine analyzes the user's tone of voice and facial expressions (e.g., using the camera function) to recognize emotions such as stress, joy, and sadness. Once the emotional state is analyzed, the results are sent to the server, which determines the appropriate response and sends instructions back to the device.

[0306] Specific examples

[0307] 1. Examples of automatic hearing aid optimization:

[0308] When a user is talking with a friend at a coffee shop, the device recognizes the surrounding noise and the conversation and sends the data to the server. The server calculates the settings to reduce noise and emphasize the conversation and sends them back to the device. The device then displays a message to the user asking, "Would you like to change the settings to reduce noise and emphasize the conversation?" and the settings are applied if the user approves.

[0309] 2. Examples of voice recognition technology:

[0310] When a user is participating in a meeting at work, the device sends the contents of the meeting to a speech recognition engine, which converts the conversation into text and displays "About the next presentation" on the device display, allowing the user to easily understand the content.

[0311] 3. Examples of providing barrier-free information:

[0312] When the user is at a station, the device collects the station's departure announcements, and the server finds and sends a video with sign language explanations to the device. The user views the video with sign language explanations, which include the message "The next train departs at 10:25," and accurately understands the information.

[0313] 4. Examples of personalization features:

[0314] When a user is relaxing at home, a specific word, "danger," is set, and if that word is heard from outside, the device will vibrate to notify the user and display "Danger detected" on the screen to warn them.

[0315] 5. Examples of Emotion Engines:

[0316] If a user feels stressed during a meeting, the emotion engine analyzes the user's facial expressions and tone of voice to detect the stress level. The server recommends audio content with a relaxing effect, and the device notifies the user, "Would you like to play relaxing music?" and plays the music if the user agrees.

[0317] In this way, the present invention provides an effective means for improving the quality of life of the deaf and hard of hearing.

[0318] The processing flow will be explained below.

[0319] Automatic hearing aid optimization

[0320] Step 1:

[0321] The user puts on the wearable device and turns it on.

[0322] Step 2:

[0323] The device automatically collects the surrounding sound environment in real time using the built-in microphone.

[0324] Step 3:

[0325] The device analyzes the sound environment data collected and identifies the type of sound (conversation, noise, alarm, etc.).

[0326] Step 4:

[0327] The device sends the analysis results to the server.

[0328] Step 5:

[0329] The server analyzes the received sound environment data and calculates appropriate hearing aid settings.

[0330] Step 6:

[0331] The server returns the calculated hearing aid settings to the device.

[0332] Step 7:

[0333] The device will prompt the user, "There are new environment-based hearing aid settings. Would you like to apply them?"

[0334] Step 8:

[0335] The user selects "Yes" or "No."

[0336] Step 9:

[0337] If the user selects "Yes," the device will automatically change the hearing aid settings.

[0338] Step 10:

[0339] The device will notify the user that the new settings have been applied.

[0340] Utilizing voice recognition technology

[0341] Step 1:

[0342] The user puts on the device and enables the speech recognition feature.

[0343] Step 2:

[0344] The device continuously collects surrounding sounds using the built-in microphone.

[0345] Step 3:

[0346] The device sends the collected voice data to the server.

[0347] Step 4:

[0348] The voice data received by the server is analyzed using a voice recognition engine and converted into text.

[0349] Step 5:

[0350] The server returns the converted text data to the terminal.

[0351] Step 6:

[0352] The text data received by the terminal is displayed on the display.

[0353] Step 7:

[0354] The device continues to convert new voice data collected into text in real time, updating the text on the display.

[0355] Barrier-free information provided

[0356] Step 1:

[0357] The user wears the device and enables the barrier-free information provision function.

[0358] Step 2:

[0359] The device uses a built-in microphone to collect public announcements and guidance broadcasts.

[0360] Step 3:

[0361] The device sends the collected voice data to the server.

[0362] Step 4:

[0363] The server analyzes the audio data and searches a database for relevant video content with sign language commentary.

[0364] Step 5:

[0365] The server sends a link to the appropriate video content back to the device.

[0366] Step 6:

[0367] The device notifies the user, "There is a related sign language explanation video. Would you like to play it?"

[0368] Step 7:

[0369] The user selects "Yes" or "No."

[0370] Step 8:

[0371] If the user selects "Yes," the device will open the video link and play the sign language explanation video.

[0372] Personal Settings Function

[0373] Step 1:

[0374] The user puts on the device and accesses the personalization screen.

[0375] Step 2:

[0376] The user sets and saves a specific sound or word (such as a name or "danger").

[0377] Step 3:

[0378] The device collects surrounding sounds in real time using a built-in microphone.

[0379] Step 4:

[0380] The device continuously compares the sounds it collects with preset sounds and words.

[0381] Step 5:

[0382] When the device detects a specific sound or word, it sends that data to the server.

[0383] Step 6:

[0384] The server analyzes the received data and confirms the detection results.

[0385] Step 7:

[0386] The server sends a notification instruction to the terminal.

[0387] Step 8:

[0388] The device will notify the user with a visual alert and / or vibration.

[0389] Step 9:

[0390] The device will display "A specific sound has been detected. Please be careful."

[0391] Utilizing the Emotion Engine

[0392] Step 1:

[0393] The user puts on the device and activates the emotion engine function.

[0394] Step 2:

[0395] The device collects the user's voice tone and facial expressions using the built-in microphone and camera.

[0396] Step 3:

[0397] The device sends the collected voice tone and facial expression data to the server.

[0398] Step 4:

[0399] The server analyzes the received data using an emotion engine to identify the user's emotional state.

[0400] Step 5:

[0401] The server returns the analysis results to the device and instructs it on the appropriate response.

[0402] Step 6:

[0403] The device will notify the user, for example, "Emotional state detected. Would you like to relax?"

[0404] Step 7:

[0405] The user selects "Yes" or "No."

[0406] Step 8:

[0407] If the user selects "Yes," the terminal plays back audio content that has a relaxing effect.

[0408] Step 9:

[0409] The device again monitors changes in the user's emotional state and makes adjustments as needed.

[0410] Example 2

[0411] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0412] Among the challenges faced by people with hearing impairments and those with hearing loss in their daily lives are adjusting their hearing aid settings to changes in the surrounding sound environment, understanding conversations and important announcements, obtaining information in public facilities, and recognizing emergency situations in real time. Effectively resolving these challenges requires a system that can collect and analyze various voice and environmental data in real time and provide appropriate support. Previously, no system offered all of these functions comprehensively, forcing users to use multiple devices and services at the same time, which was inconvenient.

[0413] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0414] In this invention, the server includes a means for analyzing the sound environment in real time and suggesting optimal hearing aid settings, a means for converting voice data into text, a means for providing visual content with sign language explanations, a means for detecting specific sounds and words and issuing a warning, and a means for analyzing the user's emotional state and suggesting appropriate responses. This allows users to receive a variety of support from a single system, comprehensively resolving various issues in daily life.

[0415] "Analyzing the sound environment in real time and proposing optimal hearing aid settings" means collecting surrounding sounds using a microphone or other device, analyzing the collected sound data, and calculating and providing hearing aid settings that optimize the way the user hears sound.

[0416] "Converting voice data to text" means collecting surrounding voices as digital signals, analyzing the voice data, converting it into text information, and displaying it.

[0417] "Providing visual content with sign language explanations" means providing users with videos or animations with added sign language explanations to supplement visual means of communication.

[0418] "Detecting specific sounds or words and issuing a warning" means recognizing specific sounds or words that have been set in advance, and issuing a notification or warning to the user when they are detected.

[0419] "Analyzing the user's emotional state and suggesting appropriate responses" means analyzing the user's tone of voice, facial expressions, etc. to recognize their emotional state, and then suggesting relaxation content or other appropriate support based on the results.

[0420] The present invention is a comprehensive system for solving various problems faced by the deaf and hard of hearing in daily life. The system includes means for analyzing the sound environment in real time and suggesting optimal hearing aid settings, means for converting speech data into text, means for providing visual content with sign language explanations, means for detecting specific sounds and words and issuing warnings, and means for analyzing the user's emotional state and suggesting appropriate responses.

[0421] Automatic hearing aid optimization

[0422] 1. Acoustic environment data collection

[0423] The user puts on the wearable device and turns it on.

[0424] The device uses a built-in microphone to collect the surrounding sound environment in real time.

[0425] Hardware used: Wearable devices (hearing aids)

[0426] 2. Analysis of sound environment data

[0427] The device analyzes the sound data collected and identifies speech, noise, alarm sounds, etc.

[0428] Software used: Audio analysis algorithm (e.g., FFT algorithm)

[0429] 3. Calculation and recommendation of hearing aid settings

[0430] The device sends the analyzed sound environment data to the server, which then calculates the optimal hearing aid settings.

[0431] The server sends the calculated settings back to the terminal and suggests them to the user.

[0432] Specific examples

[0433] When a user is talking with a friend at a coffee shop, the device recognizes the surrounding noise and the conversation and sends the data to the server. The server calculates the settings to reduce noise and emphasize the conversation and sends them back to the device. The device then displays a message to the user asking, "Would you like to change the settings to reduce noise and emphasize the conversation?" and the settings are applied if the user approves.

[0434] Utilizing voice recognition technology

[0435] 1. Collection of audio data

[0436] The user puts on the device and enables the speech recognition feature.

[0437] Your device uses a built-in microphone to collect ambient sounds.

[0438] 2. Analysis and display of audio data

[0439] The voice data collected by the device is sent to the server in real time, where it is analyzed by the server's voice recognition engine and converted into text.

[0440] The converted text data is returned to the terminal and displayed on the display.

[0441] Specific examples

[0442] When a user is participating in a meeting at work, the device sends the contents of the meeting to a speech recognition engine, which converts the conversation into text and displays "About the next presentation" on the device display.

[0443] Barrier-free information provided

[0444] 1. Collection of public information

[0445] The user uses the device at a train station or public facility.

[0446] The device uses a built-in microphone to collect public announcements and guidance broadcasts.

[0447] 2. Providing related videos

[0448] The terminal sends the collected information to a server, which searches a database of video content with sign language commentary.

[0449] When related video content is found, the server returns the link to the terminal and notifies the user.

[0450] Specific examples

[0451] When the user is at a station, the device collects the station's departure announcements, and the server finds and sends a video with sign language explanations to the device. The user views the video with sign language explanations, which include the message "The next train departs at 10:25," and accurately understands the information.

[0452] Personal Settings Function

[0453] 1. Monitor specific sounds and words in real time

[0454] The user pre-sets specific sounds and words.

[0455] The device constantly monitors surrounding sounds and, when it detects preset sounds or words, sends them to the server.

[0456] 2. Issuing a warning

[0457] After the server confirms, it sends a notification instruction to the terminal.

[0458] The device will alert the user with a visual alert and / or vibration.

[0459] Specific examples

[0460] When a user is relaxing at home and has set a specific word, "danger," and hears that word from outside, the device will vibrate to notify the user and display "Danger detected" on the screen to warn them.

[0461] Utilizing the Emotion Engine

[0462] 1. Monitoring your emotional state

[0463] The device monitors the user's emotional state in real time while they are wearing it.

[0464] The emotion engine analyzes the user's tone of voice and facial expressions to recognize their emotional state.

[0465] 2. Proposal of appropriate action

[0466] The analyzed emotional state is sent to a server, which determines the appropriate response.

[0467] The terminal notifies the user of the suggestion.

[0468] Specific examples

[0469] If a user feels stressed during a meeting, the emotion engine analyzes the user's facial expressions and tone of voice to detect the stress level. The server recommends audio content with a relaxing effect, and the device notifies the user, "Would you like to play relaxing music?" and plays the music if the user agrees.

[0470] Prompt Sentence Examples

[0471] Automatic hearing aid optimization prompts

[0472] Describe a scenario where a user is talking with a friend at a coffee shop and their hearing aids analyze the surrounding sound environment and suggest optimal settings to enhance the conversation and reduce noise.

[0473] Voice recognition technology prompts

[0474] Please give a specific example of a situation where a user is participating in a meeting at work and speech recognition technology is used to convert the meeting content into text and display it on a display.

[0475] As can be seen, the present invention can provide an effective means for improving the quality of life of the deaf and hard of hearing in a variety of situations.

[0476] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0477] Automatic hearing aid optimization

[0478] Step 1:

[0479] The user puts on the wearable device and turns it on.

[0480] The specific operation is that the user physically switches on the device and confirms the connection. As input, the device accepts the power-on state. As output, the device starts up and prepares for the next processing step.

[0481] Step 2:

[0482] The device uses a built-in microphone to collect the surrounding sound environment in real time.

[0483] Specifically, the microphone converts ambient sound into a digital signal and acquires volume and frequency data. The input is the surrounding sound environment, and the output is sound data generated as a digital signal.

[0484] Step 3:

[0485] Analyzes the sound environment data collected by the device and distinguishes between speech, noise, alarm sounds, etc.

[0486] Specifically, it uses a sound analysis algorithm (e.g., FFT algorithm) to extract and classify sound features. As input, it accepts the sound data collected in step 2. As output, it generates data on the identified sound types and their features.

[0487] Step 4:

[0488] The device sends the analysis data to the server.

[0489] Specifically, it transmits data via an internet connection using the HTTPS protocol. As input, it accepts analyzed sound environment data. As output, it sends the data to a server.

[0490] Step 5:

[0491] The server calculates the optimal hearing aid settings

[0492] Specifically, it uses an AI model (machine learning model) to calculate settings, accepts analysis data sent from the device as input, and generates optimal hearing aid setting data as output.

[0493] Step 6:

[0494] The server sends the settings back to the device

[0495] Specifically, the system packages the configuration data and transmits it via HTTPS. As input, it accepts the hearing aid configuration data from the AI ​​model. As output, it sends the configuration data to the device.

[0496] Step 7:

[0497] The device suggests settings to the user

[0498] Specifically, it displays the settings on the display and asks for the user's approval. It accepts the settings data received from the server as input. It displays a dialog box with suggested settings to the user as output.

[0499] Step 8:

[0500] The user approves the settings

[0501] The specific action is to press the approval button on the display, the input is to accept the setting proposal dialog, and the output is to notify the terminal of the approval input.

[0502] Step 9:

[0503] The device adjusts the hearing aid

[0504] Specifically, the system writes the settings data to the hearing aid's internal memory and updates the acoustic profile. As input, it accepts user approval data. As output, the hearing aid settings are changed in real time.

[0505] Utilizing voice recognition technology

[0506] Step 1:

[0507] The user puts on the device and enables the speech recognition feature

[0508] Specifically, the operation is to turn on the device's voice recognition function, to accept a user instruction to enable the voice recognition function as input, and to activate the voice recognition function as output.

[0509] Step 2:

[0510] Your device uses the built-in microphone to collect ambient sounds

[0511] Specifically, the microphone converts audio data into a digital signal and generates a data stream. As input, it collects ambient audio. As output, it generates a digital audio signal.

[0512] Step 3:

[0513] The device sends the collected voice data to the server in real time.

[0514] Specifically, it buffers audio data and periodically sends packets to the server, accepts collected audio data as input, and transmits audio data to the server as output.

[0515] Step 4:

[0516] The server analyzes the voice data and converts it into text

[0517] Specifically, it converts voice data into text using a speech recognition engine (e.g., Google Speech-to-Text), accepts voice data sent from the device as input, and generates analyzed text data as output.

[0518] Step 5:

[0519] The server sends the converted text data to the terminal.

[0520] Specifically, it sends text data via the HTTPS protocol, accepts text data converted into character information as input, and sends the text data to the terminal as output.

[0521] Step 6:

[0522] The device displays the text data.

[0523] Specifically, it renders text data on the display in a format that is easy for the user to view, accepts text data from the server as input, and displays the text data on the display as output.

[0524] Barrier-free information provided

[0525] Step 1:

[0526] Users use their devices at train stations and public facilities

[0527] Specifically, the operation is to enable the device's audio collection function. As input, an instruction to enable the audio collection function by a user operation is accepted. As output, the audio collection function is started.

[0528] Step 2:

[0529] The device uses its built-in microphone to capture public announcements and announcements.

[0530] Specifically, the microphone converts audio data into a digital signal and detects specific frequency bands and patterns. It collects public announcements and announcements as inputs and generates a digital audio signal as output.

[0531] Step 3:

[0532] The device sends the collected information to a server

[0533] Specifically, it compresses the audio data appropriately and transmits it to the server in real time. As input, it accepts collected public announcements and information announcements. As output, it transmits the audio data to the server.

[0534] Step 4:

[0535] The server searches a database of video content with sign language commentary

[0536] Specifically, it queries the database to retrieve the relevant video links, accepts audio data sent from the device as input, and generates the associated video links as output.

[0537] Step 5:

[0538] The server sends the video link back to the device

[0539] Specifically, the video link is packaged and sent in a predetermined format, the relevant video link is accepted as input, and the video link is sent to the terminal as output.

[0540] Step 6:

[0541] The device notifies the user

[0542] Specifically, it displays a pop-up notification on the display to notify the user of the video link. As input, it accepts the video link received from the server. As output, it notifies the user of the video link.

[0543] Step 7:

[0544] The user selects the notification and the video plays on the device.

[0545] Specifically, the video player loads the URL and streams the video with sign language descriptions. As input, it accepts the user's notification selection. As output, it plays the video with sign language descriptions.

[0546] Personal Settings Function

[0547] Step 1:

[0548] User pre-sets specific sounds and words

[0549] Specifically, the user inputs specific sounds or words from the device's settings screen. The input accepts the user's settings for specific sounds or words. The output saves the setting data.

[0550] Step 2:

[0551] The device monitors surrounding sounds in real time

[0552] Specifically, it uses a built-in microphone to collect ambient sounds and then runs an algorithm to identify predefined sounds and words. The input is ambient sounds, and the output is the specific sounds or words detected.

[0553] Step 3:

[0554] When a specific sound or word is detected, it is sent to the server.

[0555] The specific operation is to send the detected sound or word data to the server. The input is to accept the detected data of a specific sound or word. The output is to send the data to the server.

[0556] Step 4:

[0557] After the server confirms, it sends a notification instruction to the device.

[0558] Specifically, it analyzes the received data, generates and sends appropriate notification instructions, accepts detection data from the device as input, and sends notification instructions to the device as output.

[0559] Step 5:

[0560] The device will alert the user with a visual alert and / or vibration.

[0561] As a specific operation, it activates the specified notification method (visual alert or vibration). As input, it accepts notification instructions from the server. As output, it warns the user.

[0562] Utilizing the Emotion Engine

[0563] Step 1:

[0564] Monitors emotional state in real time while the user is wearing the device

[0565] Specifically, it uses the device's built-in camera and microphone to collect the user's voice tone and facial expression data. As input, it collects the user's voice and facial expressions. As output, it generates emotion data.

[0566] Step 2:

[0567] The emotion engine analyzes the user's tone of voice and facial expressions to recognize their emotional state.

[0568] Specifically, it uses an analysis algorithm (e.g., Affectiva's SDK) to determine the emotional state. As input, it accepts collected voice and facial expression data. As output, it generates emotional state data.

[0569] Step 3:

[0570] Send the analyzed emotional state to the server

[0571] Specifically, the data acquired by the emotion engine is sent to the server. As input, emotional state data is accepted. As output, data is sent to the server.

[0572] Step 4:

[0573] The server determines the appropriate response and returns instructions to the device.

[0574] Specifically, the system analyzes the emotional state data and selects relaxing audio content or other appropriate responses. It accepts the emotional state data as input and generates and transmits instruction data to the device as output.

[0575] Step 5:

[0576] The device notifies the user of the suggestion

[0577] Specific actions include displaying suggestions to the user on a screen or by voice message, accepting instruction data from the server as input, and notifying the user as output.

[0578] (Application example 2)

[0579] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0580] Conventional hearing aids have limited means for properly analyzing the surrounding sound environment, and in particular lack systems for recognizing dangerous or suspicious sounds in real time and immediately notifying the user. This makes it difficult for people with hearing impairments or hard of hearing to quickly obtain information that will help them live safely.

[0581] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0582] In this invention, the server includes means for analyzing the sound environment in real time and recommending appropriate audio device settings, means for converting audio data into text, means for providing video content with sign language commentary, means for detecting specific sounds and words and notifying the user, and means for analyzing dangerous sounds and suspicious sounds and notifying the user. This makes it possible to detect dangerous sounds and suspicious sounds, and enables the user to obtain the information necessary for living safely in real time.

[0583] "Sound environment" refers to all the sound states and conditions that exist in the surroundings.

[0584] "Real-time" means that data collection and processing occurs in real time.

[0585] "Analysis" refers to examining the collected data in detail and clarifying its meaning and characteristics.

[0586] "Sound equipment" is a general term for equipment and devices used to collect, amplify, and reproduce sound.

[0587] "Voice data" means any digital recording of a human voice or other sound.

[0588] "Text" is audio or video converted into written information.

[0589] "Video content with sign language commentary" refers to video media that includes commentary in sign language.

[0590] "Specific sounds and words" refers to pre-set important sounds and keywords.

[0591] "Notification" refers to an alert or message that informs the user of specific information.

[0592] "Danger sounds" refer to sounds that may cause some kind of harm to the user.

[0593] "Suspicious audio" refers to audio that may indicate an abnormality or danger.

[0594] "User" refers to individuals, primarily deaf or hard of hearing, who use the system.

[0595] A system for implementing this invention includes an application installed on a device such as a smartphone or a security robot. The system configuration includes the following hardware and software:

[0596] The hardware includes a microphone for collecting sound, a display and vibrator for notifying the user, and in the case of security robots, a movement mechanism. Examples include smartphones and general-purpose robotic devices.

[0597] The software includes a speech analysis engine (e.g., Google Speech-to-Text API) that analyzes the sound environment in real time, an emotion analysis engine (e.g., Amazon Rekognition), and a cloud server (e.g., Amazon Web Services or Google Cloud Platform). Additionally, it runs a notification system (e.g., Firebase Cloud Messaging) to send notifications to users.

[0598] First, when a user turns on the device, the microphone collects the surrounding sound environment. The collected audio data is sent to a cloud server and analyzed using a voice analysis engine. During the analysis process, the audio data is converted into text and identified to determine whether it contains specific dangerous or suspicious sounds.

[0599] Additionally, in some cases, an emotion analysis engine is used to analyze facial expressions and vocal tones of people around you, collecting additional information if an abnormal emotional state is detected, and this data is also sent to a cloud server.

[0600] The analysis results are sent back to the user's device from the cloud server, whereupon they are instantly notified, if applicable, in the form of a visual alert on the display or a physical alert via a vibrator.

[0601] For example, if a security robot is used at home and detects any suspicious sounds at night, it will immediately send a notification to the user's smartphone, and the robot will automatically start recording and, if necessary, call the police.

[0602] An example of a prompt for the generative AI model is as follows:

[0603] "Please analyze the audio data below and identify any dangerous sounds or suspicious voices. Also, perform emotion analysis, and if any facial expressions of anger or fear are recognized, please provide that information as well."

[0604] This allows users to immediately sense danger and take necessary measures. This system provides an effective means to improve the quality of life for the deaf and hard of hearing.

[0605] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0606] Step 1:

[0607] Audio data collection

[0608] Input: Sound environment collected by a microphone installed on the device

[0609] Processing: The device collects surrounding sounds in real time and records them as audio data.

[0610] Output: The collected audio data is saved on the device.

[0611] Step 2:

[0612] Sending audio data

[0613] Input: Audio data collected in step 1

[0614] Processing: The device sends the voice data to the cloud server.

[0615] Output: The audio data is sent to the cloud server.

[0616] Step 3:

[0617] Audio analysis

[0618] Input: Audio data sent to the server

[0619] Processing: The server uses a speech analysis engine (such as the Google Speech-to-Text API) to convert the audio data into text, and then uses a generative AI model to analyze it to identify specific dangerous or suspicious sounds.

[0620] Output: Text data and analysis results are generated.

[0621] Step 4:

[0622] Emotion analysis

[0623] Input: Text data from step 3 and image data as needed

[0624] Processing: The server uses an emotion analysis engine (such as Amazon Rekognition) to analyze the emotional state of people around you. Specifically, it analyzes facial expressions and tone of voice to detect emotions such as stress and anger.

[0625] Output: The result of sentiment analysis is obtained.

[0626] Step 5:

[0627] Integration of analysis results

[0628] Input: Results of Step 3 and Step 4

[0629] Processing: The server integrates the results of the speech analysis and emotion analysis and determines the appropriate response for the user.

[0630] Output: As a result of the integration, information is generated that should be communicated to the user.

[0631] Step 6:

[0632] Sending notifications

[0633] Input: Notification information generated in step 5

[0634] Processing: The server uses a notification system (such as Firebase Cloud Messaging) to send a notification to the device.

[0635] Output: A notification will be displayed on the device.

[0636] Step 7:

[0637] Alert the user

[0638] Input: Notification information from step 6

[0639] Action: The device notifies the user of the notification using a visual alert and / or vibration, specifically by displaying a warning message on the device display and vibrating to get the user's attention.

[0640] Output: The user can view the notification and take action.

[0641] This allows the user to immediately detect dangerous or suspicious sounds and take necessary measures promptly.

[0642] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0643] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0644] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0645] [Second embodiment]

[0646] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0647] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0648] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0649] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0650] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0651] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0652] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0653] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0654] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0655] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0656] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0657] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0658] The present invention is a system for solving various problems that hearing-impaired and hard-of-hearing people encounter in their daily lives, and specifically includes a means for analyzing the sound environment in real time and recommending appropriate hearing aid settings, a means for converting audio data into text, a means for providing video content with sign language commentary, and a means for detecting and notifying specific sounds and words. Each function of this system is described in detail below.

[0659] Automatic hearing aid optimization

[0660] When a user puts on the wearable device and turns it on, the device automatically collects the surrounding sound environment using a microphone. The sound environment data is analyzed by the device to identify sound types such as speech, noise, and alarms. The analyzed data is sent to a server, which calculates the most suitable hearing aid settings. The calculated settings are sent back to the device and suggested to the user. If the user approves, the device automatically adjusts the hearing aid settings.

[0661] Utilizing voice recognition technology

[0662] When a user wears the device and activates the voice recognition function, the device continuously collects surrounding sounds. The collected voice data is sent to a server in real time, where a voice recognition engine analyzes the voice data and converts it into text. The converted text data is displayed on the device's display, allowing the user to visually check the content of conversations around them.

[0663] Barrier-free information provided

[0664] When a user uses a device in a station or public facility, the device collects public information such as announcements and guidance announcements. The collected information is sent to a server, which searches a database of video content with sign language commentary. When relevant video content is found, the server returns a link to the device and notifies the user. When the user selects the link, the video with sign language commentary is played on the device.

[0665] Personal Settings Function

[0666] The user can set specific sounds or words (such as their name or "danger") in advance, and the device will monitor the surrounding sounds in real time. When the set sound or word is detected, the corresponding data is sent to the server. After confirmation, the server sends a notification instruction to the device. The device will then alert the user using a visual alert or vibration.

[0667] Specific examples

[0668] 1. Examples of automatic hearing aid optimization:

[0669] When a user is talking with a friend at a coffee shop, the device recognizes the surrounding noise and the conversation and sends the data to the server. The server calculates the settings to reduce noise and emphasize the conversation and sends them back to the device. The device then displays a message to the user asking, "Would you like to change the settings to reduce noise and emphasize the conversation?" and the settings are applied if the user approves.

[0670] 2. Examples of voice recognition technology:

[0671] When a user is participating in a meeting at work, the device sends the contents of the meeting to a speech recognition engine, which converts the conversation into text and displays "About the next presentation" on the device display, allowing the user to easily understand the content.

[0672] 3. Examples of providing barrier-free information:

[0673] When the user is at a station, the device collects the station's departure announcements, and the server finds and sends a video with sign language explanations to the device. The user views the video with sign language explanations, which include the message "The next train departs at 10:25," and accurately understands the information.

[0674] 4. Examples of personalization features:

[0675] When a user is relaxing at home, a specific word, "danger," is set, and if that word is heard from outside, the device will vibrate to notify the user and display "Danger detected" on the screen to warn them.

[0676] In this way, the present invention provides an effective means for improving the quality of life of the deaf and hard of hearing.

[0677] The processing flow will be explained below.

[0678] Automatic hearing aid optimization

[0679] Step 1:

[0680] The device detects that the user is wearing a wearable device and has it turned on.

[0681] Step 2:

[0682] The device collects the surrounding sound environment in real time using its built-in microphone and generates sound environment data.

[0683] Step 3:

[0684] The device analyzes the sound environment data collected and identifies conversation sounds, noise, alarm sounds, etc.

[0685] Step 4:

[0686] The device sends the analysis results to the server.

[0687] Step 5:

[0688] The server analyzes the received sound environment data and calculates the optimal hearing aid settings.

[0689] Step 6:

[0690] The server returns the calculation results to the terminal.

[0691] Step 7:

[0692] The device will prompt the user, "There are new environment-based hearing aid settings. Would you like to apply them?"

[0693] Step 8:

[0694] The user selects "Yes" or "No."

[0695] Step 9:

[0696] If the user selects "Yes," the device will automatically change the hearing aid settings.

[0697] Step 10:

[0698] The device will notify the user that the new settings have been applied.

[0699] Utilizing voice recognition technology

[0700] Step 1:

[0701] The user puts on the device and enables the speech recognition feature.

[0702] Step 2:

[0703] The device continuously collects surrounding sounds using the built-in microphone.

[0704] Step 3:

[0705] The device sends the collected voice data to the server.

[0706] Step 4:

[0707] The server analyzes the received voice data and converts it into text using a voice recognition engine.

[0708] Step 5:

[0709] The server returns the converted text data to the terminal.

[0710] Step 6:

[0711] The device displays the received text on the display.

[0712] Step 7:

[0713] The device updates the text display in real time as new audio data is collected.

[0714] Barrier-free information provided

[0715] Step 1:

[0716] The user wears the device and enables the barrier-free information provision function.

[0717] Step 2:

[0718] The device uses a built-in microphone to collect public announcements and guidance broadcasts.

[0719] Step 3:

[0720] The device sends the collected voice data to the server.

[0721] Step 4:

[0722] The server analyzes the audio data and searches a database for relevant video content with sign language commentary.

[0723] Step 5:

[0724] The server sends a link to the appropriate video content back to the device.

[0725] Step 6:

[0726] The device notifies the user, "There is a related sign language explanation video. Would you like to play it?"

[0727] Step 7:

[0728] The user selects "Yes" or "No."

[0729] Step 8:

[0730] If the user selects "Yes," the device will open the video link and play the sign language explanation video.

[0731] Personal Settings Function

[0732] Step 1:

[0733] The user puts on the device and accesses the personalization screen.

[0734] Step 2:

[0735] The user sets and saves a specific sound or word (such as a name or "danger").

[0736] Step 3:

[0737] The device collects surrounding sounds in real time using a built-in microphone.

[0738] Step 4:

[0739] The device continuously compares the sounds it collects with preset sounds and words.

[0740] Step 5:

[0741] When the device detects a specific sound or word, it sends that data to the server.

[0742] Step 6:

[0743] The server analyzes the received data and confirms the detection results.

[0744] Step 7:

[0745] The server sends a notification instruction to the terminal.

[0746] Step 8:

[0747] The device will notify the user with a visual alert and / or vibration.

[0748] Step 9:

[0749] The device will display "A specific sound has been detected. Please be careful."

[0750] Example 1

[0751] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0752] This invention aims to solve various problems that the hearing impaired and hard of hearing face in their daily lives. Specifically, it aims to improve the quality of life for the hearing impaired by providing technology that efficiently analyzes surrounding sounds and automatically configures optimal hearing aids, as well as technology that converts speech to text, provides video content that incorporates sign language, and detects and notifies users of specific sounds and words.

[0753] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0754] In this invention, the server includes a means for analyzing the sound environment in real time and recommending appropriate hearing aid settings, a means for converting recorded audio data into text, and a means for providing video content with sign language commentary, thereby enabling hearing-impaired or hard-of-hearing people to efficiently grasp the surrounding sound environment and receive appropriate hearing assistance.

[0755] "Sound environment" refers to the state and situation of sound waves such as surrounding voices and noise.

[0756] "Analyzing in real time" means analyzing data immediately after it is collected, without any time delay.

[0757] A "hearing assistive device" is a device that assists the hearing of people with hearing loss or hearing impairments, and has the function of amplifying sound or emphasizing specific sounds.

[0758] "Audio recording data" refers to data that has been digitized and saved from audio collected by a microphone or other device.

[0759] "Converting to text" refers to the process of converting data such as audio and video into text information.

[0760] "Video content with sign language commentary" refers to videos or video materials that include visual sign language commentary.

[0761] "Detecting specific sounds or words" refers to the recognition system identifying and detecting predefined sounds or words.

[0762] "Notify" refers to informing a user of specific information or events.

[0763] "Device" refers to electronic equipment or devices, including hearing aids and wearable devices.

[0764] "Central server" refers to the main computer system that receives data from multiple devices, processes and analyzes it, and returns the results.

[0765] "Collect" refers to gathering data or information.

[0766] "Transmit" refers to sending data or information to another device or system.

[0767] "Calculating" refers to the process of deriving a specific result based on analytical results or algorithms.

[0768] "Automatically adjust" means that the device is automatically set to the optimum state without requiring any user operation.

[0769] "Public facilities" refers to buildings and places for use by the general public, including stations and libraries.

[0770] "Linking" means associating and making accessible particular information or content.

[0771] "Monitor" refers to watching for a particular condition or state.

[0772] This system analyzes the sound environment, converts speech into text, provides video content with sign language commentary, and detects and notifies specific sounds and words to solve various problems faced by the hearing impaired and hard of hearing in their daily lives. The system of this invention functions mainly through cooperation between the device (wearable terminal) and a central server.

[0773] Hardware and software used

[0774] 1. Device (wearable device)

[0775] microphone

[0776] Built-in processor

[0777] Communication modules (Wi-Fi, 5G, etc.)

[0778] display

[0779] Vibration Motor

[0780] battery

[0781] 2. Central Server

[0782] Generative AI Models

[0783] Speech Recognition Engine

[0784] Database (including video content with sign language commentary)

[0785] Communication Interface

[0786] Program processing

[0787] Analysis of sound environments and optimization of hearing aids

[0788] When a user puts on the device and turns it on, the device automatically collects the surrounding sound environment using a microphone. The sound environment data is analyzed by the device to identify sound types such as speech, noise, and alarms. This data is sent to a central server, where a generative AI model calculates optimal hearing aid settings. The calculation results are sent back to the device and recommended to the user. If the user approves, the device automatically adjusts the hearing aid settings.

[0789] Examples:

[0790] When a user is talking with a friend at a coffee shop, the device recognizes the surrounding noise and the conversation and sends the data to a central server. The server calculates the settings to reduce noise and enhance the conversation and sends them back to the device. The device then displays a message asking, "Do you want to change the settings to reduce noise and enhance the conversation?" and the settings are applied if the user confirms.

[0791] Utilizing voice recognition technology

[0792] When a user enables the device's voice recognition function, the device will continuously collect surrounding sounds. The collected voice data is sent to a central server in real time, where the voice recognition engine converts the speech into text. The converted text data is then displayed on the device's display, allowing the user to visually check the content of the conversations around them.

[0793] Examples:

[0794] When a user is participating in a meeting at work, the device sends the contents of the meeting to a speech recognition engine, and the converted text is displayed on the device's display as "About the next presentation," allowing the user to easily understand the content.

[0795] Providing video content with sign language commentary

[0796] When a user uses a device in a station or public facility, the device collects public information such as announcements and guidance announcements. The collected information is sent to a central server, which searches a database of videos with sign language descriptions. When relevant video content is found, the central server sends a link back to the device and notifies the user. When the user selects the link, the video with sign language descriptions is played on the device.

[0797] Examples:

[0798] When a user is at a station, the terminal collects the station's departure announcements, and the central server finds the video with sign language explanations and sends it to the terminal. The user views the video with sign language explanations, which include the message "The next train departs at 10:25," and accurately understands the information.

[0799] Detecting and announcing specific sounds and words

[0800] Users can pre-set specific sounds or words (e.g., "danger"), and the device will monitor surrounding sounds in real time. When the device detects a specific sound or word, the corresponding data is sent to a central server, which then confirms the detection and sends a notification instruction to the device. The device will then alert the user with a visual alert or vibration.

[0801] Examples:

[0802] If the user is relaxing at home and the word "danger" is set and hears that word from outside, the device will vibrate to notify the user and display "Danger detected" on the screen to warn them.

[0803] As described above, the present invention integrates multiple technical means to provide an effective method for the hearing impaired and hard of hearing to efficiently obtain information necessary in daily life.

[0804] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0805] Analysis of sound environments and optimization of hearing aids

[0806] Step 1:

[0807] When a user wears the device and turns it on, the device automatically collects the surrounding sound environment with its microphone. The input here is the surrounding sound environment, and the output is digitized sound environment data. The device's microphone converts analog audio into digital data.

[0808] Step 2:

[0809] The device analyzes the sound environment data collected in real time to identify the type of sound, such as speech, noise, or alarm sounds. The input is digitized sound environment data, and the output is information about the identified sound type (voice pattern data). The device's built-in processor runs an algorithm to analyze the data.

[0810] Step 3:

[0811] The device sends the analyzed voice pattern data to a central server. The input is the voice pattern data, and the output is the transmitted data. The data is sent to a cloud server using the device's communication module (Wi-Fi, 5G, etc.).

[0812] Step 4:

[0813] Based on the voice pattern data received by the server, a generative AI model calculates the optimal hearing aid settings. The input is voice pattern data, and the output is optimal setting data. The generative AI model analyzes the voice data and derives the appropriate parameters.

[0814] Step 5:

[0815] The server returns the optimal setting data to the terminal. The input is the optimal setting data, and the output is the setting data sent to the terminal. The data is sent using the server's communication interface.

[0816] Step 6:

[0817] The device receives the optimal setting data and displays the suggestion to the user. The input is the received setting data, and the output is the display of the suggestion. The device display shows "Would you like to change the settings to reduce noise and enhance speech?"

[0818] Step 7:

[0819] When the user presses the approval button, the device automatically adjusts the settings of the hearing aid. The input is the user's approval, and the output is the adjusted settings of the hearing aid. The device provides feedback to the user by vibration or sound.

[0820] Utilizing voice recognition technology

[0821] Step 1:

[0822] When a user enables the voice recognition function of the device, the device continuously collects surrounding sounds. The input is the surrounding sounds, and the output is digitized voice data. The microphone collects the voice and converts it into digital data.

[0823] Step 2:

[0824] The terminal transmits the collected voice data to a central server in real time. The input is the digitized voice data, and the output is the transmitted voice data. The data is transmitted using a communication module.

[0825] Step 3:

[0826] The server analyzes the voice data using a voice recognition engine and converts it into text data. The input is voice data and the output is text data. The voice recognition engine analyzes the voice pattern and generates the corresponding text.

[0827] Step 4:

[0828] Text data is sent from the central server to the terminal. The input is the text data, and the output is the text data sent to the terminal. The data is sent using the server's communication interface.

[0829] Step 5:

[0830] Text data is displayed on the terminal display. The input is the received text data, and the output is the text content displayed on the display. The user can check the content of the surrounding conversation by looking at the displayed text.

[0831] Providing video content with sign language commentary

[0832] Step 1:

[0833] When a user uses a device in a station or public facility, the terminal collects announcements and guidance announcements. The input is the public facility announcements, and the output is digitized voice data. The microphone collects the voice and converts it into digital data.

[0834] Step 2:

[0835] The collected voice data is transmitted to a central server. The input is the digitized voice data and the output is the transmitted voice data. A communication module is used to transmit the data to the server.

[0836] Step 3:

[0837] The server analyzes the audio data and searches a database of video content with sign language commentary. The input is the audio data, and the output is a link to related video content. The server's generative AI model analyzes the audio and finds related videos from the database.

[0838] Step 4:

[0839] The server returns a link to the associated video content to the terminal. The input is the link to the video content, and the output is the link sent to the terminal. The link is sent using a communication interface.

[0840] Step 5:

[0841] When the user selects a link, the device plays a video with sign language descriptions. The input is the link selected by the user and the output is the video that is played. The video with sign language descriptions appears on the device's display.

[0842] Detecting and announcing specific sounds and words

[0843] Step 1:

[0844] The user sets specific sounds or words (e.g., "danger") in advance on the device. The input is the user's settings for the sounds or words, and the output is the saved setting data. The sounds or words are entered using the device's setting screen and saved in a database.

[0845] Step 2:

[0846] The device monitors the surrounding sound in real time. The input is the surrounding sound, and the output is digitized audio data. The microphone collects the audio and converts it into digital data.

[0847] Step 3:

[0848] When a preset sound or word is detected, the device sends data to a central server. The input is the detected voice data, and the output is the transmitted data. The data is sent to the server using a communication module.

[0849] Step 4:

[0850] The server checks the received data and sends a notification instruction to the device. The input is the received voice data, and the output is the notification instruction. The server's generative AI model analyzes the data and determines the appropriate response.

[0851] Step 5:

[0852] The device alerts the user with a visual alert or vibration. The input is a notification instruction, and the output is a visual alert or vibration notification. The device activates the vibration motor and displays "Danger detected" on the display.

[0853] (Application example 1)

[0854] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0855] The present invention aims to solve the problem that hearing-impaired and hard-of-hearing people have difficulty recognizing voices and grasping important information in daily life and in certain situations (e.g., when using self-driving vehicles). In particular, there is a problem that emergency situations and important announcements are difficult to convey to hearing-impaired and hard-of-hearing people. This may result in situations where the safety and comfort of users are lacking.

[0856] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0857] In this invention, the server includes a means for analyzing the sound environment in real time and recommending appropriate hearing aid settings, a means for converting voice data into text, and a means for providing video content with sign language commentary. This allows hearing-impaired and hard-of-hearing individuals to properly understand the surrounding sound environment and automatically optimize their hearing aid settings. Furthermore, by providing a means for collecting voice data in real time and notifying the user based on specific keywords, it becomes possible to reliably communicate emergencies and important announcements. Furthermore, by including a means for displaying voice recognition results on an in-vehicle display, users can visually confirm surrounding voice information, thereby improving safety and comfort when using autonomous vehicles.

[0858] "Analyzing the sound environment in real time and recommending appropriate hearing aid settings" means continuously monitoring the sound environment using sensors and microphones, classifying the type and level of sound using an analysis device, and providing optimal hearing aid settings based on that.

[0859] "Converting voice data to text" refers to the process of analyzing collected voice data using speech recognition technology and converting it into a corresponding text format.

[0860] "Providing video content with sign language commentary" is a service that allows deaf and hard of hearing people to obtain information visually by providing video information with sign language interpretation.

[0861] "Detecting and notifying specific sounds and words" refers to a mechanism that detects specific pre-set sounds or utterances and notifies the user.

[0862] "Collecting voice data in real time and notifying the user based on specific keywords" means accumulating voice data in real time, identifying specific keywords in the data, and providing the user with appropriate alerts.

[0863] "Displaying voice recognition results on an in-vehicle display" means displaying text data obtained through voice recognition technology on a display inside an autonomous vehicle to provide information visually to the user.

[0864] The present invention is a system for solving various problems that hearing-impaired and hard-of-hearing people encounter in their daily lives and when using self-driving vehicles. This system includes a means for analyzing the sound environment in real time and recommending appropriate hearing aid settings, a means for converting voice data into text, a means for providing video content with sign language commentary, and a means for detecting and notifying specific sounds and words.

[0865] System configuration

[0866] The system consists of the following main components:

[0867] Hardware

[0868] 1. Microphone: A device that collects ambient sounds, allowing for real-time monitoring of the sound environment.

[0869] 2. Server: Analyzes the sound data and provides instructions for optimal settings for the hearing aid and text conversion.

[0870] 3. Autonomous vehicle display: A device for displaying voice recognition results and system notifications.

[0871] 4. Vibration motor: To notify the user when a specific sound or keyword is detected.

[0872] software

[0873] 1. Speech recognition engine (e.g., Google Speech Recognition API): Collects speech in real time and converts it into text.

[0874] 2. Analysis algorithm: Analyzes the sound environment and recommends appropriate hearing aid settings.

[0875] 3. Notification system: Detects specific keywords or voice anomalies and notifies the user.

[0876] System Operation

[0877] 1. Real-time analysis of sound environments

[0878] When the user operates the system, the microphone begins to collect the surrounding sound environment. The collected sound environment data is sent to the server, where an analysis algorithm identifies the type of sound, such as speech, noise, or alarm. Based on the analyzed data, the server calculates the optimal hearing aid settings and sends them back to the device. If the user approves, the device automatically adjusts the hearing aid settings.

[0879] 2. Use of voice recognition technology

[0880] When the user wears the device and activates the voice recognition function, the device continuously collects surrounding sounds. The collected voice data is sent to a server in real time, where a voice recognition engine analyzes the voice data and converts it into text. The converted text data is displayed on the vehicle's display, allowing the user to visually check the content of surrounding conversations and announcements.

[0881] 3. Providing barrier-free information

[0882] When a user is using an autonomous vehicle, the device collects information such as in-car announcements and emergency alerts. The collected information is sent to a server, which searches a database of video content with sign language commentary. When relevant video content is found, the server returns a link to the device and notifies the user. When the user selects the link, the video with sign language commentary is played on the in-car display.

[0883] 4. Personal settings function

[0884] The user can set specific sounds or words (such as their name or "danger") in advance, and the device will monitor the surrounding sounds in real time. When the set sound or word is detected, the corresponding data is sent to the server. After confirmation, the server sends a notification instruction to the device. The device will then alert the user using a visual alert or vibration.

[0885] Specific examples

[0886] 1. A concrete example of automatic hearing aid optimization

[0887] When a user is talking with a friend at a coffee shop, the device recognizes the surrounding noise and the conversation and sends the data to the server. The server calculates the settings to reduce noise and emphasize the conversation and sends them back to the device. The device then displays a message to the user asking, "Would you like to change the settings to reduce noise and emphasize the conversation?" and the settings are applied if the user approves.

[0888] 2. Specific examples of voice recognition technology

[0889] When a user is participating in a meeting at work, the device sends the contents of the meeting to a speech recognition engine, which converts the conversation into text and displays "About the next presentation" on the device display, allowing the user to easily understand the content.

[0890] 3. Examples of providing barrier-free information

[0891] When the user is at a station, the device collects the station's departure announcements, and the server finds and sends a video with sign language explanations to the device. The user views the video with sign language explanations, which include the message "The next train departs at 10:25," and accurately understands the information.

[0892] 4. Examples of personal settings functions

[0893] When a user is relaxing at home, a specific word, "danger," is set, and if that word is heard from outside, the device will vibrate to notify the user and display "Danger detected" on the screen to warn them.

[0894] Example prompts for generative AI models

[0895] "You are designing a self-driving car system for the hearing impaired. The system uses speech recognition to convert emergency situations and announcements into text and notify the user. Extract the emergency content from the following audio data and display it as text."

[0896] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0897] Step 1:

[0898] The user operates the system. The user gets into the autonomous vehicle and a microphone connected to the terminal collects surrounding sounds in real time. The collected sound data is input into the terminal.

[0899] Step 2:

[0900] The device analyzes the collected sound data and classifies the sound environment. The device's internal analysis algorithm classifies the sound data into categories such as speech, noise, and alarm sounds. This analyzed data is sent to the server.

[0901] Step 3:

[0902] The server receives the sound environment data and calculates the optimal hearing aid settings. Based on the analyzed sound data, calculations are performed to automatically optimize the hearing aid settings. These optimized settings are then sent from the server to the device.

[0903] Step 4:

[0904] The device presents the hearing aid settings to the user. The device asks the user whether to apply the new settings and obtains their approval. When the user presses the approval button, the device changes the hearing aid settings.

[0905] Step 5:

[0906] The device continuously collects voice data and sends it to a voice recognition engine. Based on the voice data input, real-time voice recognition processing is performed and the data is converted into text. This text data is then sent back to the device.

[0907] Step 6:

[0908] The device displays the voice recognition results on the vehicle display, and text data is output to the display so that the user can visually check the content of surrounding conversations and announcements.

[0909] Step 7:

[0910] The device monitors specific keywords and sounds to detect abnormalities. When a pre-defined keyword (e.g., "danger" or "emergency") is detected, the data is sent to the server. After the server confirms the detection, it sends a notification instruction to the device.

[0911] Step 8:

[0912] The device will alert the user. If an abnormal sound or specific keyword is detected, the device will alert the user with a visual alert or vibration notification.

[0913] The above process will create a system that improves the safety and convenience of people who are deaf or hard of hearing when using self-driving vehicles.

[0914] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0915] The present invention is a system for solving various problems encountered in daily life by people with hearing impairments or hard of hearing. Specifically, it includes a means for analyzing the sound environment in real time and recommending appropriate hearing aid settings, a means for converting audio data into text, a means for providing video content with sign language commentary, and a means for detecting and notifying specific sounds or words. Furthermore, by combining it with an emotion engine that recognizes and analyzes the user's emotions, further user assistance is provided. Each function of this system is described in detail below.

[0916] Automatic hearing aid optimization

[0917] When a user puts on the wearable device and turns it on, the device automatically collects the surrounding sound environment using a microphone. The sound environment data is analyzed by the device to identify sound types such as speech, noise, and alarms. The analyzed data is sent to a server, which calculates the most suitable hearing aid settings. The calculated settings are sent back to the device and suggested to the user. If the user approves, the device automatically adjusts the hearing aid settings.

[0918] Utilizing voice recognition technology

[0919] When a user wears the device and activates the voice recognition function, the device continuously collects surrounding sounds. The collected voice data is sent to a server in real time, where a voice recognition engine analyzes the voice data and converts it into text. The converted text data is displayed on the device's display, allowing the user to visually check the content of conversations around them.

[0920] Barrier-free information provided

[0921] When a user uses a device in a station or public facility, the device collects public information such as announcements and guidance announcements. The collected information is sent to a server, which searches a database of video content with sign language commentary. When relevant video content is found, the server returns a link to the device and notifies the user. When the user selects the link, the video with sign language commentary is played on the device.

[0922] Personal Settings Function

[0923] The user can set specific sounds or words (such as their name or "danger") in advance, and the device will monitor the surrounding sounds in real time. When the set sound or word is detected, the corresponding data is sent to the server. After confirmation, the server sends a notification instruction to the device. The device will then alert the user using a visual alert or vibration.

[0924] Utilizing the Emotion Engine

[0925] The emotion engine monitors the user's emotional state in real time while the user is wearing the device. The emotion engine analyzes the user's tone of voice and facial expressions (e.g., using the camera function) to recognize emotions such as stress, joy, and sadness. Once the emotional state is analyzed, the results are sent to the server, which determines the appropriate response and sends instructions back to the device.

[0926] Specific examples

[0927] 1. Examples of automatic hearing aid optimization:

[0928] When a user is talking with a friend at a coffee shop, the device recognizes the surrounding noise and the conversation and sends the data to the server. The server calculates the settings to reduce noise and emphasize the conversation and sends them back to the device. The device then displays a message to the user asking, "Would you like to change the settings to reduce noise and emphasize the conversation?" and the settings are applied if the user approves.

[0929] 2. Examples of voice recognition technology:

[0930] When a user is participating in a meeting at work, the device sends the contents of the meeting to a speech recognition engine, which converts the conversation into text and displays "About the next presentation" on the device display, allowing the user to easily understand the content.

[0931] 3. Examples of providing barrier-free information:

[0932] When the user is at a station, the device collects the station's departure announcements, and the server finds and sends a video with sign language explanations to the device. The user views the video with sign language explanations, which include the message "The next train departs at 10:25," and accurately understands the information.

[0933] 4. Examples of personalization features:

[0934] When a user is relaxing at home, a specific word, "danger," is set, and if that word is heard from outside, the device will vibrate to notify the user and display "Danger detected" on the screen to warn them.

[0935] 5. Examples of Emotion Engines:

[0936] If a user feels stressed during a meeting, the emotion engine analyzes the user's facial expressions and tone of voice to detect the stress level. The server recommends audio content with a relaxing effect, and the device notifies the user, "Would you like to play relaxing music?" and plays the music if the user agrees.

[0937] In this way, the present invention provides an effective means for improving the quality of life of the deaf and hard of hearing.

[0938] The processing flow will be explained below.

[0939] Automatic hearing aid optimization

[0940] Step 1:

[0941] The user puts on the wearable device and turns it on.

[0942] Step 2:

[0943] The device automatically collects the surrounding sound environment in real time using the built-in microphone.

[0944] Step 3:

[0945] The device analyzes the sound environment data collected and identifies the type of sound (conversation, noise, alarm, etc.).

[0946] Step 4:

[0947] The device sends the analysis results to the server.

[0948] Step 5:

[0949] The server analyzes the received sound environment data and calculates appropriate hearing aid settings.

[0950] Step 6:

[0951] The server returns the calculated hearing aid settings to the device.

[0952] Step 7:

[0953] The device will prompt the user, "There are new environment-based hearing aid settings. Would you like to apply them?"

[0954] Step 8:

[0955] The user selects "Yes" or "No."

[0956] Step 9:

[0957] If the user selects "Yes," the device will automatically change the hearing aid settings.

[0958] Step 10:

[0959] The device will notify the user that the new settings have been applied.

[0960] Utilizing voice recognition technology

[0961] Step 1:

[0962] The user puts on the device and enables the speech recognition feature.

[0963] Step 2:

[0964] The device continuously collects surrounding sounds using the built-in microphone.

[0965] Step 3:

[0966] The device sends the collected voice data to the server.

[0967] Step 4:

[0968] The voice data received by the server is analyzed using a voice recognition engine and converted into text.

[0969] Step 5:

[0970] The server returns the converted text data to the terminal.

[0971] Step 6:

[0972] The text data received by the terminal is displayed on the display.

[0973] Step 7:

[0974] The device continues to convert new voice data collected into text in real time, updating the text on the display.

[0975] Barrier-free information provided

[0976] Step 1:

[0977] The user wears the device and enables the barrier-free information provision function.

[0978] Step 2:

[0979] The device uses a built-in microphone to collect public announcements and guidance broadcasts.

[0980] Step 3:

[0981] The device sends the collected voice data to the server.

[0982] Step 4:

[0983] The server analyzes the audio data and searches a database for relevant video content with sign language commentary.

[0984] Step 5:

[0985] The server sends a link to the appropriate video content back to the device.

[0986] Step 6:

[0987] The device notifies the user, "There is a related sign language explanation video. Would you like to play it?"

[0988] Step 7:

[0989] The user selects "Yes" or "No."

[0990] Step 8:

[0991] If the user selects "Yes," the device will open the video link and play the sign language explanation video.

[0992] Personal Settings Function

[0993] Step 1:

[0994] The user puts on the device and accesses the personalization screen.

[0995] Step 2:

[0996] The user sets and saves a specific sound or word (such as a name or "danger").

[0997] Step 3:

[0998] The device collects surrounding sounds in real time using a built-in microphone.

[0999] Step 4:

[1000] The device continuously compares the sounds it collects with preset sounds and words.

[1001] Step 5:

[1002] When the device detects a specific sound or word, it sends that data to the server.

[1003] Step 6:

[1004] The server analyzes the received data and confirms the detection results.

[1005] Step 7:

[1006] The server sends a notification instruction to the terminal.

[1007] Step 8:

[1008] The device will notify the user with a visual alert and / or vibration.

[1009] Step 9:

[1010] The device will display "A specific sound has been detected. Please be careful."

[1011] Utilizing the Emotion Engine

[1012] Step 1:

[1013] The user puts on the device and activates the emotion engine function.

[1014] Step 2:

[1015] The device collects the user's voice tone and facial expressions using the built-in microphone and camera.

[1016] Step 3:

[1017] The device sends the collected voice tone and facial expression data to the server.

[1018] Step 4:

[1019] The server analyzes the received data using an emotion engine to identify the user's emotional state.

[1020] Step 5:

[1021] The server returns the analysis results to the device and instructs it on the appropriate response.

[1022] Step 6:

[1023] The device will notify the user, for example, "Emotional state detected. Would you like to relax?"

[1024] Step 7:

[1025] The user selects "Yes" or "No."

[1026] Step 8:

[1027] If the user selects "Yes," the terminal plays back audio content that has a relaxing effect.

[1028] Step 9:

[1029] The device again monitors changes in the user's emotional state and makes adjustments as needed.

[1030] Example 2

[1031] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1032] Among the challenges faced by people with hearing impairments and those with hearing loss in their daily lives are adjusting their hearing aid settings to changes in the surrounding sound environment, understanding conversations and important announcements, obtaining information in public facilities, and recognizing emergency situations in real time. Effectively resolving these challenges requires a system that can collect and analyze various voice and environmental data in real time and provide appropriate support. Previously, no system offered all of these functions comprehensively, forcing users to use multiple devices and services at the same time, which was inconvenient.

[1033] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1034] In this invention, the server includes a means for analyzing the sound environment in real time and suggesting optimal hearing aid settings, a means for converting voice data into text, a means for providing visual content with sign language explanations, a means for detecting specific sounds and words and issuing a warning, and a means for analyzing the user's emotional state and suggesting appropriate responses. This allows users to receive a variety of support from a single system, comprehensively resolving various issues in daily life.

[1035] "Analyzing the sound environment in real time and proposing optimal hearing aid settings" means collecting surrounding sounds using a microphone or other device, analyzing the collected sound data, and calculating and providing hearing aid settings that optimize the way the user hears sound.

[1036] "Converting voice data to text" means collecting surrounding voices as digital signals, analyzing the voice data, converting it into text information, and displaying it.

[1037] "Providing visual content with sign language explanations" means providing users with videos or animations with added sign language explanations to supplement visual means of communication.

[1038] "Detecting specific sounds or words and issuing a warning" means recognizing specific sounds or words that have been set in advance, and issuing a notification or warning to the user when they are detected.

[1039] "Analyzing the user's emotional state and suggesting appropriate responses" means analyzing the user's tone of voice, facial expressions, etc. to recognize their emotional state, and then suggesting relaxation content or other appropriate support based on the results.

[1040] The present invention is a comprehensive system for solving various problems faced by the deaf and hard of hearing in daily life. The system includes means for analyzing the sound environment in real time and suggesting optimal hearing aid settings, means for converting speech data into text, means for providing visual content with sign language explanations, means for detecting specific sounds and words and issuing warnings, and means for analyzing the user's emotional state and suggesting appropriate responses.

[1041] Automatic hearing aid optimization

[1042] 1. Acoustic environment data collection

[1043] The user puts on the wearable device and turns it on.

[1044] The device uses a built-in microphone to collect the surrounding sound environment in real time.

[1045] Hardware used: Wearable devices (hearing aids)

[1046] 2. Analysis of sound environment data

[1047] The device analyzes the sound data collected and identifies speech, noise, alarm sounds, etc.

[1048] Software used: Audio analysis algorithm (e.g., FFT algorithm)

[1049] 3. Calculation and recommendation of hearing aid settings

[1050] The device sends the analyzed sound environment data to the server, which then calculates the optimal hearing aid settings.

[1051] The server sends the calculated settings back to the terminal and suggests them to the user.

[1052] Specific examples

[1053] When a user is talking with a friend at a coffee shop, the device recognizes the surrounding noise and the conversation and sends the data to the server. The server calculates the settings to reduce noise and emphasize the conversation and sends them back to the device. The device then displays a message to the user asking, "Would you like to change the settings to reduce noise and emphasize the conversation?" and the settings are applied if the user approves.

[1054] Utilizing voice recognition technology

[1055] 1. Collection of audio data

[1056] The user puts on the device and enables the speech recognition feature.

[1057] Your device uses a built-in microphone to collect ambient sounds.

[1058] 2. Analysis and display of audio data

[1059] The voice data collected by the device is sent to the server in real time, where it is analyzed by the server's voice recognition engine and converted into text.

[1060] The converted text data is returned to the terminal and displayed on the display.

[1061] Specific examples

[1062] When a user is participating in a meeting at work, the device sends the contents of the meeting to a speech recognition engine, which converts the conversation into text and displays "About the next presentation" on the device display.

[1063] Barrier-free information provided

[1064] 1. Collection of public information

[1065] The user uses the device at a train station or public facility.

[1066] The device uses a built-in microphone to collect public announcements and guidance broadcasts.

[1067] 2. Providing related videos

[1068] The terminal sends the collected information to a server, which searches a database of video content with sign language commentary.

[1069] When related video content is found, the server returns the link to the terminal and notifies the user.

[1070] Specific examples

[1071] When the user is at a station, the device collects the station's departure announcements, and the server finds and sends a video with sign language explanations to the device. The user views the video with sign language explanations, which include the message "The next train departs at 10:25," and accurately understands the information.

[1072] Personal Settings Function

[1073] 1. Monitor specific sounds and words in real time

[1074] The user pre-sets specific sounds and words.

[1075] The device constantly monitors surrounding sounds and, when it detects preset sounds or words, sends them to the server.

[1076] 2. Issuing a warning

[1077] After the server confirms, it sends a notification instruction to the terminal.

[1078] The device will alert the user with a visual alert and / or vibration.

[1079] Specific examples

[1080] When a user is relaxing at home and has set a specific word, "danger," and hears that word from outside, the device will vibrate to notify the user and display "Danger detected" on the screen to warn them.

[1081] Utilizing the Emotion Engine

[1082] 1. Monitoring your emotional state

[1083] The device monitors the user's emotional state in real time while they are wearing it.

[1084] The emotion engine analyzes the user's tone of voice and facial expressions to recognize their emotional state.

[1085] 2. Proposal of appropriate action

[1086] The analyzed emotional state is sent to a server, which determines the appropriate response.

[1087] The terminal notifies the user of the suggestion.

[1088] Specific examples

[1089] If a user feels stressed during a meeting, the emotion engine analyzes the user's facial expressions and tone of voice to detect the stress level. The server recommends audio content with a relaxing effect, and the device notifies the user, "Would you like to play relaxing music?" and plays the music if the user agrees.

[1090] Prompt Sentence Examples

[1091] Automatic hearing aid optimization prompts

[1092] Describe a scenario where a user is talking with a friend at a coffee shop and their hearing aids analyze the surrounding sound environment and suggest optimal settings to enhance the conversation and reduce noise.

[1093] Voice recognition technology prompts

[1094] Please give a specific example of a situation where a user is participating in a meeting at work and speech recognition technology is used to convert the meeting content into text and display it on a display.

[1095] As can be seen, the present invention can provide an effective means for improving the quality of life of the deaf and hard of hearing in a variety of situations.

[1096] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1097] Automatic hearing aid optimization

[1098] Step 1:

[1099] The user puts on the wearable device and turns it on.

[1100] The specific operation is that the user physically switches on the device and confirms the connection. As input, the device accepts the power-on state. As output, the device starts up and prepares for the next processing step.

[1101] Step 2:

[1102] The device uses a built-in microphone to collect the surrounding sound environment in real time.

[1103] Specifically, the microphone converts ambient sound into a digital signal and acquires volume and frequency data. The input is the surrounding sound environment, and the output is sound data generated as a digital signal.

[1104] Step 3:

[1105] Analyzes the sound environment data collected by the device and distinguishes between speech, noise, alarm sounds, etc.

[1106] Specifically, it uses a sound analysis algorithm (e.g., FFT algorithm) to extract and classify sound features. As input, it accepts the sound data collected in step 2. As output, it generates data on the identified sound types and their features.

[1107] Step 4:

[1108] The device sends the analysis data to the server.

[1109] Specifically, it transmits data via an internet connection using the HTTPS protocol. As input, it accepts analyzed sound environment data. As output, it sends the data to a server.

[1110] Step 5:

[1111] The server calculates the optimal hearing aid settings

[1112] Specifically, it uses an AI model (machine learning model) to calculate settings, accepts analysis data sent from the device as input, and generates optimal hearing aid setting data as output.

[1113] Step 6:

[1114] The server sends the settings back to the device

[1115] Specifically, the system packages the configuration data and transmits it via HTTPS. As input, it accepts the hearing aid configuration data from the AI ​​model. As output, it sends the configuration data to the device.

[1116] Step 7:

[1117] The device suggests settings to the user

[1118] Specifically, it displays the settings on the display and asks for the user's approval. It accepts the settings data received from the server as input. It displays a dialog box with suggested settings to the user as output.

[1119] Step 8:

[1120] The user approves the settings

[1121] The specific action is to press the approval button on the display, the input is to accept the setting proposal dialog, and the output is to notify the terminal of the approval input.

[1122] Step 9:

[1123] The device adjusts the hearing aid

[1124] Specifically, the system writes the settings data to the hearing aid's internal memory and updates the acoustic profile. As input, it accepts user approval data. As output, the hearing aid settings are changed in real time.

[1125] Utilizing voice recognition technology

[1126] Step 1:

[1127] The user puts on the device and enables the speech recognition feature

[1128] Specifically, the operation is to turn on the device's voice recognition function, to accept a user instruction to enable the voice recognition function as input, and to activate the voice recognition function as output.

[1129] Step 2:

[1130] Your device uses the built-in microphone to collect ambient sounds

[1131] Specifically, the microphone converts audio data into a digital signal and generates a data stream. As input, it collects ambient audio. As output, it generates a digital audio signal.

[1132] Step 3:

[1133] The device sends the collected voice data to the server in real time.

[1134] Specifically, it buffers audio data and periodically sends packets to the server, accepts collected audio data as input, and transmits audio data to the server as output.

[1135] Step 4:

[1136] The server analyzes the voice data and converts it into text

[1137] Specifically, it converts voice data into text using a speech recognition engine (e.g., Google Speech-to-Text), accepts voice data sent from the device as input, and generates analyzed text data as output.

[1138] Step 5:

[1139] The server sends the converted text data to the terminal.

[1140] Specifically, it sends text data via the HTTPS protocol, accepts text data converted into character information as input, and sends the text data to the terminal as output.

[1141] Step 6:

[1142] The device displays the text data.

[1143] Specifically, it renders text data on the display in a format that is easy for the user to view, accepts text data from the server as input, and displays the text data on the display as output.

[1144] Barrier-free information provided

[1145] Step 1:

[1146] Users use their devices at train stations and public facilities

[1147] Specifically, the operation is to enable the device's audio collection function. As input, an instruction to enable the audio collection function by a user operation is accepted. As output, the audio collection function is started.

[1148] Step 2:

[1149] The device uses its built-in microphone to capture public announcements and announcements.

[1150] Specifically, the microphone converts audio data into a digital signal and detects specific frequency bands and patterns. It collects public announcements and announcements as inputs and generates a digital audio signal as output.

[1151] Step 3:

[1152] The device sends the collected information to a server

[1153] Specifically, it compresses the audio data appropriately and transmits it to the server in real time. As input, it accepts collected public announcements and information announcements. As output, it transmits the audio data to the server.

[1154] Step 4:

[1155] The server searches a database of video content with sign language commentary

[1156] Specifically, it queries the database to retrieve the relevant video links, accepts audio data sent from the device as input, and generates the associated video links as output.

[1157] Step 5:

[1158] The server sends the video link back to the device

[1159] Specifically, the video link is packaged and sent in a predetermined format, the relevant video link is accepted as input, and the video link is sent to the terminal as output.

[1160] Step 6:

[1161] The device notifies the user

[1162] Specifically, it displays a pop-up notification on the display to notify the user of the video link. As input, it accepts the video link received from the server. As output, it notifies the user of the video link.

[1163] Step 7:

[1164] The user selects the notification and the video plays on the device.

[1165] Specifically, the video player loads the URL and streams the video with sign language descriptions. As input, it accepts the user's notification selection. As output, it plays the video with sign language descriptions.

[1166] Personal Settings Function

[1167] Step 1:

[1168] User pre-sets specific sounds and words

[1169] Specifically, the user inputs specific sounds or words from the device's settings screen. The input accepts the user's settings for specific sounds or words. The output saves the setting data.

[1170] Step 2:

[1171] The device monitors surrounding sounds in real time

[1172] Specifically, it uses a built-in microphone to collect ambient sounds and then runs an algorithm to identify predefined sounds and words. The input is ambient sounds, and the output is the specific sounds or words detected.

[1173] Step 3:

[1174] When a specific sound or word is detected, it is sent to the server.

[1175] The specific operation is to send the detected sound or word data to the server. The input is to accept the detected data of a specific sound or word. The output is to send the data to the server.

[1176] Step 4:

[1177] After the server confirms, it sends a notification instruction to the device.

[1178] Specifically, it analyzes the received data, generates and sends appropriate notification instructions, accepts detection data from the device as input, and sends notification instructions to the device as output.

[1179] Step 5:

[1180] The device will alert the user with a visual alert and / or vibration.

[1181] As a specific operation, it activates the specified notification method (visual alert or vibration). As input, it accepts notification instructions from the server. As output, it warns the user.

[1182] Utilizing the Emotion Engine

[1183] Step 1:

[1184] Monitors emotional state in real time while the user is wearing the device

[1185] Specifically, it uses the device's built-in camera and microphone to collect the user's voice tone and facial expression data. As input, it collects the user's voice and facial expressions. As output, it generates emotion data.

[1186] Step 2:

[1187] The emotion engine analyzes the user's tone of voice and facial expressions to recognize their emotional state.

[1188] Specifically, it uses an analysis algorithm (e.g., Affectiva's SDK) to determine the emotional state. As input, it accepts collected voice and facial expression data. As output, it generates emotional state data.

[1189] Step 3:

[1190] Send the analyzed emotional state to the server

[1191] Specifically, the data acquired by the emotion engine is sent to the server. As input, emotional state data is accepted. As output, data is sent to the server.

[1192] Step 4:

[1193] The server determines the appropriate response and returns instructions to the device.

[1194] Specifically, the system analyzes the emotional state data and selects relaxing audio content or other appropriate responses. It accepts the emotional state data as input and generates and transmits instruction data to the device as output.

[1195] Step 5:

[1196] The device notifies the user of the suggestion

[1197] Specific actions include displaying suggestions to the user on a screen or by voice message, accepting instruction data from the server as input, and notifying the user as output.

[1198] (Application example 2)

[1199] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1200] Conventional hearing aids have limited means for properly analyzing the surrounding sound environment, and in particular lack systems for recognizing dangerous or suspicious sounds in real time and immediately notifying the user. This makes it difficult for people with hearing impairments or hard of hearing to quickly obtain information that will help them live safely.

[1201] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1202] In this invention, the server includes means for analyzing the sound environment in real time and recommending appropriate audio device settings, means for converting audio data into text, means for providing video content with sign language commentary, means for detecting specific sounds and words and notifying the user, and means for analyzing dangerous sounds and suspicious sounds and notifying the user. This makes it possible to detect dangerous sounds and suspicious sounds, and enables the user to obtain the information necessary for living safely in real time.

[1203] "Sound environment" refers to all the sound states and conditions that exist in the surroundings.

[1204] "Real-time" means that data collection and processing occurs in real time.

[1205] "Analysis" refers to examining the collected data in detail and clarifying its meaning and characteristics.

[1206] "Sound equipment" is a general term for equipment and devices used to collect, amplify, and reproduce sound.

[1207] "Voice data" means any digital recording of a human voice or other sound.

[1208] "Text" is audio or video converted into written information.

[1209] "Video content with sign language commentary" refers to video media that includes commentary in sign language.

[1210] "Specific sounds and words" refers to pre-set important sounds and keywords.

[1211] "Notification" refers to an alert or message that informs the user of specific information.

[1212] "Danger sounds" refer to sounds that may cause some kind of harm to the user.

[1213] "Suspicious audio" refers to audio that may indicate an abnormality or danger.

[1214] "User" refers to individuals, primarily deaf or hard of hearing, who use the system.

[1215] A system for implementing this invention includes an application installed on a device such as a smartphone or a security robot. The system configuration includes the following hardware and software:

[1216] The hardware includes a microphone for collecting sound, a display and vibrator for notifying the user, and in the case of security robots, a movement mechanism. Examples include smartphones and general-purpose robotic devices.

[1217] The software includes a speech analysis engine (e.g., Google Speech-to-Text API) that analyzes the sound environment in real time, an emotion analysis engine (e.g., Amazon Rekognition), and a cloud server (e.g., Amazon Web Services or Google Cloud Platform). Additionally, it runs a notification system (e.g., Firebase Cloud Messaging) to send notifications to users.

[1218] First, when a user turns on the device, the microphone collects the surrounding sound environment. The collected audio data is sent to a cloud server and analyzed using a voice analysis engine. During the analysis process, the audio data is converted into text and identified to determine whether it contains specific dangerous or suspicious sounds.

[1219] Additionally, in some cases, an emotion analysis engine is used to analyze facial expressions and vocal tones of people around you, collecting additional information if an abnormal emotional state is detected, and this data is also sent to a cloud server.

[1220] The analysis results are sent back to the user's device from the cloud server, whereupon they are instantly notified, if applicable, in the form of a visual alert on the display or a physical alert via a vibrator.

[1221] For example, if a security robot is used at home and detects any suspicious sounds at night, it will immediately send a notification to the user's smartphone, and the robot will automatically start recording and, if necessary, call the police.

[1222] An example of a prompt for the generative AI model is as follows:

[1223] "Please analyze the audio data below and identify any dangerous sounds or suspicious voices. Also, perform emotion analysis, and if any facial expressions of anger or fear are recognized, please provide that information as well."

[1224] This allows users to immediately sense danger and take necessary measures. This system provides an effective means to improve the quality of life for the deaf and hard of hearing.

[1225] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1226] Step 1:

[1227] Audio data collection

[1228] Input: Sound environment collected by a microphone installed on the device

[1229] Processing: The device collects surrounding sounds in real time and records them as audio data.

[1230] Output: The collected audio data is saved on the device.

[1231] Step 2:

[1232] Sending audio data

[1233] Input: Audio data collected in step 1

[1234] Processing: The device sends the voice data to the cloud server.

[1235] Output: The audio data is sent to the cloud server.

[1236] Step 3:

[1237] Audio analysis

[1238] Input: Audio data sent to the server

[1239] Processing: The server uses a speech analysis engine (such as the Google Speech-to-Text API) to convert the audio data into text, and then uses a generative AI model to analyze it to identify specific dangerous or suspicious sounds.

[1240] Output: Text data and analysis results are generated.

[1241] Step 4:

[1242] Emotion analysis

[1243] Input: Text data from step 3 and image data as needed

[1244] Processing: The server uses an emotion analysis engine (such as Amazon Rekognition) to analyze the emotional state of people around you. Specifically, it analyzes facial expressions and tone of voice to detect emotions such as stress and anger.

[1245] Output: The result of sentiment analysis is obtained.

[1246] Step 5:

[1247] Integration of analysis results

[1248] Input: Results of Step 3 and Step 4

[1249] Processing: The server integrates the results of the speech analysis and emotion analysis and determines the appropriate response for the user.

[1250] Output: As a result of the integration, information is generated that should be communicated to the user.

[1251] Step 6:

[1252] Sending notifications

[1253] Input: Notification information generated in step 5

[1254] Processing: The server uses a notification system (such as Firebase Cloud Messaging) to send a notification to the device.

[1255] Output: A notification will be displayed on the device.

[1256] Step 7:

[1257] Alert the user

[1258] Input: Notification information from step 6

[1259] Action: The device notifies the user of the notification using a visual alert and / or vibration, specifically by displaying a warning message on the device display and vibrating to get the user's attention.

[1260] Output: The user can view the notification and take action.

[1261] This allows the user to immediately detect dangerous or suspicious sounds and take necessary measures promptly.

[1262] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1263] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1264] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1265] [Third embodiment]

[1266] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1267] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1268] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1269] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1270] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1271] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1272] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1273] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1274] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1275] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1276] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1277] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1278] The present invention is a system for solving various problems that hearing-impaired and hard-of-hearing people encounter in their daily lives, and specifically includes a means for analyzing the sound environment in real time and recommending appropriate hearing aid settings, a means for converting audio data into text, a means for providing video content with sign language commentary, and a means for detecting and notifying specific sounds and words. Each function of this system is described in detail below.

[1279] Automatic hearing aid optimization

[1280] When a user puts on the wearable device and turns it on, the device automatically collects the surrounding sound environment using a microphone. The sound environment data is analyzed by the device to identify sound types such as speech, noise, and alarms. The analyzed data is sent to a server, which calculates the most suitable hearing aid settings. The calculated settings are sent back to the device and suggested to the user. If the user approves, the device automatically adjusts the hearing aid settings.

[1281] Utilizing voice recognition technology

[1282] When a user wears the device and activates the voice recognition function, the device continuously collects surrounding sounds. The collected voice data is sent to a server in real time, where a voice recognition engine analyzes the voice data and converts it into text. The converted text data is displayed on the device's display, allowing the user to visually check the content of conversations around them.

[1283] Barrier-free information provided

[1284] When a user uses a device in a station or public facility, the device collects public information such as announcements and guidance announcements. The collected information is sent to a server, which searches a database of video content with sign language commentary. When relevant video content is found, the server returns a link to the device and notifies the user. When the user selects the link, the video with sign language commentary is played on the device.

[1285] Personal Settings Function

[1286] The user can set specific sounds or words (such as their name or "danger") in advance, and the device will monitor the surrounding sounds in real time. When the set sound or word is detected, the corresponding data is sent to the server. After confirmation, the server sends a notification instruction to the device. The device will then alert the user using a visual alert or vibration.

[1287] Specific examples

[1288] 1. Examples of automatic hearing aid optimization:

[1289] When a user is talking with a friend at a coffee shop, the device recognizes the surrounding noise and the conversation and sends the data to the server. The server calculates the settings to reduce noise and emphasize the conversation and sends them back to the device. The device then displays a message to the user asking, "Would you like to change the settings to reduce noise and emphasize the conversation?" and the settings are applied if the user approves.

[1290] 2. Examples of voice recognition technology:

[1291] When a user is participating in a meeting at work, the device sends the contents of the meeting to a speech recognition engine, which converts the conversation into text and displays "About the next presentation" on the device display, allowing the user to easily understand the content.

[1292] 3. Examples of providing barrier-free information:

[1293] When the user is at a station, the device collects the station's departure announcements, and the server finds and sends a video with sign language explanations to the device. The user views the video with sign language explanations, which include the message "The next train departs at 10:25," and accurately understands the information.

[1294] 4. Examples of personalization features:

[1295] When a user is relaxing at home, a specific word, "danger," is set, and if that word is heard from outside, the device will vibrate to notify the user and display "Danger detected" on the screen to warn them.

[1296] In this way, the present invention provides an effective means for improving the quality of life of the deaf and hard of hearing.

[1297] The processing flow will be explained below.

[1298] Automatic hearing aid optimization

[1299] Step 1:

[1300] The device detects that the user is wearing a wearable device and has it turned on.

[1301] Step 2:

[1302] The device collects the surrounding sound environment in real time using its built-in microphone and generates sound environment data.

[1303] Step 3:

[1304] The device analyzes the sound environment data collected and identifies conversation sounds, noise, alarm sounds, etc.

[1305] Step 4:

[1306] The device sends the analysis results to the server.

[1307] Step 5:

[1308] The server analyzes the received sound environment data and calculates the optimal hearing aid settings.

[1309] Step 6:

[1310] The server returns the calculation results to the terminal.

[1311] Step 7:

[1312] The device will prompt the user, "There are new environment-based hearing aid settings. Would you like to apply them?"

[1313] Step 8:

[1314] The user selects "Yes" or "No."

[1315] Step 9:

[1316] If the user selects "Yes," the device will automatically change the hearing aid settings.

[1317] Step 10:

[1318] The device will notify the user that the new settings have been applied.

[1319] Utilizing voice recognition technology

[1320] Step 1:

[1321] The user puts on the device and enables the speech recognition feature.

[1322] Step 2:

[1323] The device continuously collects surrounding sounds using the built-in microphone.

[1324] Step 3:

[1325] The device sends the collected voice data to the server.

[1326] Step 4:

[1327] The server analyzes the received voice data and converts it into text using a voice recognition engine.

[1328] Step 5:

[1329] The server returns the converted text data to the terminal.

[1330] Step 6:

[1331] The device displays the received text on the display.

[1332] Step 7:

[1333] The device updates the text display in real time as new audio data is collected.

[1334] Barrier-free information provided

[1335] Step 1:

[1336] The user wears the device and enables the barrier-free information provision function.

[1337] Step 2:

[1338] The device uses a built-in microphone to collect public announcements and guidance broadcasts.

[1339] Step 3:

[1340] The device sends the collected voice data to the server.

[1341] Step 4:

[1342] The server analyzes the audio data and searches a database for relevant video content with sign language commentary.

[1343] Step 5:

[1344] The server sends a link to the appropriate video content back to the device.

[1345] Step 6:

[1346] The device notifies the user, "There is a related sign language explanation video. Would you like to play it?"

[1347] Step 7:

[1348] The user selects "Yes" or "No."

[1349] Step 8:

[1350] If the user selects "Yes," the device will open the video link and play the sign language explanation video.

[1351] Personal Settings Function

[1352] Step 1:

[1353] The user puts on the device and accesses the personalization screen.

[1354] Step 2:

[1355] The user sets and saves a specific sound or word (such as a name or "danger").

[1356] Step 3:

[1357] The device collects surrounding sounds in real time using a built-in microphone.

[1358] Step 4:

[1359] The device continuously compares the sounds it collects with preset sounds and words.

[1360] Step 5:

[1361] When the device detects a specific sound or word, it sends that data to the server.

[1362] Step 6:

[1363] The server analyzes the received data and confirms the detection results.

[1364] Step 7:

[1365] The server sends a notification instruction to the terminal.

[1366] Step 8:

[1367] The device will notify the user with a visual alert and / or vibration.

[1368] Step 9:

[1369] The device will display "A specific sound has been detected. Please be careful."

[1370] Example 1

[1371] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1372] This invention aims to solve various problems that the hearing impaired and hard of hearing face in their daily lives. Specifically, it aims to improve the quality of life for the hearing impaired by providing technology that efficiently analyzes surrounding sounds and automatically configures optimal hearing aids, as well as technology that converts speech to text, provides video content that incorporates sign language, and detects and notifies users of specific sounds and words.

[1373] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1374] In this invention, the server includes a means for analyzing the sound environment in real time and recommending appropriate hearing aid settings, a means for converting recorded audio data into text, and a means for providing video content with sign language commentary, thereby enabling hearing-impaired or hard-of-hearing people to efficiently grasp the surrounding sound environment and receive appropriate hearing assistance.

[1375] "Sound environment" refers to the state and situation of sound waves such as surrounding voices and noise.

[1376] "Analyzing in real time" means analyzing data immediately after it is collected, without any time delay.

[1377] A "hearing assistive device" is a device that assists the hearing of people with hearing loss or hearing impairments, and has the function of amplifying sound or emphasizing specific sounds.

[1378] "Audio recording data" refers to data that has been digitized and saved from audio collected by a microphone or other device.

[1379] "Converting to text" refers to the process of converting data such as audio and video into text information.

[1380] "Video content with sign language commentary" refers to videos or video materials that include visual sign language commentary.

[1381] "Detecting specific sounds or words" refers to the recognition system identifying and detecting predefined sounds or words.

[1382] "Notify" refers to informing a user of specific information or events.

[1383] "Device" refers to electronic equipment or devices, including hearing aids and wearable devices.

[1384] "Central server" refers to the main computer system that receives data from multiple devices, processes and analyzes it, and returns the results.

[1385] "Collect" refers to gathering data or information.

[1386] "Transmit" refers to sending data or information to another device or system.

[1387] "Calculating" refers to the process of deriving a specific result based on analytical results or algorithms.

[1388] "Automatically adjust" means that the device is automatically set to the optimum state without requiring any user operation.

[1389] "Public facilities" refers to buildings and places for use by the general public, including stations and libraries.

[1390] "Linking" means associating and making accessible particular information or content.

[1391] "Monitor" refers to watching for a particular condition or state.

[1392] This system analyzes the sound environment, converts speech into text, provides video content with sign language commentary, and detects and notifies specific sounds and words to solve various problems faced by the hearing impaired and hard of hearing in their daily lives. The system of this invention functions mainly through cooperation between the device (wearable terminal) and a central server.

[1393] Hardware and software used

[1394] 1. Device (wearable device)

[1395] microphone

[1396] Built-in processor

[1397] Communication modules (Wi-Fi, 5G, etc.)

[1398] display

[1399] Vibration Motor

[1400] battery

[1401] 2. Central Server

[1402] Generative AI Models

[1403] Speech Recognition Engine

[1404] Database (including video content with sign language commentary)

[1405] Communication Interface

[1406] Program processing

[1407] Analysis of sound environments and optimization of hearing aids

[1408] When a user puts on the device and turns it on, the device automatically collects the surrounding sound environment using a microphone. The sound environment data is analyzed by the device to identify sound types such as speech, noise, and alarms. This data is sent to a central server, where a generative AI model calculates optimal hearing aid settings. The calculation results are sent back to the device and recommended to the user. If the user approves, the device automatically adjusts the hearing aid settings.

[1409] Examples:

[1410] When a user is talking with a friend at a coffee shop, the device recognizes the surrounding noise and the conversation and sends the data to a central server. The server calculates the settings to reduce noise and enhance the conversation and sends them back to the device. The device then displays a message asking, "Do you want to change the settings to reduce noise and enhance the conversation?" and the settings are applied if the user confirms.

[1411] Utilizing voice recognition technology

[1412] When a user enables the device's voice recognition function, the device will continuously collect surrounding sounds. The collected voice data is sent to a central server in real time, where the voice recognition engine converts the speech into text. The converted text data is then displayed on the device's display, allowing the user to visually check the content of the conversations around them.

[1413] Examples:

[1414] When a user is participating in a meeting at work, the device sends the contents of the meeting to a speech recognition engine, and the converted text is displayed on the device's display as "About the next presentation," allowing the user to easily understand the content.

[1415] Providing video content with sign language commentary

[1416] When a user uses a device in a station or public facility, the device collects public information such as announcements and guidance announcements. The collected information is sent to a central server, which searches a database of videos with sign language descriptions. When relevant video content is found, the central server sends a link back to the device and notifies the user. When the user selects the link, the video with sign language descriptions is played on the device.

[1417] Examples:

[1418] When a user is at a station, the terminal collects the station's departure announcements, and the central server finds the video with sign language explanations and sends it to the terminal. The user views the video with sign language explanations, which include the message "The next train departs at 10:25," and accurately understands the information.

[1419] Detecting and announcing specific sounds and words

[1420] Users can pre-set specific sounds or words (e.g., "danger"), and the device will monitor surrounding sounds in real time. When the device detects a specific sound or word, the corresponding data is sent to a central server, which then confirms the detection and sends a notification instruction to the device. The device will then alert the user with a visual alert or vibration.

[1421] Examples:

[1422] If the user is relaxing at home and the word "danger" is set and hears that word from outside, the device will vibrate to notify the user and display "Danger detected" on the screen to warn them.

[1423] As described above, the present invention integrates multiple technical means to provide an effective method for the hearing impaired and hard of hearing to efficiently obtain information necessary in daily life.

[1424] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1425] Analysis of sound environments and optimization of hearing aids

[1426] Step 1:

[1427] When a user wears the device and turns it on, the device automatically collects the surrounding sound environment with its microphone. The input here is the surrounding sound environment, and the output is digitized sound environment data. The device's microphone converts analog audio into digital data.

[1428] Step 2:

[1429] The device analyzes the sound environment data collected in real time to identify the type of sound, such as speech, noise, or alarm sounds. The input is digitized sound environment data, and the output is information about the identified sound type (voice pattern data). The device's built-in processor runs an algorithm to analyze the data.

[1430] Step 3:

[1431] The device sends the analyzed voice pattern data to a central server. The input is the voice pattern data, and the output is the transmitted data. The data is sent to a cloud server using the device's communication module (Wi-Fi, 5G, etc.).

[1432] Step 4:

[1433] Based on the voice pattern data received by the server, a generative AI model calculates the optimal hearing aid settings. The input is voice pattern data, and the output is optimal setting data. The generative AI model analyzes the voice data and derives the appropriate parameters.

[1434] Step 5:

[1435] The server returns the optimal setting data to the terminal. The input is the optimal setting data, and the output is the setting data sent to the terminal. The data is sent using the server's communication interface.

[1436] Step 6:

[1437] The device receives the optimal setting data and displays the suggestion to the user. The input is the received setting data, and the output is the display of the suggestion. The device display shows "Would you like to change the settings to reduce noise and enhance speech?"

[1438] Step 7:

[1439] When the user presses the approval button, the device automatically adjusts the settings of the hearing aid. The input is the user's approval, and the output is the adjusted settings of the hearing aid. The device provides feedback to the user by vibration or sound.

[1440] Utilizing voice recognition technology

[1441] Step 1:

[1442] When a user enables the voice recognition function of the device, the device continuously collects surrounding sounds. The input is the surrounding sounds, and the output is digitized voice data. The microphone collects the voice and converts it into digital data.

[1443] Step 2:

[1444] The terminal transmits the collected voice data to a central server in real time. The input is the digitized voice data, and the output is the transmitted voice data. The data is transmitted using a communication module.

[1445] Step 3:

[1446] The server analyzes the voice data using a voice recognition engine and converts it into text data. The input is voice data and the output is text data. The voice recognition engine analyzes the voice pattern and generates the corresponding text.

[1447] Step 4:

[1448] Text data is sent from the central server to the terminal. The input is the text data, and the output is the text data sent to the terminal. The data is sent using the server's communication interface.

[1449] Step 5:

[1450] Text data is displayed on the terminal display. The input is the received text data, and the output is the text content displayed on the display. The user can check the content of the surrounding conversation by looking at the displayed text.

[1451] Providing video content with sign language commentary

[1452] Step 1:

[1453] When a user uses a device in a station or public facility, the terminal collects announcements and guidance announcements. The input is the public facility announcements, and the output is digitized voice data. The microphone collects the voice and converts it into digital data.

[1454] Step 2:

[1455] The collected voice data is transmitted to a central server. The input is the digitized voice data and the output is the transmitted voice data. A communication module is used to transmit the data to the server.

[1456] Step 3:

[1457] The server analyzes the audio data and searches a database of video content with sign language commentary. The input is the audio data, and the output is a link to related video content. The server's generative AI model analyzes the audio and finds related videos from the database.

[1458] Step 4:

[1459] The server returns a link to the associated video content to the terminal. The input is the link to the video content, and the output is the link sent to the terminal. The link is sent using a communication interface.

[1460] Step 5:

[1461] When the user selects a link, the device plays a video with sign language descriptions. The input is the link selected by the user and the output is the video that is played. The video with sign language descriptions appears on the device's display.

[1462] Detecting and announcing specific sounds and words

[1463] Step 1:

[1464] The user sets specific sounds or words (e.g., "danger") in advance on the device. The input is the user's settings for the sounds or words, and the output is the saved setting data. The sounds or words are entered using the device's setting screen and saved in a database.

[1465] Step 2:

[1466] The device monitors the surrounding sound in real time. The input is the surrounding sound, and the output is digitized audio data. The microphone collects the audio and converts it into digital data.

[1467] Step 3:

[1468] When a preset sound or word is detected, the device sends data to a central server. The input is the detected voice data, and the output is the transmitted data. The data is sent to the server using a communication module.

[1469] Step 4:

[1470] The server checks the received data and sends a notification instruction to the device. The input is the received voice data, and the output is the notification instruction. The server's generative AI model analyzes the data and determines the appropriate response.

[1471] Step 5:

[1472] The device alerts the user with a visual alert or vibration. The input is a notification instruction, and the output is a visual alert or vibration notification. The device activates the vibration motor and displays "Danger detected" on the display.

[1473] (Application example 1)

[1474] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1475] The present invention aims to solve the problem that hearing-impaired and hard-of-hearing people have difficulty recognizing voices and grasping important information in daily life and in certain situations (e.g., when using self-driving vehicles). In particular, there is a problem that emergency situations and important announcements are difficult to convey to hearing-impaired and hard-of-hearing people. This may result in situations where the safety and comfort of users are lacking.

[1476] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1477] In this invention, the server includes a means for analyzing the sound environment in real time and recommending appropriate hearing aid settings, a means for converting voice data into text, and a means for providing video content with sign language commentary. This allows hearing-impaired and hard-of-hearing individuals to properly understand the surrounding sound environment and automatically optimize their hearing aid settings. Furthermore, by providing a means for collecting voice data in real time and notifying the user based on specific keywords, it becomes possible to reliably communicate emergencies and important announcements. Furthermore, by including a means for displaying voice recognition results on an in-vehicle display, users can visually confirm surrounding voice information, thereby improving safety and comfort when using autonomous vehicles.

[1478] "Analyzing the sound environment in real time and recommending appropriate hearing aid settings" means continuously monitoring the sound environment using sensors and microphones, classifying the type and level of sound using an analysis device, and providing optimal hearing aid settings based on that.

[1479] "Converting voice data to text" refers to the process of analyzing collected voice data using speech recognition technology and converting it into a corresponding text format.

[1480] "Providing video content with sign language commentary" is a service that allows deaf and hard of hearing people to obtain information visually by providing video information with sign language interpretation.

[1481] "Detecting and notifying specific sounds and words" refers to a mechanism that detects specific pre-set sounds or utterances and notifies the user.

[1482] "Collecting voice data in real time and notifying the user based on specific keywords" means accumulating voice data in real time, identifying specific keywords in the data, and providing the user with appropriate alerts.

[1483] "Displaying voice recognition results on an in-vehicle display" means displaying text data obtained through voice recognition technology on a display inside an autonomous vehicle to provide information visually to the user.

[1484] The present invention is a system for solving various problems that hearing-impaired and hard-of-hearing people encounter in their daily lives and when using self-driving vehicles. This system includes a means for analyzing the sound environment in real time and recommending appropriate hearing aid settings, a means for converting voice data into text, a means for providing video content with sign language commentary, and a means for detecting and notifying specific sounds and words.

[1485] System configuration

[1486] The system consists of the following main components:

[1487] Hardware

[1488] 1. Microphone: A device that collects ambient sounds, allowing for real-time monitoring of the sound environment.

[1489] 2. Server: Analyzes the sound data and provides instructions for optimal settings for the hearing aid and text conversion.

[1490] 3. Autonomous vehicle display: A device for displaying voice recognition results and system notifications.

[1491] 4. Vibration motor: To notify the user when a specific sound or keyword is detected.

[1492] software

[1493] 1. Speech recognition engine (e.g., Google Speech Recognition API): Collects speech in real time and converts it into text.

[1494] 2. Analysis algorithm: Analyzes the sound environment and recommends appropriate hearing aid settings.

[1495] 3. Notification system: Detects specific keywords or voice anomalies and notifies the user.

[1496] System Operation

[1497] 1. Real-time analysis of sound environments

[1498] When the user operates the system, the microphone begins to collect the surrounding sound environment. The collected sound environment data is sent to the server, where an analysis algorithm identifies the type of sound, such as speech, noise, or alarm. Based on the analyzed data, the server calculates the optimal hearing aid settings and sends them back to the device. If the user approves, the device automatically adjusts the hearing aid settings.

[1499] 2. Use of voice recognition technology

[1500] When the user wears the device and activates the voice recognition function, the device continuously collects surrounding sounds. The collected voice data is sent to a server in real time, where a voice recognition engine analyzes the voice data and converts it into text. The converted text data is displayed on the vehicle's display, allowing the user to visually check the content of surrounding conversations and announcements.

[1501] 3. Providing barrier-free information

[1502] When a user is using an autonomous vehicle, the device collects information such as in-car announcements and emergency alerts. The collected information is sent to a server, which searches a database of video content with sign language commentary. When relevant video content is found, the server returns a link to the device and notifies the user. When the user selects the link, the video with sign language commentary is played on the in-car display.

[1503] 4. Personal settings function

[1504] The user can set specific sounds or words (such as their name or "danger") in advance, and the device will monitor the surrounding sounds in real time. When the set sound or word is detected, the corresponding data is sent to the server. After confirmation, the server sends a notification instruction to the device. The device will then alert the user using a visual alert or vibration.

[1505] Specific examples

[1506] 1. A concrete example of automatic hearing aid optimization

[1507] When a user is talking with a friend at a coffee shop, the device recognizes the surrounding noise and the conversation and sends the data to the server. The server calculates the settings to reduce noise and emphasize the conversation and sends them back to the device. The device then displays a message to the user asking, "Would you like to change the settings to reduce noise and emphasize the conversation?" and the settings are applied if the user approves.

[1508] 2. Specific examples of voice recognition technology

[1509] When a user is participating in a meeting at work, the device sends the contents of the meeting to a speech recognition engine, which converts the conversation into text and displays "About the next presentation" on the device display, allowing the user to easily understand the content.

[1510] 3. Examples of providing barrier-free information

[1511] When the user is at a station, the device collects the station's departure announcements, and the server finds and sends a video with sign language explanations to the device. The user views the video with sign language explanations, which include the message "The next train departs at 10:25," and accurately understands the information.

[1512] 4. Examples of personal settings functions

[1513] When a user is relaxing at home, a specific word, "danger," is set, and if that word is heard from outside, the device will vibrate to notify the user and display "Danger detected" on the screen to warn them.

[1514] Example prompts for generative AI models

[1515] "You are designing a self-driving car system for the hearing impaired. The system uses speech recognition to convert emergency situations and announcements into text and notify the user. Extract the emergency content from the following audio data and display it as text."

[1516] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1517] Step 1:

[1518] The user operates the system. The user gets into the autonomous vehicle and a microphone connected to the terminal collects surrounding sounds in real time. The collected sound data is input into the terminal.

[1519] Step 2:

[1520] The device analyzes the collected sound data and classifies the sound environment. The device's internal analysis algorithm classifies the sound data into categories such as speech, noise, and alarm sounds. This analyzed data is sent to the server.

[1521] Step 3:

[1522] The server receives the sound environment data and calculates the optimal hearing aid settings. Based on the analyzed sound data, calculations are performed to automatically optimize the hearing aid settings. These optimized settings are then sent from the server to the device.

[1523] Step 4:

[1524] The device presents the hearing aid settings to the user. The device asks the user whether to apply the new settings and obtains their approval. When the user presses the approval button, the device changes the hearing aid settings.

[1525] Step 5:

[1526] The device continuously collects voice data and sends it to a voice recognition engine. Based on the voice data input, real-time voice recognition processing is performed and the data is converted into text. This text data is then sent back to the device.

[1527] Step 6:

[1528] The device displays the voice recognition results on the vehicle display, and text data is output to the display so that the user can visually check the content of surrounding conversations and announcements.

[1529] Step 7:

[1530] The device monitors specific keywords and sounds to detect abnormalities. When a pre-defined keyword (e.g., "danger" or "emergency") is detected, the data is sent to the server. After the server confirms the detection, it sends a notification instruction to the device.

[1531] Step 8:

[1532] The device will alert the user. If an abnormal sound or specific keyword is detected, the device will alert the user with a visual alert or vibration notification.

[1533] The above process will create a system that improves the safety and convenience of people who are deaf or hard of hearing when using self-driving vehicles.

[1534] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1535] The present invention is a system for solving various problems encountered in daily life by people with hearing impairments or hard of hearing. Specifically, it includes a means for analyzing the sound environment in real time and recommending appropriate hearing aid settings, a means for converting audio data into text, a means for providing video content with sign language commentary, and a means for detecting and notifying specific sounds or words. Furthermore, by combining it with an emotion engine that recognizes and analyzes the user's emotions, further user assistance is provided. Each function of this system is described in detail below.

[1536] Automatic hearing aid optimization

[1537] When a user puts on the wearable device and turns it on, the device automatically collects the surrounding sound environment using a microphone. The sound environment data is analyzed by the device to identify sound types such as speech, noise, and alarms. The analyzed data is sent to a server, which calculates the most suitable hearing aid settings. The calculated settings are sent back to the device and suggested to the user. If the user approves, the device automatically adjusts the hearing aid settings.

[1538] Utilizing voice recognition technology

[1539] When a user wears the device and activates the voice recognition function, the device continuously collects surrounding sounds. The collected voice data is sent to a server in real time, where a voice recognition engine analyzes the voice data and converts it into text. The converted text data is displayed on the device's display, allowing the user to visually check the content of conversations around them.

[1540] Barrier-free information provided

[1541] When a user uses a device in a station or public facility, the device collects public information such as announcements and guidance announcements. The collected information is sent to a server, which searches a database of video content with sign language commentary. When relevant video content is found, the server returns a link to the device and notifies the user. When the user selects the link, the video with sign language commentary is played on the device.

[1542] Personal Settings Function

[1543] The user can set specific sounds or words (such as their name or "danger") in advance, and the device will monitor the surrounding sounds in real time. When the set sound or word is detected, the corresponding data is sent to the server. After confirmation, the server sends a notification instruction to the device. The device will then alert the user using a visual alert or vibration.

[1544] Utilizing the Emotion Engine

[1545] The emotion engine monitors the user's emotional state in real time while the user is wearing the device. The emotion engine analyzes the user's tone of voice and facial expressions (e.g., using the camera function) to recognize emotions such as stress, joy, and sadness. Once the emotional state is analyzed, the results are sent to the server, which determines the appropriate response and sends instructions back to the device.

[1546] Specific examples

[1547] 1. Examples of automatic hearing aid optimization:

[1548] When a user is talking with a friend at a coffee shop, the device recognizes the surrounding noise and the conversation and sends the data to the server. The server calculates the settings to reduce noise and emphasize the conversation and sends them back to the device. The device then displays a message to the user asking, "Would you like to change the settings to reduce noise and emphasize the conversation?" and the settings are applied if the user approves.

[1549] 2. Examples of voice recognition technology:

[1550] When a user is participating in a meeting at work, the device sends the contents of the meeting to a speech recognition engine, which converts the conversation into text and displays "About the next presentation" on the device display, allowing the user to easily understand the content.

[1551] 3. Examples of providing barrier-free information:

[1552] When the user is at a station, the device collects the station's departure announcements, and the server finds and sends a video with sign language explanations to the device. The user views the video with sign language explanations, which include the message "The next train departs at 10:25," and accurately understands the information.

[1553] 4. Examples of personalization features:

[1554] When a user is relaxing at home, a specific word, "danger," is set, and if that word is heard from outside, the device will vibrate to notify the user and display "Danger detected" on the screen to warn them.

[1555] 5. Examples of Emotion Engines:

[1556] If a user feels stressed during a meeting, the emotion engine analyzes the user's facial expressions and tone of voice to detect the stress level. The server recommends audio content with a relaxing effect, and the device notifies the user, "Would you like to play relaxing music?" and plays the music if the user agrees.

[1557] In this way, the present invention provides an effective means for improving the quality of life of the deaf and hard of hearing.

[1558] The processing flow will be explained below.

[1559] Automatic hearing aid optimization

[1560] Step 1:

[1561] The user puts on the wearable device and turns it on.

[1562] Step 2:

[1563] The device automatically collects the surrounding sound environment in real time using the built-in microphone.

[1564] Step 3:

[1565] The device analyzes the sound environment data collected and identifies the type of sound (conversation, noise, alarm, etc.).

[1566] Step 4:

[1567] The device sends the analysis results to the server.

[1568] Step 5:

[1569] The server analyzes the received sound environment data and calculates appropriate hearing aid settings.

[1570] Step 6:

[1571] The server returns the calculated hearing aid settings to the device.

[1572] Step 7:

[1573] The device will prompt the user, "There are new environment-based hearing aid settings. Would you like to apply them?"

[1574] Step 8:

[1575] The user selects "Yes" or "No."

[1576] Step 9:

[1577] If the user selects "Yes," the device will automatically change the hearing aid settings.

[1578] Step 10:

[1579] The device will notify the user that the new settings have been applied.

[1580] Utilizing voice recognition technology

[1581] Step 1:

[1582] The user puts on the device and enables the speech recognition feature.

[1583] Step 2:

[1584] The device continuously collects surrounding sounds using the built-in microphone.

[1585] Step 3:

[1586] The device sends the collected voice data to the server.

[1587] Step 4:

[1588] The voice data received by the server is analyzed using a voice recognition engine and converted into text.

[1589] Step 5:

[1590] The server returns the converted text data to the terminal.

[1591] Step 6:

[1592] The text data received by the terminal is displayed on the display.

[1593] Step 7:

[1594] The device continues to convert new voice data collected into text in real time, updating the text on the display.

[1595] Barrier-free information provided

[1596] Step 1:

[1597] The user wears the device and enables the barrier-free information provision function.

[1598] Step 2:

[1599] The device uses a built-in microphone to collect public announcements and guidance broadcasts.

[1600] Step 3:

[1601] The device sends the collected voice data to the server.

[1602] Step 4:

[1603] The server analyzes the audio data and searches a database for relevant video content with sign language commentary.

[1604] Step 5:

[1605] The server sends a link to the appropriate video content back to the device.

[1606] Step 6:

[1607] The device notifies the user, "There is a related sign language explanation video. Would you like to play it?"

[1608] Step 7:

[1609] The user selects "Yes" or "No."

[1610] Step 8:

[1611] If the user selects "Yes," the device will open the video link and play the sign language explanation video.

[1612] Personal Settings Function

[1613] Step 1:

[1614] The user puts on the device and accesses the personalization screen.

[1615] Step 2:

[1616] The user sets and saves a specific sound or word (such as a name or "danger").

[1617] Step 3:

[1618] The device collects surrounding sounds in real time using a built-in microphone.

[1619] Step 4:

[1620] The device continuously compares the sounds it collects with preset sounds and words.

[1621] Step 5:

[1622] When the device detects a specific sound or word, it sends that data to the server.

[1623] Step 6:

[1624] The server analyzes the received data and confirms the detection results.

[1625] Step 7:

[1626] The server sends a notification instruction to the terminal.

[1627] Step 8:

[1628] The device will notify the user with a visual alert and / or vibration.

[1629] Step 9:

[1630] The device will display "A specific sound has been detected. Please be careful."

[1631] Utilizing the Emotion Engine

[1632] Step 1:

[1633] The user puts on the device and activates the emotion engine function.

[1634] Step 2:

[1635] The device collects the user's voice tone and facial expressions using the built-in microphone and camera.

[1636] Step 3:

[1637] The device sends the collected voice tone and facial expression data to the server.

[1638] Step 4:

[1639] The server analyzes the received data using an emotion engine to identify the user's emotional state.

[1640] Step 5:

[1641] The server returns the analysis results to the device and instructs it on the appropriate response.

[1642] Step 6:

[1643] The device will notify the user, for example, "Emotional state detected. Would you like to relax?"

[1644] Step 7:

[1645] The user selects "Yes" or "No."

[1646] Step 8:

[1647] If the user selects "Yes," the terminal plays back audio content that has a relaxing effect.

[1648] Step 9:

[1649] The device again monitors changes in the user's emotional state and makes adjustments as needed.

[1650] Example 2

[1651] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1652] Among the challenges faced by people with hearing impairments and those with hearing loss in their daily lives are adjusting their hearing aid settings to changes in the surrounding sound environment, understanding conversations and important announcements, obtaining information in public facilities, and recognizing emergency situations in real time. Effectively resolving these challenges requires a system that can collect and analyze various voice and environmental data in real time and provide appropriate support. Previously, no system offered all of these functions comprehensively, forcing users to use multiple devices and services at the same time, which was inconvenient.

[1653] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1654] In this invention, the server includes a means for analyzing the sound environment in real time and suggesting optimal hearing aid settings, a means for converting voice data into text, a means for providing visual content with sign language explanations, a means for detecting specific sounds and words and issuing a warning, and a means for analyzing the user's emotional state and suggesting appropriate responses. This allows users to receive a variety of support from a single system, comprehensively resolving various issues in daily life.

[1655] "Analyzing the sound environment in real time and proposing optimal hearing aid settings" means collecting surrounding sounds using a microphone or other device, analyzing the collected sound data, and calculating and providing hearing aid settings that optimize the way the user hears sound.

[1656] "Converting voice data to text" means collecting surrounding voices as digital signals, analyzing the voice data, converting it into text information, and displaying it.

[1657] "Providing visual content with sign language explanations" means providing users with videos or animations with added sign language explanations to supplement visual means of communication.

[1658] "Detecting specific sounds or words and issuing a warning" means recognizing specific sounds or words that have been set in advance, and issuing a notification or warning to the user when they are detected.

[1659] "Analyzing the user's emotional state and suggesting appropriate responses" means analyzing the user's tone of voice, facial expressions, etc. to recognize their emotional state, and then suggesting relaxation content or other appropriate support based on the results.

[1660] The present invention is a comprehensive system for solving various problems faced by the deaf and hard of hearing in daily life. The system includes means for analyzing the sound environment in real time and suggesting optimal hearing aid settings, means for converting speech data into text, means for providing visual content with sign language explanations, means for detecting specific sounds and words and issuing warnings, and means for analyzing the user's emotional state and suggesting appropriate responses.

[1661] Automatic hearing aid optimization

[1662] 1. Acoustic environment data collection

[1663] The user puts on the wearable device and turns it on.

[1664] The device uses a built-in microphone to collect the surrounding sound environment in real time.

[1665] Hardware used: Wearable devices (hearing aids)

[1666] 2. Analysis of sound environment data

[1667] The device analyzes the sound data collected and identifies speech, noise, alarm sounds, etc.

[1668] Software used: Audio analysis algorithm (e.g., FFT algorithm)

[1669] 3. Calculation and recommendation of hearing aid settings

[1670] The device sends the analyzed sound environment data to the server, which then calculates the optimal hearing aid settings.

[1671] The server sends the calculated settings back to the terminal and suggests them to the user.

[1672] Specific examples

[1673] When a user is talking with a friend at a coffee shop, the device recognizes the surrounding noise and the conversation and sends the data to the server. The server calculates the settings to reduce noise and emphasize the conversation and sends them back to the device. The device then displays a message to the user asking, "Would you like to change the settings to reduce noise and emphasize the conversation?" and the settings are applied if the user approves.

[1674] Utilizing voice recognition technology

[1675] 1. Collection of audio data

[1676] The user puts on the device and enables the speech recognition feature.

[1677] Your device uses a built-in microphone to collect ambient sounds.

[1678] 2. Analysis and display of audio data

[1679] The voice data collected by the device is sent to the server in real time, where it is analyzed by the server's voice recognition engine and converted into text.

[1680] The converted text data is returned to the terminal and displayed on the display.

[1681] Specific examples

[1682] When a user is participating in a meeting at work, the device sends the contents of the meeting to a speech recognition engine, which converts the conversation into text and displays "About the next presentation" on the device display.

[1683] Barrier-free information provided

[1684] 1. Collection of public information

[1685] The user uses the device at a train station or public facility.

[1686] The device uses a built-in microphone to collect public announcements and guidance broadcasts.

[1687] 2. Providing related videos

[1688] The terminal sends the collected information to a server, which searches a database of video content with sign language commentary.

[1689] When related video content is found, the server returns the link to the terminal and notifies the user.

[1690] Specific examples

[1691] When the user is at a station, the device collects the station's departure announcements, and the server finds and sends a video with sign language explanations to the device. The user views the video with sign language explanations, which include the message "The next train departs at 10:25," and accurately understands the information.

[1692] Personal Settings Function

[1693] 1. Monitor specific sounds and words in real time

[1694] The user pre-sets specific sounds and words.

[1695] The device constantly monitors surrounding sounds and, when it detects preset sounds or words, sends them to the server.

[1696] 2. Issuing a warning

[1697] After the server confirms, it sends a notification instruction to the terminal.

[1698] The device will alert the user with a visual alert and / or vibration.

[1699] Specific examples

[1700] When a user is relaxing at home and has set a specific word, "danger," and hears that word from outside, the device will vibrate to notify the user and display "Danger detected" on the screen to warn them.

[1701] Utilizing the Emotion Engine

[1702] 1. Monitoring your emotional state

[1703] The device monitors the user's emotional state in real time while they are wearing it.

[1704] The emotion engine analyzes the user's tone of voice and facial expressions to recognize their emotional state.

[1705] 2. Proposal of appropriate action

[1706] The analyzed emotional state is sent to a server, which determines the appropriate response.

[1707] The terminal notifies the user of the suggestion.

[1708] Specific examples

[1709] If a user feels stressed during a meeting, the emotion engine analyzes the user's facial expressions and tone of voice to detect the stress level. The server recommends audio content with a relaxing effect, and the device notifies the user, "Would you like to play relaxing music?" and plays the music if the user agrees.

[1710] Prompt Sentence Examples

[1711] Automatic hearing aid optimization prompts

[1712] Describe a scenario where a user is talking with a friend at a coffee shop and their hearing aids analyze the surrounding sound environment and suggest optimal settings to enhance the conversation and reduce noise.

[1713] Voice recognition technology prompts

[1714] Please give a specific example of a situation where a user is participating in a meeting at work and speech recognition technology is used to convert the meeting content into text and display it on a display.

[1715] As can be seen, the present invention can provide an effective means for improving the quality of life of the deaf and hard of hearing in a variety of situations.

[1716] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1717] Automatic hearing aid optimization

[1718] Step 1:

[1719] The user puts on the wearable device and turns it on.

[1720] The specific operation is that the user physically switches on the device and confirms the connection. As input, the device accepts the power-on state. As output, the device starts up and prepares for the next processing step.

[1721] Step 2:

[1722] The device uses a built-in microphone to collect the surrounding sound environment in real time.

[1723] Specifically, the microphone converts ambient sound into a digital signal and acquires volume and frequency data. The input is the surrounding sound environment, and the output is sound data generated as a digital signal.

[1724] Step 3:

[1725] Analyzes the sound environment data collected by the device and distinguishes between speech, noise, alarm sounds, etc.

[1726] Specifically, it uses a sound analysis algorithm (e.g., FFT algorithm) to extract and classify sound features. As input, it accepts the sound data collected in step 2. As output, it generates data on the identified sound types and their features.

[1727] Step 4:

[1728] The device sends the analysis data to the server.

[1729] Specifically, it transmits data via an internet connection using the HTTPS protocol. As input, it accepts analyzed sound environment data. As output, it sends the data to a server.

[1730] Step 5:

[1731] The server calculates the optimal hearing aid settings

[1732] Specifically, it uses an AI model (machine learning model) to calculate settings, accepts analysis data sent from the device as input, and generates optimal hearing aid setting data as output.

[1733] Step 6:

[1734] The server sends the settings back to the device

[1735] Specifically, the system packages the configuration data and transmits it via HTTPS. As input, it accepts the hearing aid configuration data from the AI ​​model. As output, it sends the configuration data to the device.

[1736] Step 7:

[1737] The device suggests settings to the user

[1738] Specifically, it displays the settings on the display and asks for the user's approval. It accepts the settings data received from the server as input. It displays a dialog box with suggested settings to the user as output.

[1739] Step 8:

[1740] The user approves the settings

[1741] The specific action is to press the approval button on the display, the input is to accept the setting proposal dialog, and the output is to notify the terminal of the approval input.

[1742] Step 9:

[1743] The device adjusts the hearing aid

[1744] Specifically, the system writes the settings data to the hearing aid's internal memory and updates the acoustic profile. As input, it accepts user approval data. As output, the hearing aid settings are changed in real time.

[1745] Utilizing voice recognition technology

[1746] Step 1:

[1747] The user puts on the device and enables the speech recognition feature

[1748] Specifically, the operation is to turn on the device's voice recognition function, to accept a user instruction to enable the voice recognition function as input, and to activate the voice recognition function as output.

[1749] Step 2:

[1750] Your device uses the built-in microphone to collect ambient sounds

[1751] Specifically, the microphone converts audio data into a digital signal and generates a data stream. As input, it collects ambient audio. As output, it generates a digital audio signal.

[1752] Step 3:

[1753] The device sends the collected voice data to the server in real time.

[1754] Specifically, it buffers audio data and periodically sends packets to the server, accepts collected audio data as input, and transmits audio data to the server as output.

[1755] Step 4:

[1756] The server analyzes the voice data and converts it into text

[1757] Specifically, it converts voice data into text using a speech recognition engine (e.g., Google Speech-to-Text), accepts voice data sent from the device as input, and generates analyzed text data as output.

[1758] Step 5:

[1759] The server sends the converted text data to the terminal.

[1760] Specifically, it sends text data via the HTTPS protocol, accepts text data converted into character information as input, and sends the text data to the terminal as output.

[1761] Step 6:

[1762] The device displays the text data.

[1763] Specifically, it renders text data on the display in a format that is easy for the user to view, accepts text data from the server as input, and displays the text data on the display as output.

[1764] Barrier-free information provided

[1765] Step 1:

[1766] Users use their devices at train stations and public facilities

[1767] Specifically, the operation is to enable the device's audio collection function. As input, an instruction to enable the audio collection function by a user operation is accepted. As output, the audio collection function is started.

[1768] Step 2:

[1769] The device uses its built-in microphone to capture public announcements and announcements.

[1770] Specifically, the microphone converts audio data into a digital signal and detects specific frequency bands and patterns. It collects public announcements and announcements as inputs and generates a digital audio signal as output.

[1771] Step 3:

[1772] The device sends the collected information to a server

[1773] Specifically, it compresses the audio data appropriately and transmits it to the server in real time. As input, it accepts collected public announcements and information announcements. As output, it transmits the audio data to the server.

[1774] Step 4:

[1775] The server searches a database of video content with sign language commentary

[1776] Specifically, it queries the database to retrieve the relevant video links, accepts audio data sent from the device as input, and generates the associated video links as output.

[1777] Step 5:

[1778] The server sends the video link back to the device

[1779] Specifically, the video link is packaged and sent in a predetermined format, the relevant video link is accepted as input, and the video link is sent to the terminal as output.

[1780] Step 6:

[1781] The device notifies the user

[1782] Specifically, it displays a pop-up notification on the display to notify the user of the video link. As input, it accepts the video link received from the server. As output, it notifies the user of the video link.

[1783] Step 7:

[1784] The user selects the notification and the video plays on the device.

[1785] Specifically, the video player loads the URL and streams the video with sign language descriptions. As input, it accepts the user's notification selection. As output, it plays the video with sign language descriptions.

[1786] Personal Settings Function

[1787] Step 1:

[1788] User pre-sets specific sounds and words

[1789] Specifically, the user inputs specific sounds or words from the device's settings screen. The input accepts the user's settings for specific sounds or words. The output saves the setting data.

[1790] Step 2:

[1791] The device monitors surrounding sounds in real time

[1792] Specifically, it uses a built-in microphone to collect ambient sounds and then runs an algorithm to identify predefined sounds and words. The input is ambient sounds, and the output is the specific sounds or words detected.

[1793] Step 3:

[1794] When a specific sound or word is detected, it is sent to the server.

[1795] The specific operation is to send the detected sound or word data to the server. The input is to accept the detected data of a specific sound or word. The output is to send the data to the server.

[1796] Step 4:

[1797] After the server confirms, it sends a notification instruction to the device.

[1798] Specifically, it analyzes the received data, generates and sends appropriate notification instructions, accepts detection data from the device as input, and sends notification instructions to the device as output.

[1799] Step 5:

[1800] The device will alert the user with a visual alert and / or vibration.

[1801] As a specific operation, it activates the specified notification method (visual alert or vibration). As input, it accepts notification instructions from the server. As output, it warns the user.

[1802] Utilizing the Emotion Engine

[1803] Step 1:

[1804] Monitors emotional state in real time while the user is wearing the device

[1805] Specifically, it uses the device's built-in camera and microphone to collect the user's voice tone and facial expression data. As input, it collects the user's voice and facial expressions. As output, it generates emotion data.

[1806] Step 2:

[1807] The emotion engine analyzes the user's tone of voice and facial expressions to recognize their emotional state.

[1808] Specifically, it uses an analysis algorithm (e.g., Affectiva's SDK) to determine the emotional state. As input, it accepts collected voice and facial expression data. As output, it generates emotional state data.

[1809] Step 3:

[1810] Send the analyzed emotional state to the server

[1811] Specifically, the data acquired by the emotion engine is sent to the server. As input, emotional state data is accepted. As output, data is sent to the server.

[1812] Step 4:

[1813] The server determines the appropriate response and returns instructions to the device.

[1814] Specifically, the system analyzes the emotional state data and selects relaxing audio content or other appropriate responses. It accepts the emotional state data as input and generates and transmits instruction data to the device as output.

[1815] Step 5:

[1816] The device notifies the user of the suggestion

[1817] Specific actions include displaying suggestions to the user on a screen or by voice message, accepting instruction data from the server as input, and notifying the user as output.

[1818] (Application example 2)

[1819] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1820] Conventional hearing aids have limited means for properly analyzing the surrounding sound environment, and in particular lack systems for recognizing dangerous or suspicious sounds in real time and immediately notifying the user. This makes it difficult for people with hearing impairments or hard of hearing to quickly obtain information that will help them live safely.

[1821] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1822] In this invention, the server includes means for analyzing the sound environment in real time and recommending appropriate audio device settings, means for converting audio data into text, means for providing video content with sign language commentary, means for detecting specific sounds and words and notifying the user, and means for analyzing dangerous sounds and suspicious sounds and notifying the user. This makes it possible to detect dangerous sounds and suspicious sounds, and enables the user to obtain the information necessary for living safely in real time.

[1823] "Sound environment" refers to all the sound states and conditions that exist in the surroundings.

[1824] "Real-time" means that data collection and processing occurs in real time.

[1825] "Analysis" refers to examining the collected data in detail and clarifying its meaning and characteristics.

[1826] "Sound equipment" is a general term for equipment and devices used to collect, amplify, and reproduce sound.

[1827] "Voice data" means any digital recording of a human voice or other sound.

[1828] "Text" is audio or video converted into written information.

[1829] "Video content with sign language commentary" refers to video media that includes commentary in sign language.

[1830] "Specific sounds and words" refers to pre-set important sounds and keywords.

[1831] "Notification" refers to an alert or message that informs the user of specific information.

[1832] "Danger sounds" refer to sounds that may cause some kind of harm to the user.

[1833] "Suspicious audio" refers to audio that may indicate an abnormality or danger.

[1834] "User" refers to individuals, primarily deaf or hard of hearing, who use the system.

[1835] A system for implementing this invention includes an application installed on a device such as a smartphone or a security robot. The system configuration includes the following hardware and software:

[1836] The hardware includes a microphone for collecting sound, a display and vibrator for notifying the user, and in the case of security robots, a movement mechanism. Examples include smartphones and general-purpose robotic devices.

[1837] The software includes a speech analysis engine (e.g., Google Speech-to-Text API) that analyzes the sound environment in real time, an emotion analysis engine (e.g., Amazon Rekognition), and a cloud server (e.g., Amazon Web Services or Google Cloud Platform). Additionally, it runs a notification system (e.g., Firebase Cloud Messaging) to send notifications to users.

[1838] First, when a user turns on the device, the microphone collects the surrounding sound environment. The collected audio data is sent to a cloud server and analyzed using a voice analysis engine. During the analysis process, the audio data is converted into text and identified to determine whether it contains specific dangerous or suspicious sounds.

[1839] Additionally, in some cases, an emotion analysis engine is used to analyze facial expressions and vocal tones of people around you, collecting additional information if an abnormal emotional state is detected, and this data is also sent to a cloud server.

[1840] The analysis results are sent back to the user's device from the cloud server, whereupon they are instantly notified, if applicable, in the form of a visual alert on the display or a physical alert via a vibrator.

[1841] For example, if a security robot is used at home and detects any suspicious sounds at night, it will immediately send a notification to the user's smartphone, and the robot will automatically start recording and, if necessary, call the police.

[1842] An example of a prompt for the generative AI model is as follows:

[1843] "Please analyze the audio data below and identify any dangerous sounds or suspicious voices. Also, perform emotion analysis, and if any facial expressions of anger or fear are recognized, please provide that information as well."

[1844] This allows users to immediately sense danger and take necessary measures. This system provides an effective means to improve the quality of life for the deaf and hard of hearing.

[1845] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1846] Step 1:

[1847] Audio data collection

[1848] Input: Sound environment collected by a microphone installed on the device

[1849] Processing: The device collects surrounding sounds in real time and records them as audio data.

[1850] Output: The collected audio data is saved on the device.

[1851] Step 2:

[1852] Sending audio data

[1853] Input: Audio data collected in step 1

[1854] Processing: The device sends the voice data to the cloud server.

[1855] Output: The audio data is sent to the cloud server.

[1856] Step 3:

[1857] Audio analysis

[1858] Input: Audio data sent to the server

[1859] Processing: The server uses a speech analysis engine (such as the Google Speech-to-Text API) to convert the audio data into text, and then uses a generative AI model to analyze it to identify specific dangerous or suspicious sounds.

[1860] Output: Text data and analysis results are generated.

[1861] Step 4:

[1862] Emotion analysis

[1863] Input: Text data from step 3 and image data as needed

[1864] Processing: The server uses an emotion analysis engine (such as Amazon Rekognition) to analyze the emotional state of people around you. Specifically, it analyzes facial expressions and tone of voice to detect emotions such as stress and anger.

[1865] Output: The result of sentiment analysis is obtained.

[1866] Step 5:

[1867] Integration of analysis results

[1868] Input: Results of Step 3 and Step 4

[1869] Processing: The server integrates the results of the speech analysis and emotion analysis and determines the appropriate response for the user.

[1870] Output: As a result of the integration, information is generated that should be communicated to the user.

[1871] Step 6:

[1872] Sending notifications

[1873] Input: Notification information generated in step 5

[1874] Processing: The server uses a notification system (such as Firebase Cloud Messaging) to send a notification to the device.

[1875] Output: A notification will be displayed on the device.

[1876] Step 7:

[1877] Alert the user

[1878] Input: Notification information from step 6

[1879] Action: The device notifies the user of the notification using a visual alert and / or vibration, specifically by displaying a warning message on the device display and vibrating to get the user's attention.

[1880] Output: The user can view the notification and take action.

[1881] This allows the user to immediately detect dangerous or suspicious sounds and take necessary measures promptly.

[1882] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1883] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1884] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1885] [Fourth embodiment]

[1886] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1887] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1888] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1889] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1890] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1891] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1892] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1893] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1894] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1895] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1896] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1897] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1898] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1899] The present invention is a system for solving various problems that hearing-impaired and hard-of-hearing people encounter in their daily lives, and specifically includes a means for analyzing the sound environment in real time and recommending appropriate hearing aid settings, a means for converting audio data into text, a means for providing video content with sign language commentary, and a means for detecting and notifying specific sounds and words. Each function of this system is described in detail below.

[1900] Automatic hearing aid optimization

[1901] When a user puts on the wearable device and turns it on, the device automatically collects the surrounding sound environment using a microphone. The sound environment data is analyzed by the device to identify sound types such as speech, noise, and alarms. The analyzed data is sent to a server, which calculates the most suitable hearing aid settings. The calculated settings are sent back to the device and suggested to the user. If the user approves, the device automatically adjusts the hearing aid settings.

[1902] Utilizing voice recognition technology

[1903] When a user wears the device and activates the voice recognition function, the device continuously collects surrounding sounds. The collected voice data is sent to a server in real time, where a voice recognition engine analyzes the voice data and converts it into text. The converted text data is displayed on the device's display, allowing the user to visually check the content of conversations around them.

[1904] Barrier-free information provided

[1905] When a user uses a device in a station or public facility, the device collects public information such as announcements and guidance announcements. The collected information is sent to a server, which searches a database of video content with sign language commentary. When relevant video content is found, the server returns a link to the device and notifies the user. When the user selects the link, the video with sign language commentary is played on the device.

[1906] Personal Settings Function

[1907] The user can set specific sounds or words (such as their name or "danger") in advance, and the device will monitor the surrounding sounds in real time. When the set sound or word is detected, the corresponding data is sent to the server. After confirmation, the server sends a notification instruction to the device. The device will then alert the user using a visual alert or vibration.

[1908] Specific examples

[1909] 1. Examples of automatic hearing aid optimization:

[1910] When a user is talking with a friend at a coffee shop, the device recognizes the surrounding noise and the conversation and sends the data to the server. The server calculates the settings to reduce noise and emphasize the conversation and sends them back to the device. The device then displays a message to the user asking, "Would you like to change the settings to reduce noise and emphasize the conversation?" and the settings are applied if the user approves.

[1911] 2. Examples of voice recognition technology:

[1912] When a user is participating in a meeting at work, the device sends the contents of the meeting to a speech recognition engine, which converts the conversation into text and displays "About the next presentation" on the device display, allowing the user to easily understand the content.

[1913] 3. Examples of providing barrier-free information:

[1914] When the user is at a station, the device collects the station's departure announcements, and the server finds and sends a video with sign language explanations to the device. The user views the video with sign language explanations, which include the message "The next train departs at 10:25," and accurately understands the information.

[1915] 4. Examples of personalization features:

[1916] When a user is relaxing at home, a specific word, "danger," is set, and if that word is heard from outside, the device will vibrate to notify the user and display "Danger detected" on the screen to warn them.

[1917] In this way, the present invention provides an effective means for improving the quality of life of the deaf and hard of hearing.

[1918] The processing flow will be explained below.

[1919] Automatic hearing aid optimization

[1920] Step 1:

[1921] The device detects that the user is wearing a wearable device and has it turned on.

[1922] Step 2:

[1923] The device collects the surrounding sound environment in real time using its built-in microphone and generates sound environment data.

[1924] Step 3:

[1925] The device analyzes the sound environment data collected and identifies conversation sounds, noise, alarm sounds, etc.

[1926] Step 4:

[1927] The device sends the analysis results to the server.

[1928] Step 5:

[1929] The server analyzes the received sound environment data and calculates the optimal hearing aid settings.

[1930] Step 6:

[1931] The server returns the calculation results to the terminal.

[1932] Step 7:

[1933] The device will prompt the user, "There are new environment-based hearing aid settings. Would you like to apply them?"

[1934] Step 8:

[1935] The user selects "Yes" or "No."

[1936] Step 9:

[1937] If the user selects "Yes," the device will automatically change the hearing aid settings.

[1938] Step 10:

[1939] The device will notify the user that the new settings have been applied.

[1940] Utilizing voice recognition technology

[1941] Step 1:

[1942] The user puts on the device and enables the speech recognition feature.

[1943] Step 2:

[1944] The device continuously collects surrounding sounds using the built-in microphone.

[1945] Step 3:

[1946] The device sends the collected voice data to the server.

[1947] Step 4:

[1948] The server analyzes the received voice data and converts it into text using a voice recognition engine.

[1949] Step 5:

[1950] The server returns the converted text data to the terminal.

[1951] Step 6:

[1952] The device displays the received text on the display.

[1953] Step 7:

[1954] The device updates the text display in real time as new audio data is collected.

[1955] Barrier-free information provided

[1956] Step 1:

[1957] The user wears the device and enables the barrier-free information provision function.

[1958] Step 2:

[1959] The device uses a built-in microphone to collect public announcements and guidance broadcasts.

[1960] Step 3:

[1961] The device sends the collected voice data to the server.

[1962] Step 4:

[1963] The server analyzes the audio data and searches a database for relevant video content with sign language commentary.

[1964] Step 5:

[1965] The server sends a link to the appropriate video content back to the device.

[1966] Step 6:

[1967] The device notifies the user, "There is a related sign language explanation video. Would you like to play it?"

[1968] Step 7:

[1969] The user selects "Yes" or "No."

[1970] Step 8:

[1971] If the user selects "Yes," the device will open the video link and play the sign language explanation video.

[1972] Personal Settings Function

[1973] Step 1:

[1974] The user puts on the device and accesses the personalization screen.

[1975] Step 2:

[1976] The user sets and saves a specific sound or word (such as a name or "danger").

[1977] Step 3:

[1978] The device collects surrounding sounds in real time using a built-in microphone.

[1979] Step 4:

[1980] The device continuously compares the sounds it collects with preset sounds and words.

[1981] Step 5:

[1982] When the device detects a specific sound or word, it sends that data to the server.

[1983] Step 6:

[1984] The server analyzes the received data and confirms the detection results.

[1985] Step 7:

[1986] The server sends a notification instruction to the terminal.

[1987] Step 8:

[1988] The device will notify the user with a visual alert and / or vibration.

[1989] Step 9:

[1990] The device will display "A specific sound has been detected. Please be careful."

[1991] Example 1

[1992] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1993] This invention aims to solve various problems that the hearing impaired and hard of hearing face in their daily lives. Specifically, it aims to improve the quality of life for the hearing impaired by providing technology that efficiently analyzes surrounding sounds and automatically configures optimal hearing aids, as well as technology that converts speech to text, provides video content that incorporates sign language, and detects and notifies users of specific sounds and words.

[1994] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1995] In this invention, the server includes a means for analyzing the sound environment in real time and recommending appropriate hearing aid settings, a means for converting recorded audio data into text, and a means for providing video content with sign language commentary, thereby enabling hearing-impaired or hard-of-hearing people to efficiently grasp the surrounding sound environment and receive appropriate hearing assistance.

[1996] "Sound environment" refers to the state and situation of sound waves such as surrounding voices and noise.

[1997] "Analyzing in real time" means analyzing data immediately after it is collected, without any time delay.

[1998] A "hearing assistive device" is a device that assists the hearing of people with hearing loss or hearing impairments, and has the function of amplifying sound or emphasizing specific sounds.

[1999] "Audio recording data" refers to data that has been digitized and saved from audio collected by a microphone or other device.

[2000] "Converting to text" refers to the process of converting data such as audio and video into text information.

[2001] "Video content with sign language commentary" refers to videos or video materials that include visual sign language commentary.

[2002] "Detecting specific sounds or words" refers to the recognition system identifying and detecting predefined sounds or words.

[2003] "Notify" refers to informing a user of specific information or events.

[2004] "Device" refers to electronic equipment or devices, including hearing aids and wearable devices.

[2005] "Central server" refers to the main computer system that receives data from multiple devices, processes and analyzes it, and returns the results.

[2006] "Collect" refers to gathering data or information.

[2007] "Transmit" refers to sending data or information to another device or system.

[2008] "Calculating" refers to the process of deriving a specific result based on analytical results or algorithms.

[2009] "Automatically adjust" means that the device is automatically set to the optimum state without requiring any user operation.

[2010] "Public facilities" refers to buildings and places for use by the general public, including stations and libraries.

[2011] "Linking" means associating and making accessible particular information or content.

[2012] "Monitor" refers to watching for a particular condition or state.

[2013] This system analyzes the sound environment, converts speech into text, provides video content with sign language commentary, and detects and notifies specific sounds and words to solve various problems faced by the hearing impaired and hard of hearing in their daily lives. The system of this invention functions mainly through cooperation between the device (wearable terminal) and a central server.

[2014] Hardware and software used

[2015] 1. Device (wearable device)

[2016] microphone

[2017] Built-in processor

[2018] Communication modules (Wi-Fi, 5G, etc.)

[2019] display

[2020] Vibration Motor

[2021] battery

[2022] 2. Central Server

[2023] Generative AI Models

[2024] Speech Recognition Engine

[2025] Database (including video content with sign language commentary)

[2026] Communication Interface

[2027] Program processing

[2028] Analysis of sound environments and optimization of hearing aids

[2029] When a user puts on the device and turns it on, the device automatically collects the surrounding sound environment using a microphone. The sound environment data is analyzed by the device to identify sound types such as speech, noise, and alarms. This data is sent to a central server, where a generative AI model calculates optimal hearing aid settings. The calculation results are sent back to the device and recommended to the user. If the user approves, the device automatically adjusts the hearing aid settings.

[2030] Examples:

[2031] When a user is talking with a friend at a coffee shop, the device recognizes the surrounding noise and the conversation and sends the data to a central server. The server calculates the settings to reduce noise and enhance the conversation and sends them back to the device. The device then displays a message asking, "Do you want to change the settings to reduce noise and enhance the conversation?" and the settings are applied if the user confirms.

[2032] Utilizing voice recognition technology

[2033] When a user enables the device's voice recognition function, the device will continuously collect surrounding sounds. The collected voice data is sent to a central server in real time, where the voice recognition engine converts the speech into text. The converted text data is then displayed on the device's display, allowing the user to visually check the content of the conversations around them.

[2034] Examples:

[2035] When a user is participating in a meeting at work, the device sends the contents of the meeting to a speech recognition engine, and the converted text is displayed on the device's display as "About the next presentation," allowing the user to easily understand the content.

[2036] Providing video content with sign language commentary

[2037] When a user uses a device in a station or public facility, the device collects public information such as announcements and guidance announcements. The collected information is sent to a central server, which searches a database of videos with sign language descriptions. When relevant video content is found, the central server sends a link back to the device and notifies the user. When the user selects the link, the video with sign language descriptions is played on the device.

[2038] Examples:

[2039] When a user is at a station, the terminal collects the station's departure announcements, and the central server finds the video with sign language explanations and sends it to the terminal. The user views the video with sign language explanations, which include the message "The next train departs at 10:25," and accurately understands the information.

[2040] Detecting and announcing specific sounds and words

[2041] Users can pre-set specific sounds or words (e.g., "danger"), and the device will monitor surrounding sounds in real time. When the device detects a specific sound or word, the corresponding data is sent to a central server, which then confirms the detection and sends a notification instruction to the device. The device will then alert the user with a visual alert or vibration.

[2042] Examples:

[2043] If the user is relaxing at home and the word "danger" is set and hears that word from outside, the device will vibrate to notify the user and display "Danger detected" on the screen to warn them.

[2044] As described above, the present invention integrates multiple technical means to provide an effective method for the hearing impaired and hard of hearing to efficiently obtain information necessary in daily life.

[2045] The flow of the identification process in the first embodiment will be described with reference to FIG.

[2046] Analysis of sound environments and optimization of hearing aids

[2047] Step 1:

[2048] When a user wears the device and turns it on, the device automatically collects the surrounding sound environment with its microphone. The input here is the surrounding sound environment, and the output is digitized sound environment data. The device's microphone converts analog audio into digital data.

[2049] Step 2:

[2050] The device analyzes the sound environment data collected in real time to identify the type of sound, such as speech, noise, or alarm sounds. The input is digitized sound environment data, and the output is information about the identified sound type (voice pattern data). The device's built-in processor runs an algorithm to analyze the data.

[2051] Step 3:

[2052] The device sends the analyzed voice pattern data to a central server. The input is the voice pattern data, and the output is the transmitted data. The data is sent to a cloud server using the device's communication module (Wi-Fi, 5G, etc.).

[2053] Step 4:

[2054] Based on the voice pattern data received by the server, a generative AI model calculates the optimal hearing aid settings. The input is voice pattern data, and the output is optimal setting data. The generative AI model analyzes the voice data and derives the appropriate parameters.

[2055] Step 5:

[2056] The server returns the optimal setting data to the terminal. The input is the optimal setting data, and the output is the setting data sent to the terminal. The data is sent using the server's communication interface.

[2057] Step 6:

[2058] The device receives the optimal setting data and displays the suggestion to the user. The input is the received setting data, and the output is the display of the suggestion. The device display shows "Would you like to change the settings to reduce noise and enhance speech?"

[2059] Step 7:

[2060] When the user presses the approval button, the device automatically adjusts the settings of the hearing aid. The input is the user's approval, and the output is the adjusted settings of the hearing aid. The device provides feedback to the user by vibration or sound.

[2061] Utilizing voice recognition technology

[2062] Step 1:

[2063] When a user enables the voice recognition function of the device, the device continuously collects surrounding sounds. The input is the surrounding sounds, and the output is digitized voice data. The microphone collects the voice and converts it into digital data.

[2064] Step 2:

[2065] The terminal transmits the collected voice data to a central server in real time. The input is the digitized voice data, and the output is the transmitted voice data. The data is transmitted using a communication module.

[2066] Step 3:

[2067] The server analyzes the voice data using a voice recognition engine and converts it into text data. The input is voice data and the output is text data. The voice recognition engine analyzes the voice pattern and generates the corresponding text.

[2068] Step 4:

[2069] Text data is sent from the central server to the terminal. The input is the text data, and the output is the text data sent to the terminal. The data is sent using the server's communication interface.

[2070] Step 5:

[2071] Text data is displayed on the terminal display. The input is the received text data, and the output is the text content displayed on the display. The user can check the content of the surrounding conversation by looking at the displayed text.

[2072] Providing video content with sign language commentary

[2073] Step 1:

[2074] When a user uses a device in a station or public facility, the terminal collects announcements and guidance announcements. The input is the public facility announcements, and the output is digitized voice data. The microphone collects the voice and converts it into digital data.

[2075] Step 2:

[2076] The collected voice data is transmitted to a central server. The input is the digitized voice data and the output is the transmitted voice data. A communication module is used to transmit the data to the server.

[2077] Step 3:

[2078] The server analyzes the audio data and searches a database of video content with sign language commentary. The input is the audio data, and the output is a link to related video content. The server's generative AI model analyzes the audio and finds related videos from the database.

[2079] Step 4:

[2080] The server returns a link to the associated video content to the terminal. The input is the link to the video content, and the output is the link sent to the terminal. The link is sent using a communication interface.

[2081] Step 5:

[2082] When the user selects a link, the device plays a video with sign language descriptions. The input is the link selected by the user and the output is the video that is played. The video with sign language descriptions appears on the device's display.

[2083] Detecting and announcing specific sounds and words

[2084] Step 1:

[2085] The user sets specific sounds or words (e.g., "danger") in advance on the device. The input is the user's settings for the sounds or words, and the output is the saved setting data. The sounds or words are entered using the device's setting screen and saved in a database.

[2086] Step 2:

[2087] The device monitors the surrounding sound in real time. The input is the surrounding sound, and the output is digitized audio data. The microphone collects the audio and converts it into digital data.

[2088] Step 3:

[2089] When a preset sound or word is detected, the device sends data to a central server. The input is the detected voice data, and the output is the transmitted data. The data is sent to the server using a communication module.

[2090] Step 4:

[2091] The server checks the received data and sends a notification instruction to the device. The input is the received voice data, and the output is the notification instruction. The server's generative AI model analyzes the data and determines the appropriate response.

[2092] Step 5:

[2093] The device alerts the user with a visual alert or vibration. The input is a notification instruction, and the output is a visual alert or vibration notification. The device activates the vibration motor and displays "Danger detected" on the display.

[2094] (Application example 1)

[2095] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2096] The present invention aims to solve the problem that hearing-impaired and hard-of-hearing people have difficulty recognizing voices and grasping important information in daily life and in certain situations (e.g., when using self-driving vehicles). In particular, there is a problem that emergency situations and important announcements are difficult to convey to hearing-impaired and hard-of-hearing people. This may result in situations where the safety and comfort of users are lacking.

[2097] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[2098] In this invention, the server includes a means for analyzing the sound environment in real time and recommending appropriate hearing aid settings, a means for converting voice data into text, and a means for providing video content with sign language commentary. This allows hearing-impaired and hard-of-hearing individuals to properly understand the surrounding sound environment and automatically optimize their hearing aid settings. Furthermore, by providing a means for collecting voice data in real time and notifying the user based on specific keywords, it becomes possible to reliably communicate emergencies and important announcements. Furthermore, by including a means for displaying voice recognition results on an in-vehicle display, users can visually confirm surrounding voice information, thereby improving safety and comfort when using autonomous vehicles.

[2099] "Analyzing the sound environment in real time and recommending appropriate hearing aid settings" means continuously monitoring the sound environment using sensors and microphones, classifying the type and level of sound using an analysis device, and providing optimal hearing aid settings based on that.

[2100] "Converting voice data to text" refers to the process of analyzing collected voice data using speech recognition technology and converting it into a corresponding text format.

[2101] "Providing video content with sign language commentary" is a service that allows deaf and hard of hearing people to obtain information visually by providing video information with sign language interpretation.

[2102] "Detecting and notifying specific sounds and words" refers to a mechanism that detects specific pre-set sounds or utterances and notifies the user.

[2103] "Collecting voice data in real time and notifying the user based on specific keywords" means accumulating voice data in real time, identifying specific keywords in the data, and providing the user with appropriate alerts.

[2104] "Displaying voice recognition results on an in-vehicle display" means displaying text data obtained through voice recognition technology on a display inside an autonomous vehicle to provide information visually to the user.

[2105] The present invention is a system for solving various problems that hearing-impaired and hard-of-hearing people encounter in their daily lives and when using self-driving vehicles. This system includes a means for analyzing the sound environment in real time and recommending appropriate hearing aid settings, a means for converting voice data into text, a means for providing video content with sign language commentary, and a means for detecting and notifying specific sounds and words.

[2106] System configuration

[2107] The system consists of the following main components:

[2108] Hardware

[2109] 1. Microphone: A device that collects ambient sounds, allowing for real-time monitoring of the sound environment.

[2110] 2. Server: Analyzes the sound data and provides instructions for optimal settings for the hearing aid and text conversion.

[2111] 3. Autonomous vehicle display: A device for displaying voice recognition results and system notifications.

[2112] 4. Vibration motor: To notify the user when a specific sound or keyword is detected.

[2113] software

[2114] 1. Speech recognition engine (e.g., Google Speech Recognition API): Collects speech in real time and converts it into text.

[2115] 2. Analysis algorithm: Analyzes the sound environment and recommends appropriate hearing aid settings.

[2116] 3. Notification system: Detects specific keywords or voice anomalies and notifies the user.

[2117] System Operation

[2118] 1. Real-time analysis of sound environments

[2119] When the user operates the system, the microphone begins to collect the surrounding sound environment. The collected sound environment data is sent to the server, where an analysis algorithm identifies the type of sound, such as speech, noise, or alarm. Based on the analyzed data, the server calculates the optimal hearing aid settings and sends them back to the device. If the user approves, the device automatically adjusts the hearing aid settings.

[2120] 2. Use of voice recognition technology

[2121] When the user wears the device and activates the voice recognition function, the device continuously collects surrounding sounds. The collected voice data is sent to a server in real time, where a voice recognition engine analyzes the voice data and converts it into text. The converted text data is displayed on the vehicle's display, allowing the user to visually check the content of surrounding conversations and announcements.

[2122] 3. Providing barrier-free information

[2123] When a user is using an autonomous vehicle, the device collects information such as in-car announcements and emergency alerts. The collected information is sent to a server, which searches a database of video content with sign language commentary. When relevant video content is found, the server returns a link to the device and notifies the user. When the user selects the link, the video with sign language commentary is played on the in-car display.

[2124] 4. Personal settings function

[2125] The user can set specific sounds or words (such as their name or "danger") in advance, and the device will monitor the surrounding sounds in real time. When the set sound or word is detected, the corresponding data is sent to the server. After confirmation, the server sends a notification instruction to the device. The device will then alert the user using a visual alert or vibration.

[2126] Specific examples

[2127] 1. A concrete example of automatic hearing aid optimization

[2128] When a user is talking with a friend at a coffee shop, the device recognizes the surrounding noise and the conversation and sends the data to the server. The server calculates the settings to reduce noise and emphasize the conversation and sends them back to the device. The device then displays a message to the user asking, "Would you like to change the settings to reduce noise and emphasize the conversation?" and the settings are applied if the user approves.

[2129] 2. Specific examples of voice recognition technology

[2130] When a user is participating in a meeting at work, the device sends the contents of the meeting to a speech recognition engine, which converts the conversation into text and displays "About the next presentation" on the device display, allowing the user to easily understand the content.

[2131] 3. Examples of providing barrier-free information

[2132] When the user is at a station, the device collects the station's departure announcements, and the server finds and sends a video with sign language explanations to the device. The user views the video with sign language explanations, which include the message "The next train departs at 10:25," and accurately understands the information.

[2133] 4. Examples of personal settings functions

[2134] When a user is relaxing at home, a specific word, "danger," is set, and if that word is heard from outside, the device will vibrate to notify the user and display "Danger detected" on the screen to warn them.

[2135] Example prompts for generative AI models

[2136] "You are designing a self-driving car system for the hearing impaired. The system uses speech recognition to convert emergency situations and announcements into text and notify the user. Extract the emergency content from the following audio data and display it as text."

[2137] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[2138] Step 1:

[2139] The user operates the system. The user gets into the autonomous vehicle and a microphone connected to the terminal collects surrounding sounds in real time. The collected sound data is input into the terminal.

[2140] Step 2:

[2141] The device analyzes the collected sound data and classifies the sound environment. The device's internal analysis algorithm classifies the sound data into categories such as speech, noise, and alarm sounds. This analyzed data is sent to the server.

[2142] Step 3:

[2143] The server receives the sound environment data and calculates the optimal hearing aid settings. Based on the analyzed sound data, calculations are performed to automatically optimize the hearing aid settings. These optimized settings are then sent from the server to the device.

[2144] Step 4:

[2145] The device presents the hearing aid settings to the user. The device asks the user whether to apply the new settings and obtains their approval. When the user presses the approval button, the device changes the hearing aid settings.

[2146] Step 5:

[2147] The device continuously collects voice data and sends it to a voice recognition engine. Based on the voice data input, real-time voice recognition processing is performed and the data is converted into text. This text data is then sent back to the device.

[2148] Step 6:

[2149] The device displays the voice recognition results on the vehicle display, and text data is output to the display so that the user can visually check the content of surrounding conversations and announcements.

[2150] Step 7:

[2151] The device monitors specific keywords and sounds to detect abnormalities. When a pre-defined keyword (e.g., "danger" or "emergency") is detected, the data is sent to the server. After the server confirms the detection, it sends a notification instruction to the device.

[2152] Step 8:

[2153] The device will alert the user. If an abnormal sound or specific keyword is detected, the device will alert the user with a visual alert or vibration notification.

[2154] The above process will create a system that improves the safety and convenience of people who are deaf or hard of hearing when using self-driving vehicles.

[2155] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[2156] The present invention is a system for solving various problems encountered in daily life by people with hearing impairments or hard of hearing. Specifically, it includes a means for analyzing the sound environment in real time and recommending appropriate hearing aid settings, a means for converting audio data into text, a means for providing video content with sign language commentary, and a means for detecting and notifying specific sounds or words. Furthermore, by combining it with an emotion engine that recognizes and analyzes the user's emotions, further user assistance is provided. Each function of this system is described in detail below.

[2157] Automatic hearing aid optimization

[2158] When a user puts on the wearable device and turns it on, the device automatically collects the surrounding sound environment using a microphone. The sound environment data is analyzed by the device to identify sound types such as speech, noise, and alarms. The analyzed data is sent to a server, which calculates the most suitable hearing aid settings. The calculated settings are sent back to the device and suggested to the user. If the user approves, the device automatically adjusts the hearing aid settings.

[2159] Utilizing voice recognition technology

[2160] When a user wears the device and activates the voice recognition function, the device continuously collects surrounding sounds. The collected voice data is sent to a server in real time, where a voice recognition engine analyzes the voice data and converts it into text. The converted text data is displayed on the device's display, allowing the user to visually check the content of conversations around them.

[2161] Barrier-free information provided

[2162] When a user uses a device in a station or public facility, the device collects public information such as announcements and guidance announcements. The collected information is sent to a server, which searches a database of video content with sign language commentary. When relevant video content is found, the server returns a link to the device and notifies the user. When the user selects the link, the video with sign language commentary is played on the device.

[2163] Personal Settings Function

[2164] The user can set specific sounds or words (such as their name or "danger") in advance, and the device will monitor the surrounding sounds in real time. When the set sound or word is detected, the corresponding data is sent to the server. After confirmation, the server sends a notification instruction to the device. The device will then alert the user using a visual alert or vibration.

[2165] Utilizing the Emotion Engine

[2166] The emotion engine monitors the user's emotional state in real time while the user is wearing the device. The emotion engine analyzes the user's tone of voice and facial expressions (e.g., using the camera function) to recognize emotions such as stress, joy, and sadness. Once the emotional state is analyzed, the results are sent to the server, which determines the appropriate response and sends instructions back to the device.

[2167] Specific examples

[2168] 1. Examples of automatic hearing aid optimization:

[2169] When a user is talking with a friend at a coffee shop, the device recognizes the surrounding noise and the conversation and sends the data to the server. The server calculates the settings to reduce noise and emphasize the conversation and sends them back to the device. The device then displays a message to the user asking, "Would you like to change the settings to reduce noise and emphasize the conversation?" and the settings are applied if the user approves.

[2170] 2. Examples of voice recognition technology:

[2171] When a user is participating in a meeting at work, the device sends the contents of the meeting to a speech recognition engine, which converts the conversation into text and displays "About the next presentation" on the device display, allowing the user to easily understand the content.

[2172] 3. Examples of providing barrier-free information:

[2173] When the user is at a station, the device collects the station's departure announcements, and the server finds and sends a video with sign language explanations to the device. The user views the video with sign language explanations, which include the message "The next train departs at 10:25," and accurately understands the information.

[2174] 4. Examples of personalization features:

[2175] When a user is relaxing at home, a specific word, "danger," is set, and if that word is heard from outside, the device will vibrate to notify the user and display "Danger detected" on the screen to warn them.

[2176] 5. Examples of Emotion Engines:

[2177] If a user feels stressed during a meeting, the emotion engine analyzes the user's facial expressions and tone of voice to detect the stress level. The server recommends audio content with a relaxing effect, and the device notifies the user, "Would you like to play relaxing music?" and plays the music if the user agrees.

[2178] In this way, the present invention provides an effective means for improving the quality of life of the deaf and hard of hearing.

[2179] The processing flow will be explained below.

[2180] Automatic hearing aid optimization

[2181] Step 1:

[2182] The user puts on the wearable device and turns it on.

[2183] Step 2:

[2184] The device automatically collects the surrounding sound environment in real time using the built-in microphone.

[2185] Step 3:

[2186] The device analyzes the sound environment data collected and identifies the type of sound (conversation, noise, alarm, etc.).

[2187] Step 4:

[2188] The device sends the analysis results to the server.

[2189] Step 5:

[2190] The server analyzes the received sound environment data and calculates appropriate hearing aid settings.

[2191] Step 6:

[2192] The server returns the calculated hearing aid settings to the device.

[2193] Step 7:

[2194] The device will prompt the user, "There are new environment-based hearing aid settings. Would you like to apply them?"

[2195] Step 8:

[2196] The user selects "Yes" or "No."

[2197] Step 9:

[2198] If the user selects "Yes," the device will automatically change the hearing aid settings.

[2199] Step 10:

[2200] The device will notify the user that the new settings have been applied.

[2201] Utilizing voice recognition technology

[2202] Step 1:

[2203] The user puts on the device and enables the speech recognition feature.

[2204] Step 2:

[2205] The device continuously collects surrounding sounds using the built-in microphone.

[2206] Step 3:

[2207] The device sends the collected voice data to the server.

[2208] Step 4:

[2209] The voice data received by the server is analyzed using a voice recognition engine and converted into text.

[2210] Step 5:

[2211] The server returns the converted text data to the terminal.

[2212] Step 6:

[2213] The text data received by the terminal is displayed on the display.

[2214] Step 7:

[2215] The device continues to convert new voice data collected into text in real time, updating the text on the display.

[2216] Barrier-free information provided

[2217] Step 1:

[2218] The user wears the device and enables the barrier-free information provision function.

[2219] Step 2:

[2220] The device uses a built-in microphone to collect public announcements and guidance broadcasts.

[2221] Step 3:

[2222] The device sends the collected voice data to the server.

[2223] Step 4:

[2224] The server analyzes the audio data and searches a database for relevant video content with sign language commentary.

[2225] Step 5:

[2226] The server sends a link to the appropriate video content back to the device.

[2227] Step 6:

[2228] The device notifies the user, "There is a related sign language explanation video. Would you like to play it?"

[2229] Step 7:

[2230] The user selects "Yes" or "No."

[2231] Step 8:

[2232] If the user selects "Yes," the device will open the video link and play the sign language explanation video.

[2233] Personal Settings Function

[2234] Step 1:

[2235] The user puts on the device and accesses the personalization screen.

[2236] Step 2:

[2237] The user sets and saves a specific sound or word (such as a name or "danger").

[2238] Step 3:

[2239] The device collects surrounding sounds in real time using a built-in microphone.

[2240] Step 4:

[2241] The device continuously compares the sounds it collects with preset sounds and words.

[2242] Step 5:

[2243] When the device detects a specific sound or word, it sends that data to the server.

[2244] Step 6:

[2245] The server analyzes the received data and confirms the detection results.

[2246] Step 7:

[2247] The server sends a notification instruction to the terminal.

[2248] Step 8:

[2249] The device will notify the user with a visual alert and / or vibration.

[2250] Step 9:

[2251] The device will display "A specific sound has been detected. Please be careful."

[2252] Utilizing the Emotion Engine

[2253] Step 1:

[2254] The user puts on the device and activates the emotion engine function.

[2255] Step 2:

[2256] The device collects the user's voice tone and facial expressions using the built-in microphone and camera.

[2257] Step 3:

[2258] The device sends the collected voice tone and facial expression data to the server.

[2259] Step 4:

[2260] The server analyzes the received data using an emotion engine to identify the user's emotional state.

[2261] Step 5:

[2262] The server returns the analysis results to the device and instructs it on the appropriate response.

[2263] Step 6:

[2264] The device will notify the user, for example, "Emotional state detected. Would you like to relax?"

[2265] Step 7:

[2266] The user selects "Yes" or "No."

[2267] Step 8:

[2268] If the user selects "Yes," the terminal plays back audio content that has a relaxing effect.

[2269] Step 9:

[2270] The device again monitors changes in the user's emotional state and makes adjustments as needed.

[2271] Example 2

[2272] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2273] Among the challenges faced by people with hearing impairments and those with hearing loss in their daily lives are adjusting their hearing aid settings to changes in the surrounding sound environment, understanding conversations and important announcements, obtaining information in public facilities, and recognizing emergency situations in real time. Effectively resolving these challenges requires a system that can collect and analyze various voice and environmental data in real time and provide appropriate support. Previously, no system offered all of these functions comprehensively, forcing users to use multiple devices and services at the same time, which was inconvenient.

[2274] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[2275] In this invention, the server includes a means for analyzing the sound environment in real time and suggesting optimal hearing aid settings, a means for converting voice data into text, a means for providing visual content with sign language explanations, a means for detecting specific sounds and words and issuing a warning, and a means for analyzing the user's emotional state and suggesting appropriate responses. This allows users to receive a variety of support from a single system, comprehensively resolving various issues in daily life.

[2276] "Analyzing the sound environment in real time and proposing optimal hearing aid settings" means collecting surrounding sounds using a microphone or other device, analyzing the collected sound data, and calculating and providing hearing aid settings that optimize the way the user hears sound.

[2277] "Converting voice data to text" means collecting surrounding voices as digital signals, analyzing the voice data, converting it into text information, and displaying it.

[2278] "Providing visual content with sign language explanations" means providing users with videos or animations with added sign language explanations to supplement visual means of communication.

[2279] "Detecting specific sounds or words and issuing a warning" means recognizing specific sounds or words that have been set in advance, and issuing a notification or warning to the user when they are detected.

[2280] "Analyzing the user's emotional state and suggesting appropriate responses" means analyzing the user's tone of voice, facial expressions, etc. to recognize their emotional state, and then suggesting relaxation content or other appropriate support based on the results.

[2281] The present invention is a comprehensive system for solving various problems faced by the deaf and hard of hearing in daily life. The system includes means for analyzing the sound environment in real time and suggesting optimal hearing aid settings, means for converting speech data into text, means for providing visual content with sign language explanations, means for detecting specific sounds and words and issuing warnings, and means for analyzing the user's emotional stat...

Claims

1. A means to analyze the sound environment in real time and recommend appropriate hearing aid settings; a means for converting the audio data into text; A means of providing video content with sign language commentary; a means for detecting and announcing specific sounds or words; A system including:

2. 2. The system of claim 1, wherein the system transmits sound environment data to a server and receives hearing aid settings from the server.

3. 2. The system according to claim 1, wherein the collected voice data is sent to a voice recognition engine and returned to the terminal as text data.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A