System
A system that converts ambient sounds and sign language into text/audio and provides audio guidance addresses the challenges faced by the hearing and visually impaired, enhancing their understanding and safety in daily life.
Patent Information
- Application Number
- JP2024117320
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-22
- Publication Date
- 2026-02-03
AI Technical Summary
People with hearing and visual impairments face difficulties in detecting important information and navigating their surroundings, leading to safety and communication challenges.
A system that captures ambient sounds and sign language, converts them into text or audio, and provides audio guidance based on environmental data analysis, enabling real-time understanding of surroundings for the hearing and visually impaired.
Enhances the ability of hearing and visually impaired individuals to understand their environment and communicate effectively, promoting safety and independent living.
Smart Images

Figure 2026016230000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] There is a need for technology to alleviate the difficulties faced by people with hearing and visual impairments in their daily lives. Specifically, people with hearing impairments cannot hear surrounding sounds or conversations, making it difficult for them to detect important information or dangers, and this makes safety and communication difficult. On the other hand, people with visual impairments have difficulty obtaining visual information beyond their physical reach, which increases the risk of mobility and understanding daily situations. Therefore, there is a need for systems that solve these issues and support the safety and independent living of people with disabilities. [Means for solving the problem]
[0005] In order to solve the above problems, the present invention provides the following means. First, for the hearing impaired, a system is provided that includes means for capturing ambient sounds, means for converting the captured audio data into text data, means for analyzing the type, direction, and distance of the sound, and means for displaying the analyzed information. It also includes means for capturing sign language, means for converting the captured sign language data into text data, and means for audibly converting the converted text data. Furthermore, for the visually impaired, a system is provided that includes means for capturing ambient environmental data, means for analyzing the captured environmental data, means for generating the analyzed information as audio guidance, and means for outputting the generated audio guidance. This achieves the effect of making it easier for the hearing impaired to understand ambient sounds and sign language, and easier for the visually impaired to grasp the surrounding situation.
[0006] A "means" is an element of a device, method, or system used to accomplish some purpose.
[0007] A "sound capturing means" is a microphone or other audio input device used to collect ambient sounds.
[0008] "Audio data" is a digital representation of captured sound.
[0009] "Means for converting into text data" refers to speech recognition software or algorithms for converting voice data into text information.
[0010] "Means for analyzing type, direction and distance of sound" is a processing device for analyzing the captured audio data to determine what the sound is, what direction it is coming from, and its distance.
[0011] "Means for displaying information" refers to a display or screen that visually conveys the analyzed information to the user.
[0012] "Means for capturing sign language" refers to cameras or sensors that capture the user's sign language movements and collect them as data.
[0013] "Sign language data" means a digital representation of captured sign language movements.
[0014] "Means for converting into text data (sign language)" refers to sign language recognition software or algorithms for converting sign language data into text information.
[0015] The "means for converting text data into voice" is a voice synthesis engine for converting text data into voice and playing it back.
[0016] "Means for capturing environmental data" refers to cameras and sensors used to collect information about the surrounding environment.
[0017] "Environmental data" is data that digitally represents captured information about the surrounding environment.
[0018] "Means for analyzing environmental data" refers to AI algorithms and processing devices that analyze captured environmental data and recognize objects and situations.
[0019] The "means for generating audio guidance" is a voice synthesis engine for generating the analyzed environmental information as an audio message.
[0020] The "means for outputting audio guidance" refers to speakers or earphones that audibly notify the user of the generated audio guidance. [Brief explanation of the drawings]
[0021] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0022] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0023] First, the terms used in the following description will be explained.
[0024] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0025] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0026] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0027] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0028] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0029] [First embodiment]
[0030] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0031] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0032] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0033] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0034] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0035] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0036] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0037] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0038] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0039] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0040] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0041] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0042] The present invention is an assistance system for the hearing impaired and the visually impaired, and is implemented in the following manner.
[0043] System configuration
[0044] This system consists of a user's mobile device (hereafter referred to as the device) and a cloud-based server (hereafter referred to as the server). The device is equipped with a microphone, camera, speaker, display, and various sensors, and uses this hardware to collect and analyze information about the user's surroundings and provide the necessary information.
[0045] Hearing-impaired features
[0046] Subtitling of ambient sounds
[0047] 1. Sound capture and processing
[0048] The device uses a microphone to constantly capture ambient sounds.
[0049] The captured audio data undergoes pre-processing such as noise reduction inside the device.
[0050] The preprocessed voice data is converted into text in real time by a voice recognition engine.
[0051] 2. Analysis of sound type, direction, and distance
[0052] The device identifies the type of sound from the converted text information and sound characteristics, such as "ambulance siren" or "human voice."
[0053] The device uses multiple microphone arrangements to calculate the direction and distance of the sound and notify the user.
[0054] 3. Displaying Information
[0055] The analyzed information is displayed in text format on the device's display, allowing the user to understand the surrounding situation in real time.
[0056] 4. Specific Examples
[0057] For example, if a car horn sounds near the user, the device will display "The horn is sounding from 10 meters to the right."
[0058] Automatic sign language recognition and translation
[0059] 1. Sign Language Capture and Processing
[0060] The device uses a camera to capture the user's sign language in real time.
[0061] The captured video data is pre-processed to extract the shape and movement of the hand.
[0062] 2. Sign Language Recognition and Translation
[0063] The pre-processed video data is analyzed by a sign language recognition algorithm to generate corresponding text information.
[0064] The generated text information is converted into voice data by a voice synthesis engine.
[0065] 3. Audio Output
[0066] The device plays the translated audio through a speaker, helping people who don't understand sign language to communicate with the user.
[0067] 4. Specific Examples
[0068] For example, if a user says "hello" in sign language, the terminal recognizes the sign and outputs "hello" aloud.
[0069] Features for the visually impaired
[0070] Audio guidance of surrounding objects and situations
[0071] 1. Environmental data capture and processing
[0072] The device uses a camera and various sensors to capture data about the surrounding environment.
[0073] The captured data is subjected to image recognition preprocessing to extract the contours and feature points of the object.
[0074] 2. Object Recognition and Analysis
[0075] The preprocessed data is then used by AI algorithms to recognize surrounding objects and situations.
[0076] The location and movement of recognized objects are analyzed to pick out important information.
[0077] 3. Audio guide generation and output
[0078] The analyzed information is generated as audio guide and converted into voice by a voice synthesis engine.
[0079] The terminal notifies the user of the generated audio guide through a speaker.
[0080] 4. Specific Examples
[0081] If the user is walking and there is an obstacle one meter ahead, the device will announce with a voice message, "There is an obstacle one meter ahead."
[0082] Starting and shutting down the system
[0083] 1. Initial Setup
[0084] During initial setup, the user selects the system mode (hearing impaired mode, visually impaired mode) that best suits their needs.
[0085] The server receives the user's configuration information and prepares the required services.
[0086] 2. Service launch
[0087] The device will activate the necessary sensors (microphone, camera, etc.) based on the initial settings and begin collecting data.
[0088] The server analyzes the data sent from the terminal and returns the necessary information in real time.
[0089] The terminal displays or outputs the analysis results to the user.
[0090] 3. Termination or Suspension of Service
[0091] If a user wishes to stop using the service, they can select the option to terminate or suspend the service in their device's settings screen.
[0092] The terminal receives the service termination command, shuts down all sensors, and cuts off communication with the server.
[0093] This system provides useful support to the hearing-impaired and visually impaired to overcome difficulties in daily life and live safely and independently.
[0094] The processing flow will be explained below.
[0095] Features for the hearing impaired: Subtitling of surrounding sounds
[0096] Step 1: Capture the sound
[0097] The device activates the microphone and captures ambient sounds at 0.1 second intervals.
[0098] Step 2: Preprocessing the audio data
[0099] The device performs noise reduction and normalization on the captured audio data.
[0100] Step 3: Speech to text
[0101] The device sends the preprocessed voice data to a voice recognition engine, which converts it into text data in real time.
[0102] Step 4: Identify the type of sound
[0103] The device extracts sound characteristics from the text data and identifies the type of sound based on them (e.g., "ambulance siren," "human voice," etc.).
[0104] Step 5: Analyze sound direction and distance
[0105] The device uses multiple microphone arrangements to calculate the direction and distance from which the sound is coming.
[0106] Step 6: Display subtitles
[0107] The device displays text information on the display, informing the user of the type, direction, and distance of the sound.
[0108] As a specific example, it displays "The sound of a horn is coming from 10 meters to the right."
[0109] Features for the visually impaired: Audio guide to surrounding objects and situations
[0110] Step 1: Capture environment data
[0111] The device will activate its camera and sensors to capture data on the surrounding environment every 0.5 seconds.
[0112] Step 2: Preprocessing the data
[0113] The device performs image recognition preprocessing on the captured video data to extract the contours and feature points of objects.
[0114] Step 3: Object Recognition
[0115] The device sends the pre-processed data to AI algorithms to recognize surrounding objects.
[0116] Step 4: Obtaining information about the object
[0117] The device analyzes the location and movement of recognized objects and picks out important information (e.g., obstacles ahead, approaching people).
[0118] Step 5: Generate audio guide
[0119] The device generates the analyzed information as a text message, sends it to a speech synthesis engine, and generates audio guidance.
[0120] Step 6: Audio Notifications
[0121] The terminal notifies the user of the generated audio guide through a speaker.
[0122] As a specific example, a voice message will be displayed saying, "There is an obstacle one meter ahead."
[0123] Starting and shutting down the system
[0124] Step 1: Initial Setup
[0125] The user configures the system and selects the mode that best suits their needs (e.g., hearing-impaired mode, visually-impaired mode).
[0126] The server receives the user's configuration information and prepares to provide the required services.
[0127] Step 2: Real-time analytics and notifications
[0128] Based on the selected mode, the device will activate the necessary sensors (microphone, camera, etc.) and collect data.
[0129] The server analyzes the data sent from the terminal and returns the necessary information in real time.
[0130] The device notifies the user of the analysis results in an appropriate format (subtitles, audio).
[0131] Step 3: Terminate or suspend service
[0132] If a user wants to stop using the system, they can select the option to terminate or suspend the service on their device's settings screen.
[0133] The terminal receives the service termination command, shuts down all sensors, and cuts off communication with the server.
[0134] Example 1
[0135] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0136] The purpose of this invention is to provide support for the hearing-impaired and visually-impaired to overcome the difficulties they face in daily life and live safely and independently. Specifically, the invention involves the development of a system that recognizes surrounding sounds, sign language, and environmental conditions in real time and provides the necessary information in an appropriate format.
[0137] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0138] In this invention, the server includes means for capturing ambient sounds, means for preprocessing the captured audio data, means for converting the preprocessed audio data into text data, means for analyzing the type, direction, and distance of sound, means for displaying the analyzed information, means for capturing sign language, means for preprocessing the captured sign language data, means for converting the preprocessed sign language data into text data, means for converting the converted text data into audio data and outputting it, means for capturing ambient environmental data, means for preprocessing the captured environmental data, means for analyzing the preprocessed environmental data to recognize objects, means for generating audio guidance based on the recognized object information, and means for outputting the generated audio guidance. This enables hearing-impaired and visually impaired people to grasp their surroundings in real time and live safely and independently.
[0139] The "means for capturing ambient sound" refers to a device or sensor function for collecting audio data from the user's surrounding environment.
[0140] "Means for preprocessing captured audio data" refers to processes or algorithms used to remove noise from collected audio data and prepare it for speech recognition.
[0141] The "means for converting preprocessed voice data into text data" refers to software or a system for converting voice data into text information using voice recognition technology.
[0142] "Means for analyzing the type, direction and distance of sound" refers to an algorithm or device that analyzes text data and sound feature data to identify the source, type, direction and distance of a sound.
[0143] "Means for displaying analyzed information" refers to a device or system that uses a display, screen, or the like to visually convey the analysis results to the user.
[0144] A "means for capturing sign language" is a camera or sensor that records the user's hand movements and shapes in real time.
[0145] "Means for pre-processing captured sign language data" refers to a process or system that appropriately processes captured video data for sign language recognition and extracts hand shapes and movements.
[0146] A "means for converting preprocessed sign language data into text data" is a technology or system that analyzes captured sign language movements and converts them into corresponding text information.
[0147] "Means for converting the converted text data into audio data and outputting it" refers to software or a device for converting character information into audio and reproducing it via a speaker or the like.
[0148] The "means for capturing surrounding environmental data" refers to a device or function for collecting information about the user's surrounding environment using a camera or sensor.
[0149] "Means for pre-processing captured environmental data" refers to processes or algorithms that process collected environmental data into a form suitable for image recognition.
[0150] "Means for analyzing preprocessed environmental data to recognize objects" refers to AI algorithms or systems that identify and recognize the shape and position of objects from environmental data.
[0151] The "means for generating audio guidance based on recognized object information" refers to a system or process for creating audio guidance to be notified to the user based on information about the identified object.
[0152] The "means for outputting the generated audio guide" refers to a device or system for conveying the generated audio guide to the user using a speaker or the like.
[0153] This invention is an assistance system for hearing-impaired and visually-impaired people to overcome difficulties in daily life, and is implemented as a system including a user's mobile terminal (hereinafter referred to as "terminal") and a cloud-based server (hereinafter referred to as "server"). The terminal is equipped with a microphone, camera, speaker, display, and various sensors. These hardware components are used to collect and analyze information about the user's surroundings and provide the necessary information.
[0154] Hearing-impaired features
[0155] Subtitling of ambient sounds
[0156] The device uses a microphone to capture ambient sounds and preprocesses the audio data. It then uses noise reduction technology to remove background noise and converts the audio into text using a speech recognition engine. The device uses multiple microphone arrangements to analyze the type, direction, and distance of the sound and displays the analyzed information on the display. For example, if a user is walking on the sidewalk and hears a car horn honking from the right at a distance of 10 meters, the device will display the message, "A car horn is honking from the right at a distance of 10 meters."
[0157] Automatic sign language recognition and translation
[0158] The device uses a camera to capture the user's sign language in real time and preprocesses the video data. After extracting the hand shape and movement, a sign language recognition algorithm generates corresponding text information. Next, a speech synthesis engine converts the text information into audio data and outputs it from the speaker. For example, if the user signs "hello," the device will recognize the sign and output "hello" aloud.
[0159] Features for the visually impaired
[0160] Audio guidance of surrounding objects and situations
[0161] The device uses cameras and various sensors to capture and preprocess data on the surrounding environment. It uses image recognition technology to extract the contours and feature points of objects, and uses AI algorithms to recognize surrounding objects and situations. It analyzes the location and movement of recognized objects and generates important information as audio guidance. This audio guidance is converted into voice by a speech synthesis engine and notified to the user through the speaker. For example, if a user is walking and there is an obstacle one meter ahead, the device will announce, "There is an obstacle one meter ahead."
[0162] Starting and shutting down the system
[0163] During the initial setup, the user selects the system mode (deaf mode, visually impaired mode) that best suits their needs. The server receives the user's configuration information and prepares the required services. The device activates the required sensors (microphone, camera, etc.) based on the initial setup and begins collecting data. The server analyzes the data sent from the device and returns the required information in real time. If the user wants to terminate or pause the service, they select the option to terminate or pause the service on the device's settings screen. The device receives the service termination command, stops all sensors, and cuts off communication with the server.
[0164] Examples of concrete examples and prompts
[0165] Example prompt for speech recognition: "Transcribe the speech that says 'hello' to text."
[0166] Example prompt for sign language recognition: "Translate this sign language video into text."
[0167] This provides useful support to hearing-impaired and visually impaired people to overcome difficulties in daily life and live safely and independently.
[0168] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0169] Hearing-impaired features
[0170] Subtitling of ambient sounds
[0171] Step 1: Capture the sound
[0172] Subject: Device
[0173] How it works: The device uses a microphone to capture ambient sound, taking ambient audio data as input and raw captured data as output.
[0174] Input and Output: The input is the ambient audio data, and the output is the captured audio data.
[0175] Step 2: Pre-processing the sound
[0176] Subject: Device
[0177] Specific operation: The device preprocesses the captured audio data and removes noise. The input is the captured audio data, and the output is the noise-removed audio data.
[0178] Input and Output: The input is the captured audio data, and the output is the pre-processed audio data.
[0179] Step 3: Voice Recognition
[0180] Subject: Device
[0181] Specific operation: The device sends the preprocessed voice data to the voice recognition engine and converts it into text data. The preprocessed voice data is input, and the converted text data is obtained as output.
[0182] Input and Output: The input is preprocessed audio data, and the output is text data.
[0183] Step 4: Analyzing the sound information
[0184] Subject: Device
[0185] Specific operation: The device analyzes the text data and sound feature data to identify the type, direction, and distance of the sound. The text data and sound feature data are input, and the analysis results are obtained as output.
[0186] Input and output: The input is text data and sound feature data, and the output is analyzed sound information.
[0187] Step 5: Viewing information
[0188] Subject: Device
[0189] Specific operation: The terminal displays the analyzed information in text format on the display. The analyzed sound information is input and displayed on the display as output.
[0190] Input and output: The input is the analyzed sound information, and the output is the text information displayed on the screen.
[0191] Automatic sign language recognition and translation
[0192] Step 1: Capture sign language
[0193] Subject: Device
[0194] Specific operation: The device uses a camera to capture the user's sign language actions in real time. The sign language actions are input and the captured video data is obtained as output.
[0195] Input and Output: The input is sign language actions, and the output is captured video data.
[0196] Step 2: Preprocessing the video data
[0197] Subject: Device
[0198] Specific operation: The device preprocesses the captured video data and extracts the hand shape and movement. The captured video data is input and the preprocessed video data is output.
[0199] Input and Output: The input is the captured video data, and the output is the pre-processed video data.
[0200] Step 3: Sign Language Recognition and Text Conversion
[0201] Subject: Device
[0202] Specific operation: The device sends the preprocessed video data to a sign language recognition algorithm to generate corresponding text data. The preprocessed video data is input, and the generated text data is output.
[0203] Input and Output: The input is the preprocessed video data, and the output is the generated text data.
[0204] Step 4: Speech synthesis and output
[0205] Subject: Device
[0206] Specific operation: The device sends the generated text data to a speech synthesis engine, converts it into voice data, and outputs the converted voice data to the user through the speaker. Text data is input, and voice data is obtained as output.
[0207] Input and output: The input is the generated text data, and the output is the audio output from the speaker.
[0208] Features for the visually impaired
[0209] Audio guidance of surrounding objects and situations
[0210] Step 1: Capture environment data
[0211] Subject: Device
[0212] Specific operation: The device uses a camera and various sensors to capture surrounding environmental data. The surrounding environmental information is input, and the captured environmental data is output.
[0213] Input and Output: The input is the surrounding environment information, and the output is the captured environment data.
[0214] Step 2: Preprocessing the image data
[0215] Subject: Device
[0216] Specific operation: The device preprocesses the captured environmental data and extracts object contours and feature points. The captured environmental data is input, and the preprocessed environmental data is output.
[0217] Input and Output: The input is the captured environmental data, and the output is the preprocessed environmental data.
[0218] Step 3: Object Recognition
[0219] Subject: Device
[0220] How it works: The device analyzes the preprocessed data using AI algorithms to recognize surrounding objects and situations. The preprocessed data is input, and the recognized object information is output.
[0221] Input and Output: The input is the preprocessed data, and the output is the recognized object information.
[0222] Step 4: Location Analysis
[0223] Subject: Device
[0224] Specific operation: The device analyzes the location and movement of the recognized object and extracts important information. The recognized object information is input, and the location and movement analysis results are obtained as output.
[0225] Input and Output: The input is the recognized object information, and the output is the analysis result of the position and movement.
[0226] Step 5: Audio guide generation and output
[0227] Subject: Device
[0228] Specific operation: The device generates the analyzed information as an audio guide and converts it into speech using a speech synthesis engine. The converted audio guide is then notified to the user through the speaker. The analyzed information is input, and the audio guide is obtained as output.
[0229] Input and Output: The input is the parsed information, and the output is the audio guide from the speaker.
[0230] Starting and shutting down the system
[0231] Step 1: Initial Setup
[0232] Subject: User
[0233] Specific operation: The user selects the system mode (deaf mode, visually impaired mode) according to their needs in the initial setting. The user's needs are input, and the setting information is sent to the server as output.
[0234] Input and Output: Input is user needs, output is configuration information.
[0235] Step 2: Start the service
[0236] Subject: Server
[0237] Specific operation: The server receives the user's configuration information and prepares the necessary services. The device activates the necessary sensors (microphone, camera, etc.) based on the initial settings and begins collecting data. The configuration information is received as input and the service is started as output.
[0238] Input and Output: Input is the configuration information, output is the started service.
[0239] Step 3: Data collection and analysis
[0240] Subject: Device
[0241] Specific operation: The terminal starts collecting data, and the server analyzes the data sent from the terminal. The collected data is input, and the analysis results are obtained as output.
[0242] Input and Output: Input is the collected data and output is the analysis result.
[0243] Step 4: Notification of information
[0244] Subject: Device
[0245] Specific operation: The device displays or outputs the analysis results to the user. The analysis results are input and notified to the user as output.
[0246] Input and output: The input is the analysis result, and the output is the notification to the user.
[0247] Step 5: Terminate or suspend service
[0248] Subject: User
[0249] Specific behavior: When a user wants to stop using a service, they select the option to terminate or pause the service in the device settings screen. The user's operation is the input, and the service is terminated or paused as the output.
[0250] Input and Output: Input is a user action, output is a service that has been terminated or suspended.
[0251] (Application example 1)
[0252] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0253] The purpose of this invention is to provide a means to resolve the lack of information and communication difficulties faced by hearing-impaired and visually-impaired people when shopping comfortably and safely in a brick-and-mortar store. Specifically, the objective is to provide a system that supports people with disabilities to act independently by acquiring in-store announcements, conversations with store staff, and product information in real time and conveying this information appropriately to the user.
[0254] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0255] In this invention, the server includes means for capturing ambient sounds, means for converting the captured audio data into text data, means for analyzing the type, direction, and distance of the sound, means for displaying the analyzed information, means for displaying the captured text data on a smart device display, means for capturing sign language, means for converting the captured sign language data into text data, means for converting the converted text data into audio, means for capturing ambient environmental data, means for analyzing the captured environmental data, means for generating the analyzed information as an audio guide, and means for outputting the generated audio guide. This allows hearing-impaired people to understand in-store announcements and conversations in real time through subtitles, and makes it easy to convert sign language communication into audio. Furthermore, visually impaired people can receive information about surrounding objects and products as an audio guide, enabling safe and efficient shopping.
[0256] "Means for capturing ambient sounds" is a function that collects environmental sounds using the microphone of the smart device.
[0257] "Means for converting captured voice data into text data" refers to a function that transcribes collected voice data using voice recognition technology.
[0258] "Means for analyzing the type, direction and distance of sound" is a function that analyzes collected audio data to identify the type, location and distance of the sound source.
[0259] "Means for displaying analyzed information" refers to a function that visually displays the analysis results on the display of a smart device.
[0260] The "means for displaying captured text data on the display of the smart device" is a function for displaying text obtained by speech recognition on the screen of the smart device.
[0261] A "means for capturing sign language" is a function that uses the camera of a smart device to capture sign language movements.
[0262] The "means for converting captured sign language data into text data" is a function that analyzes the video of the sign language and converts it into text information.
[0263] The "means for converting the converted text data into voice" is a function for outputting the text data as voice using voice synthesis technology.
[0264] "Means for capturing surrounding environmental data" refers to the ability to collect images and data of the environment using the smart device's camera and sensors.
[0265] "Means for analyzing captured environmental data" refers to a function that analyzes collected video and data to recognize objects and environments.
[0266] "Means for generating analyzed information as audio guide" is a function for creating audio guide based on the analysis results.
[0267] "Means for outputting the generated audio guide" is a function that transmits the generated audio guide to the user through the speaker of the smart device.
[0268] This invention is a system that helps hearing- and visually-impaired people enjoy shopping comfortably and safely in physical stores. The system consists of a user's smart device and a cloud-based server. Specifically, the system collects and analyzes surrounding information using the smart device's microphone, camera, display, speaker, and various sensors. This collected and analyzed information is then provided to the user appropriately.
[0269] System configuration
[0270] 1. Subtitling of ambient sounds
[0271] Capture method: Capture surrounding sounds using the microphone on your smart device.
[0272] Voice data conversion: The captured voice data is converted into text data using voice recognition technology. Specifically, Google Cloud's Speech-to-Text API is used.
[0273] Analysis means: Analyzes the type, direction and distance of sound and provides important information to the user.
[0274] Display method: The analyzed information is displayed in real time as subtitles on the smart device display, allowing hearing-impaired people to understand in-store announcements and conversations.
[0275] 2. Automatic Sign Language Recognition and Translation
[0276] Capture method: Sign language is captured using the smart device's camera.
[0277] Sign language data conversion: The captured sign language data is converted into text data using image processing techniques and sign language recognition algorithms.
[0278] Speech conversion method: The converted text data is converted into speech using speech synthesis technology (Google Cloud's Text-to-Speech API) and played back through the smart device's speaker, allowing hearing-impaired people to easily communicate with store staff and other customers.
[0279] 3. Audio guide to surrounding objects and situations
[0280] Capture method: Capture surrounding environmental data using the smart device's camera and various sensors.
[0281] Environmental data analysis: The captured data is analyzed using image recognition technology (Google Vision API) to identify objects and obtain their location information.
[0282] Audio guide generation: Audio guides are generated based on the analyzed information and converted into audio using Google Cloud's Text-to-Speech API.
[0283] Output means: The generated audio guide is output from the smart device's speaker to inform visually impaired people of their surroundings.
[0284] Specific examples
[0285] For example, when a user wears a smart device and enters a physical store, the system captures in-store announcements with a microphone and displays information such as "Today's special sale is XX" in real time as subtitles on the display. Also, if the user expresses "Where is it?" in sign language, the system captures the sign, converts it into text, and relays it to the store clerk as "Where is it?" Furthermore, when a visually impaired person stands in front of a shelf, the camera recognizes the product label and provides an audio guide such as "There is a fruit shelf ahead." This guide is provided through the smart device's speaker.
[0286] Prompt Sentence Examples
[0287] "Please provide real-time subtitles for in-store announcements."
[0288] "Please recognize the sign language, convert it into speech, and communicate it to the store clerk."
[0289] "Recognize surrounding objects and provide voice guidance"
[0290] As can be seen, this system provides a useful aid for the hearing and visually impaired to enjoy shopping independently.
[0291] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0292] Step 1:
[0293] Capture Processing
[0294] Using sensors such as the device's microphone and camera, surrounding audio and video data is captured in real time.
[0295] Input: Ambient audio, video such as sign language, and environmental data.
[0296] How it works: A microphone captures audio data, and a camera captures sign language and footage of the environment.
[0297] Step 2:
[0298] Pretreatment
[0299] Noise reduction and image processing are performed on the captured audio and video data, converting it into a format that is easy to analyze.
[0300] Input: Raw captured data (audio data, video data).
[0301] Data processing: Noise reduction of audio data, contour extraction of video data.
[0302] Output: Preprocessed audio data, preprocessed video data.
[0303] Operation: A noise reduction algorithm is used to remove noise from the audio data, and a contour extraction filter is applied to the video data.
[0304] Step 3:
[0305] Data transformation and analysis
[0306] The preprocessed data is analyzed using a voice recognition engine or image recognition algorithm.
[0307] Input: Preprocessed audio data, preprocessed video data.
[0308] Data calculation: Voice data is converted into text using a speech recognition algorithm (Google Cloud Speech-to-Text API), and sign language is analyzed and converted into text data using an image recognition algorithm.
[0309] Output: Text data, object recognition data.
[0310] How it works: It calls the Google Cloud Speech-to-Text API to convert speech to text, then uses a sign language recognition algorithm to convert sign language to text.
[0311] Step 4:
[0312] Displaying and Speech Output of Data
[0313] The converted text data is displayed on the smart device's display, and if necessary, a voice synthesis engine is used to output the text as speech.
[0314] Input: Converted text data, object recognition data.
[0315] Data calculation: Text data is displayed as a user interface, and speech is generated using a speech synthesis algorithm (Google Cloud Text-to-Speech API).
[0316] Output: Text displayed on the screen, audio guidance output from the speaker.
[0317] How it works: Text data is displayed as subtitles on the smart device's display, and audio data is generated by calling the Google Cloud Text-to-Speech API and output through the speaker.
[0318] Step 5:
[0319] User Feedback Processing
[0320] If there is additional input from the user, that input is passed back to the capture process for further analysis and interaction.
[0321] Input: New input from the user (speech, sign language, changes in the environment).
[0322] Data processing: Preprocess and analyze again as necessary.
[0323] Output: Updated visual and audio guides.
[0324] What it does: Captures user feedback, then analyzes it again and displays / speech outputs it.
[0325] Through these steps, hearing and visually impaired people can have a more comfortable and safer shopping experience in physical stores.
[0326] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0327] The present invention is an assistive system for the hearing and visually impaired, which has the ability to recognize a user's emotions and use that information to adjust output. Detailed description of the preferred embodiments of the present invention is provided below.
[0328] System configuration
[0329] This system consists of a user's mobile device (hereafter referred to as the device) and a cloud-based server (hereafter referred to as the server). The device is equipped with a microphone, camera, speaker, display, various sensors, and an emotion recognition engine, and uses this hardware to collect and analyze information about the user's surroundings and emotional state, and provide the necessary information.
[0330] Hearing-impaired features
[0331] Subtitling of ambient sounds
[0332] 1. Sound capture and processing
[0333] The device uses a microphone to capture ambient sounds at 0.1 second intervals.
[0334] The captured audio data undergoes pre-processing such as noise reduction inside the device.
[0335] The preprocessed voice data is converted into text in real time by a voice recognition engine.
[0336] 2. Analysis of sound type, direction, and distance
[0337] The device identifies the type of sound from the converted text information and sound characteristics (e.g., "ambulance siren," "human voice," etc.).
[0338] The device uses multiple microphone arrangements to calculate the direction and distance from which the sound is coming.
[0339] 3. Displaying Information
[0340] The analyzed information is displayed in text format on the device's display, allowing the user to understand the surrounding situation in real time.
[0341] 4. Emotion recognition and regulation
[0342] The device uses a camera and microphone to capture the user's facial expressions and tone of voice, which are then analyzed by an emotion engine.
[0343] The text displayed and notification method are adjusted based on the analyzed emotional information.
[0344] For example, when displaying "The sound of a horn is honking from 10 meters to the right," if the user appears surprised, the display will be emphasized.
[0345] Features for the visually impaired
[0346] Audio guidance of surrounding objects and situations
[0347] 1. Environmental data capture and processing
[0348] The device uses a camera and various sensors to capture data about the surrounding environment at 0.5-second intervals.
[0349] The captured data is subjected to image recognition preprocessing to extract the contours and feature points of the object.
[0350] 2. Object Recognition and Analysis
[0351] The pre-processed data is then used by AI algorithms to recognize surrounding objects.
[0352] The location and movement of recognized objects are analyzed to pick out important information.
[0353] 3. Audio guide generation and output
[0354] The analyzed information is generated as audio guide and converted into voice by a voice synthesis engine.
[0355] The terminal notifies the user of the generated audio guide through a speaker.
[0356] 4. Emotion recognition and regulation
[0357] The device captures the user's emotions and adjusts the audio guidance based on the analyzed emotion information.
[0358] As a specific example, if the user is nervous, the tone of the voice guide is softened.
[0359] Automatic sign language recognition and translation
[0360] 1. Sign Language Capture and Processing
[0361] The device uses a camera to capture the user's sign language in real time.
[0362] The captured video data is pre-processed to extract the shape and movement of the hand.
[0363] 2. Sign Language Recognition and Translation
[0364] The pre-processed video data is analyzed by a sign language recognition algorithm to generate corresponding text information.
[0365] The generated text information is converted into voice data by a voice synthesis engine.
[0366] 3. Audio Output
[0367] The device plays the translated audio through a speaker, helping people who don't understand sign language to communicate with the user.
[0368] 4. Emotion Recognition and Translation Adjustment
[0369] The device analyzes the user's emotions along with the captured sign language data.
[0370] The device adjusts the expression and tone of the text translation based on the emotional information.
[0371] As a specific example, if the user says "hello" in sign language and has a happy expression, the tone of the voice output is brightened.
[0372] Starting and shutting down the system
[0373] 1. Initial Setup
[0374] The user configures the system and selects the mode that best suits their needs (e.g., hearing-impaired mode, visually-impaired mode).
[0375] The server receives the user's configuration information and prepares to provide the required services.
[0376] 2. Real-time analysis and notifications
[0377] Based on the selected mode, the device will activate the necessary sensors (microphone, camera, etc.) and collect data.
[0378] The server analyzes the data sent from the terminal and returns the necessary information in real time.
[0379] The device notifies the user of the analysis results in an appropriate format (subtitles, audio), adjusting it according to the emotional information.
[0380] 3. Termination or Suspension of Service
[0381] If a user wants to stop using the system, they can select the option to terminate or suspend the service on their device's settings screen.
[0382] The terminal receives the service termination command, shuts down all sensors, and cuts off communication with the server.
[0383] This system provides useful support for hearing- and visually-impaired people to overcome difficulties in daily life and live safely and independently. Furthermore, by adjusting its response according to the user's emotions, it is possible to provide support that meets more individual needs.
[0384] The processing flow will be explained below.
[0385] Features for the hearing impaired: Subtitling of surrounding sounds
[0386] Step 1: Capture the sound
[0387] The device activates the microphone and captures ambient sounds at 0.1 second intervals.
[0388] Step 2: Preprocessing the audio data
[0389] The device performs noise reduction and normalization on the captured audio data.
[0390] Step 3: Speech to text
[0391] The device sends the preprocessed voice data to a voice recognition engine, which converts it into text data in real time.
[0392] Step 4: Identify the type of sound
[0393] The device extracts sound characteristics from the text data and identifies the type of sound based on them (e.g., "ambulance siren," "human voice," etc.).
[0394] Step 5: Analyze sound direction and distance
[0395] The device uses multiple microphone arrangements to calculate the direction and distance from which the sound is coming.
[0396] Step 6: Capturing Emotions
[0397] The device uses a camera and microphone to capture the user's facial expressions and tone of voice.
[0398] Step 7: Analyze the sentiment data
[0399] The device sends the captured emotion data to an emotion engine to analyze the user's emotions.
[0400] Step 8: Display subtitles
[0401] Based on the analyzed voice information and emotional data, the device displays text information on the screen, conveying the type, direction, and distance of the sound to the user in a format that is optimal for them.
[0402] For example, when displaying the message "A horn is honking from 10 meters to the right," if the user appears surprised, the warning message will be emphasized.
[0403] Features for the visually impaired: Audio guide to surrounding objects and situations
[0404] Step 1: Capture environment data
[0405] The device will activate its camera and sensors to capture data on the surrounding environment every 0.5 seconds.
[0406] Step 2: Preprocessing the data
[0407] The device performs image recognition preprocessing on the captured video data to extract the contours and feature points of objects.
[0408] Step 3: Object Recognition
[0409] The device sends the pre-processed data to AI algorithms to recognize surrounding objects.
[0410] Step 4: Obtaining information about the object
[0411] The device analyzes the location and movement of recognized objects and picks out important information (e.g., obstacles ahead, approaching people).
[0412] Step 5: Capturing emotions
[0413] The device uses a camera and microphone to capture the user's facial expressions and tone of voice.
[0414] Step 6: Analyze the sentiment data
[0415] The device sends the captured emotion data to an emotion engine to analyze the user's emotions.
[0416] Step 7: Generate audio guide
[0417] The device generates audio guidance based on the analyzed environmental information and emotional data.
[0418] For example, if the user is nervous, the tone of the voice prompt will be softened and the user will be notified, "There is an obstacle 1 meter ahead."
[0419] Step 8: Audio Notifications
[0420] The terminal notifies the user of the generated audio guide through a speaker.
[0421] Automatic sign language recognition and translation
[0422] Step 1: Capture sign language
[0423] The device uses a camera to capture the user's sign language in real time.
[0424] Step 2: Preprocessing the sign language data
[0425] The device performs pre-processing on the captured video data to extract the shape and movement of the hand.
[0426] Step 3: Sign Language Recognition
[0427] The device sends the preprocessed data to a sign language recognition algorithm to generate corresponding text information.
[0428] Step 4: Capturing emotions
[0429] The device captures the user's facial expressions with a camera along with the captured video data, and obtains emotion data.
[0430] Step 5: Analyze the sentiment data
[0431] The device analyzes the emotional data using an emotion engine and determines the user's emotions.
[0432] Step 6: Adjusting the text data
[0433] The device adjusts the text information generated from the sign language as needed based on emotional data.
[0434] Step 7: Converting text data into speech
[0435] The device sends the adjusted text data to a speech synthesis engine and converts it into voice data.
[0436] Step 8: Audio Output
[0437] The device plays audio data through a speaker, helping people who do not understand sign language to communicate with the user.
[0438] As a specific example, if a user who signs "hello" looks happy, the tone of the voice output is made brighter.
[0439] Example 2
[0440] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0441] The present invention aims to solve the problem of the lack of a function to adjust information according to the user's emotions in the conventional technology for life support systems for the hearing-impaired and visually-impaired. Specifically, the present invention aims to support a more comfortable and safe life by capturing surrounding sounds and environmental information in real time and providing it to the user, while adjusting the output content and format based on the user's emotions.
[0442] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0443] In this invention, the server includes means for capturing ambient sounds, means for preprocessing the captured audio data, means for converting the preprocessed audio data into text data, means for analyzing the type, direction, and distance of the sound, means for displaying the analyzed information, and means for recognizing the user's emotions and adjusting the display information. This allows the user to grasp the surrounding situation in real time and receive appropriate notifications according to the user's emotions. Furthermore, by including means for capturing sign language, means for preprocessing the captured sign language data, means for converting the preprocessed sign language data into text data, means for converting the converted text data into audio data and outputting it, and means for recognizing the user's emotions and adjusting the audio output, communication with sign language users can be facilitated. Furthermore, by including means for capturing ambient environmental data, means for preprocessing the captured environmental data, means for analyzing the preprocessed environmental data, means for generating the analyzed information as audio guidance, means for outputting the generated audio guidance, and means for recognizing the user's emotions and adjusting the audio guidance, it becomes easier for visually impaired people to grasp the surrounding situation and appropriate guidance according to their emotions can be provided.
[0444] The "means for capturing ambient sound" refers to a means for capturing ambient sound using an acoustic sensor such as a microphone and acquiring the audio data.
[0445] The "means for pre-processing captured audio data" refers to a means for performing processes such as noise reduction and filtering on the acquired audio data to make the audio signal clearer.
[0446] The "means for converting preprocessed voice data into text data" refers to a means for analyzing preprocessed voice data using a voice recognition technique and generating corresponding text data.
[0447] "Means for analyzing the type, direction and distance of a sound" refers to means for analyzing the characteristics of pre-processed and text-converted audio data to identify what the sound is (e.g., a car horn, a human voice, etc.), the direction from which it is coming, and how far away it is.
[0448] The "means for displaying analyzed information" refers to a means for displaying the analyzed audio information on a display or monitor so that the user can visually confirm the information.
[0449] "Means for recognizing the user's emotions and adjusting the displayed information" refers to means for analyzing the user's facial expressions and tone of voice using a camera or microphone, determining the user's emotions, and then adjusting the format and content of the displayed information.
[0450] The "means for capturing sign language" refers to a means for capturing a user's sign language actions in real time using a video capture device such as a camera, and acquiring the data.
[0451] The "means for preprocessing captured sign language data" refers to a means for processing the acquired sign language video data to make it easier to analyze, and extracting features such as hand shape and movement.
[0452] The "means for converting preprocessed sign language data into text data" refers to a means for analyzing the preprocessed sign language data using a sign language recognition algorithm and generating text data representing the meaning of the sign language.
[0453] "Means for converting the converted text data into voice data and outputting it" refers to means for generating voice data based on the text data using voice synthesis technology and notifying the user of the voice data through a speaker or the like.
[0454] The "means for capturing surrounding environmental data" refers to a means for obtaining visual and other data about surrounding objects and the environment using a camera or various sensors.
[0455] "Means for preprocessing captured environmental data" refers to means for performing image processing and data analysis on the acquired environmental data, extracting the contours and feature points of objects, and making the data easier to analyze.
[0456] "Means for analyzing preprocessed environmental data" refers to means for analyzing preprocessed environmental data using AI algorithms and recognizing objects and situations.
[0457] The "means for generating analyzed information as audio guidance" refers to a means for generating information to be provided to the user as audio guidance based on the analysis results.
[0458] The "means for outputting the generated audio guide" refers to a means for outputting the generated audio guide as audio using speech synthesis technology and notifying the user through a speaker.
[0459] The "means for recognizing the user's emotions and adjusting the audio guidance" refers to a means for analyzing the user's facial expressions and tone of voice, and adjusting the tone and content of the audio guidance based on the determined emotional information.
[0460] The present invention is an assistance system for the hearing-impaired and visually-impaired, which has the ability to recognize a user's emotions and adjust output using that information. Detailed embodiments for implementing the present invention will be described below.
[0461] System configuration
[0462] This system consists of a cloud-based server (hereafter referred to as the "server") and a mobile device (hereafter referred to as the "device") carried by the user. The device is equipped with a microphone, camera, speaker, display, various sensors, and an emotion recognition engine.
[0463] Embodiments for the hearing impaired
[0464] Subtitling of ambient sounds
[0465] The device uses a microphone to capture ambient sounds at 0.1-second intervals. The captured voice data undergoes preprocessing such as noise reduction within the device, and is then converted into text data by a voice recognition engine. This text data is then used to analyze the type, direction, and distance of the sound, and displayed on the device's display. The device also uses a camera and microphone to capture the user's emotions, which are analyzed by an emotion engine. The device adjusts the text displayed and notification method based on this emotional information.
[0466] As a specific example of how this works, if the user is surprised by the information "There is a horn honking 10 meters to the right," it will be displayed with an emphasis of "Caution!"
[0467] Embodiments for the visually impaired
[0468] Audio guidance of surrounding objects and situations
[0469] The device captures data on the surrounding environment every 0.5 seconds using a camera and various sensors. The captured data undergoes image recognition preprocessing to extract object contours and feature points. This data is analyzed using an AI algorithm to identify the location and movement of recognized objects and generate audio guidance. This audio guidance is output from the speaker via a speech synthesis engine. The device also captures the user's emotions and adjusts the audio guidance based on the analyzed emotional information.
[0470] As a specific example of how it works, if the user is nervous about an approaching car, the system will notify them in a soft tone, "There is a car approaching 10 meters ahead, please proceed slowly."
[0471] Embodiments of automatic sign language recognition and translation
[0472] The device uses a camera to capture sign language in real time and preprocesses the hand shape and movements. The preprocessed video data is then analyzed using a sign language recognition algorithm to generate corresponding text data. This text data is converted into audio data by a speech synthesis engine and output through the speaker. The device also captures the user's emotions and adjusts the tone of the voice based on the emotional information.
[0473] As a specific example of operation, if the user signs "hello" and looks happy, "hello" is output as voice in a bright tone.
[0474] Examples of prompt statements
[0475] An example prompt to be input to a generative AI model would be something like this:
[0476] 1. Prompt for the deaf:
[0477] "Capture ambient sounds and subtitle them in real time. Analyze the type, direction, and distance of the sound and display it in text format."
[0478] 2. Prompt for the visually impaired:
[0479] "Use cameras and sensors to capture data about the surrounding environment and output recognized object information as audio guidance."
[0480] 3. Prompt for sign language recognition:
[0481] "Use your camera to capture sign language in real time and translate it into text or speech."
[0482] As described above, this system adjusts the way information is displayed and notified according to the user's emotions, providing assistance that meets more individual needs.
[0483] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0484] Processing steps for the hearing impaired
[0485] Subtitling of ambient sounds
[0486] Step 1: Capture the sound
[0487] The device uses a microphone to capture ambient sounds at 0.1 second intervals.
[0488] Input: Ambient sound
[0489] Specific operation: The device picks up surrounding environmental sounds and voices using the microphone.
[0490] Output: Captured audio data
[0491] Step 2: Preprocessing the audio data
[0492] The device performs pre-processing such as noise reduction and filtering on the captured audio data.
[0493] Input: Captured audio data
[0494] What it does: The device removes background noise and emphasizes the main voice.
[0495] Output: Preprocessed audio data
[0496] Step 3: Convert audio data to text
[0497] The device sends the preprocessed voice data to a voice recognition engine, which converts it into text data in real time.
[0498] Input: Preprocessed audio data
[0499] Specific operation: The voice recognition engine converts spoken content such as "Hello, it's a nice day today" into text.
[0500] Output: Text data
[0501] Step 4: Analyze the type, direction and distance of the sound
[0502] The device uses the characteristics of text and voice data to identify the type of sound and calculates direction and distance using multiple microphone arrangements.
[0503] Input: Text data, audio data features
[0504] Specific operation: The device uses an analysis algorithm to classify sounds such as "car horns" and "human voices" and calculates direction and distance.
[0505] Output: Analyzed sound type, direction and distance
[0506] Step 5: Viewing information
[0507] The device displays the analyzed voice information on the display.
[0508] Input: Analyzed audio information
[0509] Specific operation: The device displays information such as "A horn is honking 10 meters to the right" on the screen.
[0510] Output: Information displayed on the display
[0511] Step 6: Capture and analyze emotions
[0512] The device uses a camera and microphone to capture the user's facial expressions and tone of voice, which are then analyzed by an emotion engine.
[0513] Input: User facial expressions and tone of voice
[0514] What it does: The device uses a facial expression analysis algorithm to detect emotions such as surprise or anxiety.
[0515] Output: Parsed emotion information
[0516] Step 7: Adjust notification methods
[0517] The device adjusts the text displayed and notification method based on the analyzed emotional information.
[0518] Input: Parsed emotion information
[0519] Specific behavior: If the user is startled, the message "Sound of horn 10 meters to the right -- Caution!" will be displayed in emphasis.
[0520] Output: Adjusted display information
[0521] Processing steps for the visually impaired
[0522] Audio guidance of surrounding objects and situations
[0523] Step 1: Capture environment data
[0524] The device uses a camera and various sensors to capture data on the surrounding environment every 0.5 seconds.
[0525] Input: Surrounding environment
[0526] Specific operation: The camera captures images of the road, pedestrians, cars, etc. ahead, and the sensor measures distance and movement.
[0527] Output: Captured environmental data
[0528] Step 2: Preprocessing environmental data
[0529] The device performs image processing on the captured environmental data to extract the contours and feature points of objects.
[0530] Input: Captured environmental data
[0531] How it works: Image processing algorithms extract the shape and features of objects.
[0532] Output: Preprocessed environmental data
[0533] Step 3: Object Recognition and Analysis
[0534] The device sends pre-processed environmental data to AI algorithms that recognize surrounding objects and analyze their position and movement.
[0535] Input: Preprocessed environmental data
[0536] Specific operation: The AI recognizes things like "there are two people ahead" or "a car is approaching."
[0537] Output: Position and movement of analyzed objects
[0538] Step 4: Generate audio guide
[0539] The device generates audio guidance based on the analyzed information and converts it into audio data using a speech synthesis engine.
[0540] Input: Analyzed object information
[0541] Specific operation: Generates voice guidance such as "There is a step 10 meters ahead."
[0542] Output: The generated audio guide
[0543] Step 5: Outputting voice guidance
[0544] The device will announce the generated audio guide through the speaker.
[0545] Input: Generated audio guide
[0546] Specific operation: The device will announce "There is a step 10 meters ahead" via voice.
[0547] Output: Guide output as audio
[0548] Step 6: Capture and analyze emotions
[0549] The device uses a camera and microphone to capture the user's emotions and analyzes them using an emotion engine.
[0550] Input: User facial expressions and tone of voice
[0551] Specific operation: The device analyzes facial expressions to determine whether the user is anxious or nervous.
[0552] Output: Parsed emotion information
[0553] Step 7: Adjust the audio prompts
[0554] The device adjusts the audio guidance based on the analyzed emotional information.
[0555] Input: Parsed emotion information
[0556] Specific behavior: If the user is nervous, a soft tone will be displayed saying, "There is a step ahead, please walk slowly."
[0557] Output: Adjusted voice prompts
[0558] Processing steps for automatic sign language recognition and translation
[0559] Step 1: Capture sign language
[0560] The device uses a camera to capture the user's sign language in real time.
[0561] Input: User's sign language actions
[0562] Specific operation: The camera captures the user's hand movements.
[0563] Output: Captured sign language data
[0564] Step 2: Preprocessing the sign language data
[0565] The device preprocesses the captured sign language data to make it easier to analyze, extracting the shape and movement of the hand.
[0566] Input: Captured sign language data
[0567] Specific operation: An algorithm is run to analyze the shape and movement of the hand.
[0568] Output: Preprocessed sign language data
[0569] Step 3: Recognize and translate sign language data
[0570] The terminal sends the preprocessed sign language data to a sign language recognition algorithm to generate corresponding text data.
[0571] Input: Preprocessed sign language data
[0572] Specific behavior: The algorithm converts the sign "hello" into text "hello."
[0573] Output: Generated text data
[0574] Step 4: Generate audio data
[0575] The terminal sends the generated text data to a speech synthesis engine and converts it into voice data.
[0576] Input: Generated text data
[0577] Specific operation: The speech synthesis engine converts "hello" into speech.
[0578] Output: Generated audio data
[0579] Step 5: Outputting audio data
[0580] The device will play the translated audio through the speaker.
[0581] Input: Generated audio data
[0582] Specific action: "Hello" is output from the speaker.
[0583] Output: Translation played back as audio
[0584] Step 6: Capture and analyze emotions
[0585] The device uses a camera and microphone to analyze the user's emotions.
[0586] Input: Captured sign language data, as well as the user's facial expressions and tone of voice
[0587] Specific behavior: Facial expression analysis algorithm determines emotions.
[0588] Output: Parsed emotion information
[0589] Step 7: Adjusting the audio data
[0590] The device adjusts the tone and content of the voice based on the analyzed emotional information.
[0591] Input: Parsed emotion information
[0592] Specific behavior: If the user has a happy expression, play "Hello" in a bright tone.
[0593] Output: Modified audio data
[0594] These are the specific processing steps of this system. Based on the data input at each step, it is possible to provide optimal output according to the user's situation and emotions.
[0595] (Application example 2)
[0596] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0597] In autonomous vehicles, there is a need to address the safety and information shortages faced by the hearing- and visually impaired. Another challenge is to provide optimal information in response to the user's emotions, enabling safer and more comfortable use.
[0598] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing ambient sounds, means for converting the captured voice data into text data, means for analyzing the type, direction, and distance of the sound, means for displaying the analyzed information, means for recognizing the user's emotions, and means for adjusting the display information based on the emotion information. This enables hearing-impaired and visually impaired people to use autonomous vehicles safely and comfortably.
[0599] An "ambient sound capturing means" is a device or mechanism capable of detecting sounds in the environment and recording them as digital data.
[0600] A "means for converting captured audio data into text data" is a mechanism or software that analyzes the recorded audio signal and converts it into corresponding text information.
[0601] "Means for analyzing sound type, direction and distance" means an algorithm or mechanism for calculating and identifying the source, type and distance of a sound from captured audio data.
[0602] "Means for displaying analyzed information" refers to a display or other display device for visually presenting the results of the analysis.
[0603] The "means for recognizing user emotions" is an engine or software for analyzing the user's tone of voice and facial expressions to determine their emotional state.
[0604] The "means for adjusting displayed information based on emotional information" refers to a mechanism or program that dynamically changes the content or emphasis of displayed information in response to a recognized emotional state.
[0605] A "means for capturing sign language" is a camera or sensor that captures and records the user's sign language actions.
[0606] A "means for converting captured sign language data into text data" is an algorithm or system for analyzing the recorded sign language movements and converting them into corresponding written information.
[0607] The "means for converting the converted text data into voice" refers to a voice synthesis device or software for outputting the character data as voice.
[0608] "Means for capturing surrounding environmental data" refers to devices or systems that use sensors and cameras to capture the conditions inside and around the vehicle.
[0609] The "means for analyzing captured environmental data" is an algorithm or program for extracting and analyzing useful information from the acquired environmental data.
[0610] The "means for generating the analyzed information as a voice guide" is an engine or program for generating voice guidance based on the analysis results.
[0611] The "means for outputting the generated audio guide" refers to a device for informing the user of the generated audio guide using a speaker or the like.
[0612] The "means for adjusting the tone and content of the audio description based on emotional information" refers to a system or algorithm for dynamically changing the tone and content of the audio description depending on the user's emotional state.
[0613] The present invention relates to an assistance system for the hearing impaired and the visually impaired that is applied to an autonomous vehicle. Specific embodiments of the present invention will be described below.
[0614] System configuration
[0615] The system consists of the user's mobile device (hereafter referred to as the device) and a cloud-based server (hereafter referred to as the server). The device is equipped with a microphone, camera, speaker, display, various sensors, and an emotion recognition engine. This hardware is used to collect and analyze information about the user's surroundings and emotional state, and provide the necessary information.
[0616] Hearing-impaired features
[0617] Subtitling of ambient sounds
[0618] 1. Sound capture and processing
[0619] The device uses a microphone (e.g., general hardware) to capture ambient sounds. The captured audio data undergoes pre-processing such as noise reduction within the device.
[0620] The preprocessed voice data is converted into text in real time using voice recognition software (e.g., Google Speech Recognition API).
[0621] 2. Analysis of sound type, direction, and distance
[0622] The device identifies the type of sound from the converted text information and sound characteristics, and uses the arrangement of multiple microphones to calculate the direction and distance from which the sound is coming.
[0623] 3. Displaying Information
[0624] The analyzed information is displayed in text format on the device's display, allowing the user to understand the surrounding situation in real time.
[0625] 4. Emotion recognition and regulation
[0626] The device uses a camera (e.g., general hardware) or microphone to capture the user's facial expressions and tone of voice, which are then analyzed by an emotion recognition engine (e.g., the EmotionRecognition library).
[0627] The system adjusts the displayed text and notification method based on the analyzed emotion information. For example, if a car horn is honking from the right at a distance of 10 meters, and the user shows signs of surprise, the display will be emphasized.
[0628] Features for the visually impaired
[0629] Audio guidance of surrounding objects and situations
[0630] 1. Environmental data capture and processing
[0631] The device uses a camera and various sensors to capture data on the surrounding environment, which is then subjected to image recognition preprocessing to extract object contours and feature points.
[0632] 2. Object Recognition and Analysis
[0633] The pre-processed data is then used by AI algorithms (e.g., TensorFlow) to recognize surrounding objects, analyze their location and movement, and extract important information.
[0634] 3. Audio guide generation and output
[0635] The analyzed information is generated as audio guidance and converted into speech by a speech synthesis engine (e.g., gTTS library). The device then notifies the user of the generated audio guidance through the speaker.
[0636] 4. Emotion recognition and regulation
[0637] The device captures the user's emotions and adjusts the voice guidance based on the analyzed emotion information, for example, softening the tone of the voice guidance if the user is nervous.
[0638] Automatic sign language recognition and translation
[0639] 1. Sign Language Capture and Processing
[0640] The device uses a camera to capture the user's sign language in real time, and the captured video data undergoes pre-processing to extract hand shapes and movements.
[0641] 2. Sign Language Recognition and Translation
[0642] The preprocessed video data is analyzed using a sign language recognition algorithm (e.g., OpenCV or TensorFlow) to generate corresponding text information, which is then converted into speech data by a speech synthesis engine.
[0643] 3. Audio Output
[0644] The device plays the translated audio through a speaker, helping people who don't understand sign language to communicate with the user.
[0645] 4. Emotion Recognition and Translation Adjustment
[0646] The device analyzes the user's emotions along with the captured sign language data, and adjusts the expression of the text translation and the tone of the voice based on the emotional information. For example, if the user signs "hello" and has a happy expression, the tone of the voice output will be brighter.
[0647] Prompt Sentence Examples
[0648] "Can you hear the horn now?", "There is an obstacle ahead", "If the user is nervous, generate voice prompts in a softer tone."
[0649] As a result, information and notifications are appropriately adjusted according to the user's emotional information, enabling hearing-impaired and visually-impaired people to use self-driving vehicles safely and comfortably.
[0650] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0651] Step 1:
[0652] Audio data capture and preprocessing
[0653] Input: Ambient sound
[0654] Output: Preprocessed audio data
[0655] The device uses a microphone to capture surrounding sounds. The captured audio data undergoes pre-processing such as noise reduction. This processing makes the audio data easier to analyze. The data is also updated at regular intervals.
[0656] Step 2:
[0657] Converting audio data to text
[0658] Input: Preprocessed audio data
[0659] Output: Text data
[0660] The device converts the pre-processed voice data into text in real time using speech recognition software (e.g., Google Speech Recognition API). This process captures the voice as text information.
[0661] Step 3:
[0662] Analysis of sound type, direction, and distance
[0663] Input: Text data and audio data features
[0664] Output: Information about the type, direction, and distance of the sound
[0665] The device identifies the type of sound from the converted text information and features of the voice data, and calculates the direction and distance from which the sound is coming using a multi-microphone arrangement, identifying, for example, the sound of an ambulance siren or car horn.
[0666] Step 4:
[0667] Viewing analysis information
[0668] Input: Information about the type, direction, and distance of the sound
[0669] Output: Text information displayed
[0670] The device displays the analyzed information in text format on the display, allowing the user to understand the surrounding situation in real time. For example, it may display information such as "There is a horn honking 10 meters from the right."
[0671] Step 5:
[0672] Emotion recognition
[0673] Input: Camera video, audio data (if necessary)
[0674] Output: Emotional information
[0675] The device uses a camera (and optionally a microphone) to capture the user's facial expressions and tone of voice, which are then analyzed by an emotion recognition engine (e.g., the EmotionRecognition library). In the process, the user's emotional state (such as surprise, tension, or joy) is determined.
[0676] Step 6:
[0677] Adjusting display and notification information
[0678] Input: Emotion information, analysis information
[0679] Output: Adjusted text information or audio notification
[0680] The device adjusts the displayed text and notification method based on the acquired emotional information. For example, if the user expresses surprise, the device will emphasize the text or display it in a more visually appealing way. If a voice notification is required, the device will use a speech synthesis engine to adjust the tone of the notification.
[0681] Step 7:
[0682] Sign language capture and translation
[0683] Input: Sign language video data
[0684] Output: Text data, audio data
[0685] The device uses a camera to capture the user's sign language in real time. The captured video data is preprocessed to extract hand movements and shapes. The sign language is analyzed using a sign language recognition algorithm (e.g., OpenCV, TensorFlow), and corresponding text information is generated and output as voice data using a speech synthesis engine.
[0686] Step 8:
[0687] Environmental Data Capture and Analysis
[0688] Input: Environmental data (video and sensor data)
[0689] Output: Analysis information
[0690] The device uses cameras and sensors to capture data about the surrounding environment, which is then pre-processed using image recognition to extract object contours and feature points, and AI algorithms are used to identify the object type and location, generating key analytics.
[0691] Step 9:
[0692] Audio description generation and output
[0693] Input: Analysis information and emotion information
[0694] Output: Voice prompt
[0695] The device generates audio guidance based on the analyzed information and converts it into voice data using a speech synthesis engine (e.g., gTTS). The generated audio guidance is then played back to the user through the speaker. The tone and content of the audio guidance are adjusted according to the user's emotional state.
[0696] Through the above steps, the present invention assists hearing-impaired and visually-impaired people in safely and comfortably using autonomous vehicles.
[0697] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0698] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0699] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0700] [Second embodiment]
[0701] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0702] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0703] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0704] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0705] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0706] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0707] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0708] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0709] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0710] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0711] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0712] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0713] The present invention is an assistance system for the hearing impaired and the visually impaired, and is implemented in the following manner.
[0714] System configuration
[0715] This system consists of a user's mobile device (hereafter referred to as the device) and a cloud-based server (hereafter referred to as the server). The device is equipped with a microphone, camera, speaker, display, and various sensors, and uses this hardware to collect and analyze information about the user's surroundings and provide the necessary information.
[0716] Hearing-impaired features
[0717] Subtitling of ambient sounds
[0718] 1. Sound capture and processing
[0719] The device uses a microphone to constantly capture ambient sounds.
[0720] The captured audio data undergoes pre-processing such as noise reduction inside the device.
[0721] The preprocessed voice data is converted into text in real time by a voice recognition engine.
[0722] 2. Analysis of sound type, direction, and distance
[0723] The device identifies the type of sound from the converted text information and sound characteristics, such as "ambulance siren" or "human voice."
[0724] The device uses multiple microphone arrangements to calculate the direction and distance of the sound and notify the user.
[0725] 3. Displaying Information
[0726] The analyzed information is displayed in text format on the device's display, allowing the user to understand the surrounding situation in real time.
[0727] 4. Specific Examples
[0728] For example, if a car horn sounds near the user, the device will display "The horn is sounding from 10 meters to the right."
[0729] Automatic sign language recognition and translation
[0730] 1. Sign Language Capture and Processing
[0731] The device uses a camera to capture the user's sign language in real time.
[0732] The captured video data is pre-processed to extract the shape and movement of the hand.
[0733] 2. Sign Language Recognition and Translation
[0734] The pre-processed video data is analyzed by a sign language recognition algorithm to generate corresponding text information.
[0735] The generated text information is converted into voice data by a voice synthesis engine.
[0736] 3. Audio Output
[0737] The device plays the translated audio through a speaker, helping people who don't understand sign language to communicate with the user.
[0738] 4. Specific Examples
[0739] For example, if a user says "hello" in sign language, the terminal recognizes the sign and outputs "hello" aloud.
[0740] Features for the visually impaired
[0741] Audio guidance of surrounding objects and situations
[0742] 1. Environmental data capture and processing
[0743] The device uses a camera and various sensors to capture data about the surrounding environment.
[0744] The captured data is subjected to image recognition preprocessing to extract the contours and feature points of the object.
[0745] 2. Object Recognition and Analysis
[0746] The preprocessed data is then used by AI algorithms to recognize surrounding objects and situations.
[0747] The location and movement of recognized objects are analyzed to pick out important information.
[0748] 3. Audio guide generation and output
[0749] The analyzed information is generated as audio guide and converted into voice by a voice synthesis engine.
[0750] The terminal notifies the user of the generated audio guide through a speaker.
[0751] 4. Specific Examples
[0752] If the user is walking and there is an obstacle one meter ahead, the device will announce with a voice message, "There is an obstacle one meter ahead."
[0753] Starting and shutting down the system
[0754] 1. Initial Setup
[0755] During initial setup, the user selects the system mode (hearing impaired mode, visually impaired mode) that best suits their needs.
[0756] The server receives the user's configuration information and prepares the required services.
[0757] 2. Service launch
[0758] The device will activate the necessary sensors (microphone, camera, etc.) based on the initial settings and begin collecting data.
[0759] The server analyzes the data sent from the terminal and returns the necessary information in real time.
[0760] The terminal displays or outputs the analysis results to the user.
[0761] 3. Termination or Suspension of Service
[0762] If a user wishes to stop using the service, they can select the option to terminate or suspend the service in their device's settings screen.
[0763] The terminal receives the service termination command, shuts down all sensors, and cuts off communication with the server.
[0764] This system provides useful support to the hearing-impaired and visually impaired to overcome difficulties in daily life and live safely and independently.
[0765] The processing flow will be explained below.
[0766] Features for the hearing impaired: Subtitling of surrounding sounds
[0767] Step 1: Capture the sound
[0768] The device activates the microphone and captures ambient sounds at 0.1 second intervals.
[0769] Step 2: Preprocessing the audio data
[0770] The device performs noise reduction and normalization on the captured audio data.
[0771] Step 3: Speech to text
[0772] The device sends the preprocessed voice data to a voice recognition engine, which converts it into text data in real time.
[0773] Step 4: Identify the type of sound
[0774] The device extracts sound characteristics from the text data and identifies the type of sound based on them (e.g., "ambulance siren," "human voice," etc.).
[0775] Step 5: Analyze sound direction and distance
[0776] The device uses multiple microphone arrangements to calculate the direction and distance from which the sound is coming.
[0777] Step 6: Display subtitles
[0778] The device displays text information on the display, informing the user of the type, direction, and distance of the sound.
[0779] As a specific example, it displays "The sound of a horn is coming from 10 meters to the right."
[0780] Features for the visually impaired: Audio guide to surrounding objects and situations
[0781] Step 1: Capture environment data
[0782] The device will activate its camera and sensors to capture data on the surrounding environment every 0.5 seconds.
[0783] Step 2: Preprocessing the data
[0784] The device performs image recognition preprocessing on the captured video data to extract the contours and feature points of objects.
[0785] Step 3: Object Recognition
[0786] The device sends the pre-processed data to AI algorithms to recognize surrounding objects.
[0787] Step 4: Obtaining information about the object
[0788] The device analyzes the location and movement of recognized objects and picks out important information (e.g., obstacles ahead, approaching people).
[0789] Step 5: Generate audio guide
[0790] The device generates the analyzed information as a text message, sends it to a speech synthesis engine, and generates audio guidance.
[0791] Step 6: Audio Notifications
[0792] The terminal notifies the user of the generated audio guide through a speaker.
[0793] As a specific example, a voice message will be displayed saying, "There is an obstacle one meter ahead."
[0794] Starting and shutting down the system
[0795] Step 1: Initial Setup
[0796] The user configures the system and selects the mode that best suits their needs (e.g., hearing-impaired mode, visually-impaired mode).
[0797] The server receives the user's configuration information and prepares to provide the required services.
[0798] Step 2: Real-time analytics and notifications
[0799] Based on the selected mode, the device will activate the necessary sensors (microphone, camera, etc.) and collect data.
[0800] The server analyzes the data sent from the terminal and returns the necessary information in real time.
[0801] The device notifies the user of the analysis results in an appropriate format (subtitles, audio).
[0802] Step 3: Terminate or suspend service
[0803] If a user wants to stop using the system, they can select the option to terminate or suspend the service on their device's settings screen.
[0804] The terminal receives the service termination command, shuts down all sensors, and cuts off communication with the server.
[0805] Example 1
[0806] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0807] The purpose of this invention is to provide support for the hearing-impaired and visually-impaired to overcome the difficulties they face in daily life and live safely and independently. Specifically, the invention involves the development of a system that recognizes surrounding sounds, sign language, and environmental conditions in real time and provides the necessary information in an appropriate format.
[0808] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0809] In this invention, the server includes means for capturing ambient sounds, means for preprocessing the captured audio data, means for converting the preprocessed audio data into text data, means for analyzing the type, direction, and distance of sound, means for displaying the analyzed information, means for capturing sign language, means for preprocessing the captured sign language data, means for converting the preprocessed sign language data into text data, means for converting the converted text data into audio data and outputting it, means for capturing ambient environmental data, means for preprocessing the captured environmental data, means for analyzing the preprocessed environmental data to recognize objects, means for generating audio guidance based on the recognized object information, and means for outputting the generated audio guidance. This enables hearing-impaired and visually impaired people to grasp their surroundings in real time and live safely and independently.
[0810] The "means for capturing ambient sound" refers to a device or sensor function for collecting audio data from the user's surrounding environment.
[0811] "Means for preprocessing captured audio data" refers to processes or algorithms used to remove noise from collected audio data and prepare it for speech recognition.
[0812] The "means for converting preprocessed voice data into text data" refers to software or a system for converting voice data into text information using voice recognition technology.
[0813] "Means for analyzing the type, direction and distance of sound" refers to an algorithm or device that analyzes text data and sound feature data to identify the source, type, direction and distance of a sound.
[0814] "Means for displaying analyzed information" refers to a device or system that uses a display, screen, or the like to visually convey the analysis results to the user.
[0815] A "means for capturing sign language" is a camera or sensor that records the user's hand movements and shapes in real time.
[0816] "Means for pre-processing captured sign language data" refers to a process or system that appropriately processes captured video data for sign language recognition and extracts hand shapes and movements.
[0817] A "means for converting preprocessed sign language data into text data" is a technology or system that analyzes captured sign language movements and converts them into corresponding text information.
[0818] "Means for converting the converted text data into audio data and outputting it" refers to software or a device for converting character information into audio and reproducing it via a speaker or the like.
[0819] The "means for capturing surrounding environmental data" refers to a device or function for collecting information about the user's surrounding environment using a camera or sensor.
[0820] "Means for pre-processing captured environmental data" refers to processes or algorithms that process collected environmental data into a form suitable for image recognition.
[0821] "Means for analyzing preprocessed environmental data to recognize objects" refers to AI algorithms or systems that identify and recognize the shape and position of objects from environmental data.
[0822] The "means for generating audio guidance based on recognized object information" refers to a system or process for creating audio guidance to be notified to the user based on information about the identified object.
[0823] The "means for outputting the generated audio guide" refers to a device or system for conveying the generated audio guide to the user using a speaker or the like.
[0824] This invention is an assistance system for hearing-impaired and visually-impaired people to overcome difficulties in daily life, and is implemented as a system including a user's mobile terminal (hereinafter referred to as "terminal") and a cloud-based server (hereinafter referred to as "server"). The terminal is equipped with a microphone, camera, speaker, display, and various sensors. These hardware components are used to collect and analyze information about the user's surroundings and provide the necessary information.
[0825] Hearing-impaired features
[0826] Subtitling of ambient sounds
[0827] The device uses a microphone to capture ambient sounds and preprocesses the audio data. It then uses noise reduction technology to remove background noise and converts the audio into text using a speech recognition engine. The device uses multiple microphone arrangements to analyze the type, direction, and distance of the sound and displays the analyzed information on the display. For example, if a user is walking on the sidewalk and hears a car horn honking from the right at a distance of 10 meters, the device will display the message, "A car horn is honking from the right at a distance of 10 meters."
[0828] Automatic sign language recognition and translation
[0829] The device uses a camera to capture the user's sign language in real time and preprocesses the video data. After extracting the hand shape and movement, a sign language recognition algorithm generates corresponding text information. Next, a speech synthesis engine converts the text information into audio data and outputs it from the speaker. For example, if the user signs "hello," the device will recognize the sign and output "hello" aloud.
[0830] Features for the visually impaired
[0831] Audio guidance of surrounding objects and situations
[0832] The device uses cameras and various sensors to capture and preprocess data on the surrounding environment. It uses image recognition technology to extract the contours and feature points of objects, and uses AI algorithms to recognize surrounding objects and situations. It analyzes the location and movement of recognized objects and generates important information as audio guidance. This audio guidance is converted into voice by a speech synthesis engine and notified to the user through the speaker. For example, if a user is walking and there is an obstacle one meter ahead, the device will announce, "There is an obstacle one meter ahead."
[0833] Starting and shutting down the system
[0834] During the initial setup, the user selects the system mode (deaf mode, visually impaired mode) that best suits their needs. The server receives the user's configuration information and prepares the required services. The device activates the required sensors (microphone, camera, etc.) based on the initial setup and begins collecting data. The server analyzes the data sent from the device and returns the required information in real time. If the user wants to terminate or pause the service, they select the option to terminate or pause the service on the device's settings screen. The device receives the service termination command, stops all sensors, and cuts off communication with the server.
[0835] Examples of concrete examples and prompts
[0836] Example prompt for speech recognition: "Transcribe the speech that says 'hello' to text."
[0837] Example prompt for sign language recognition: "Translate this sign language video into text."
[0838] This provides useful support to hearing-impaired and visually impaired people to overcome difficulties in daily life and live safely and independently.
[0839] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0840] Hearing-impaired features
[0841] Subtitling of ambient sounds
[0842] Step 1: Capture the sound
[0843] Subject: Device
[0844] How it works: The device uses a microphone to capture ambient sound, taking ambient audio data as input and raw captured data as output.
[0845] Input and Output: The input is the ambient audio data, and the output is the captured audio data.
[0846] Step 2: Pre-processing the sound
[0847] Subject: Device
[0848] Specific operation: The device preprocesses the captured audio data and removes noise. The input is the captured audio data, and the output is the noise-removed audio data.
[0849] Input and Output: The input is the captured audio data, and the output is the pre-processed audio data.
[0850] Step 3: Voice Recognition
[0851] Subject: Device
[0852] Specific operation: The device sends the preprocessed voice data to the voice recognition engine and converts it into text data. The preprocessed voice data is input, and the converted text data is obtained as output.
[0853] Input and Output: The input is preprocessed audio data, and the output is text data.
[0854] Step 4: Analyzing the sound information
[0855] Subject: Device
[0856] Specific operation: The device analyzes the text data and sound feature data to identify the type, direction, and distance of the sound. The text data and sound feature data are input, and the analysis results are obtained as output.
[0857] Input and output: The input is text data and sound feature data, and the output is analyzed sound information.
[0858] Step 5: Viewing information
[0859] Subject: Device
[0860] Specific operation: The terminal displays the analyzed information in text format on the display. The analyzed sound information is input and displayed on the display as output.
[0861] Input and output: The input is the analyzed sound information, and the output is the text information displayed on the screen.
[0862] Automatic sign language recognition and translation
[0863] Step 1: Capture sign language
[0864] Subject: Device
[0865] Specific operation: The device uses a camera to capture the user's sign language actions in real time. The sign language actions are input and the captured video data is obtained as output.
[0866] Input and Output: The input is sign language actions, and the output is captured video data.
[0867] Step 2: Preprocessing the video data
[0868] Subject: Device
[0869] Specific operation: The device preprocesses the captured video data and extracts the hand shape and movement. The captured video data is input and the preprocessed video data is output.
[0870] Input and Output: The input is the captured video data, and the output is the pre-processed video data.
[0871] Step 3: Sign Language Recognition and Text Conversion
[0872] Subject: Device
[0873] Specific operation: The device sends the preprocessed video data to a sign language recognition algorithm to generate corresponding text data. The preprocessed video data is input, and the generated text data is output.
[0874] Input and Output: The input is the preprocessed video data, and the output is the generated text data.
[0875] Step 4: Speech synthesis and output
[0876] Subject: Device
[0877] Specific operation: The device sends the generated text data to a speech synthesis engine, converts it into voice data, and outputs the converted voice data to the user through the speaker. Text data is input, and voice data is obtained as output.
[0878] Input and output: The input is the generated text data, and the output is the audio output from the speaker.
[0879] Features for the visually impaired
[0880] Audio guidance of surrounding objects and situations
[0881] Step 1: Capture environment data
[0882] Subject: Device
[0883] Specific operation: The device uses a camera and various sensors to capture surrounding environmental data. The surrounding environmental information is input, and the captured environmental data is output.
[0884] Input and Output: The input is the surrounding environment information, and the output is the captured environment data.
[0885] Step 2: Preprocessing the image data
[0886] Subject: Device
[0887] Specific operation: The device preprocesses the captured environmental data and extracts object contours and feature points. The captured environmental data is input, and the preprocessed environmental data is output.
[0888] Input and Output: The input is the captured environmental data, and the output is the preprocessed environmental data.
[0889] Step 3: Object Recognition
[0890] Subject: Device
[0891] How it works: The device analyzes the preprocessed data using AI algorithms to recognize surrounding objects and situations. The preprocessed data is input, and the recognized object information is output.
[0892] Input and Output: The input is the preprocessed data, and the output is the recognized object information.
[0893] Step 4: Location Analysis
[0894] Subject: Device
[0895] Specific operation: The device analyzes the location and movement of the recognized object and extracts important information. The recognized object information is input, and the location and movement analysis results are obtained as output.
[0896] Input and Output: The input is the recognized object information, and the output is the analysis result of the position and movement.
[0897] Step 5: Audio guide generation and output
[0898] Subject: Device
[0899] Specific operation: The device generates the analyzed information as an audio guide and converts it into speech using a speech synthesis engine. The converted audio guide is then notified to the user through the speaker. The analyzed information is input, and the audio guide is obtained as output.
[0900] Input and Output: The input is the parsed information, and the output is the audio guide from the speaker.
[0901] Starting and shutting down the system
[0902] Step 1: Initial Setup
[0903] Subject: User
[0904] Specific operation: The user selects the system mode (deaf mode, visually impaired mode) according to their needs in the initial setting. The user's needs are input, and the setting information is sent to the server as output.
[0905] Input and Output: Input is user needs, output is configuration information.
[0906] Step 2: Start the service
[0907] Subject: Server
[0908] Specific operation: The server receives the user's configuration information and prepares the necessary services. The device activates the necessary sensors (microphone, camera, etc.) based on the initial settings and begins collecting data. The configuration information is received as input and the service is started as output.
[0909] Input and Output: Input is the configuration information, output is the started service.
[0910] Step 3: Data collection and analysis
[0911] Subject: Device
[0912] Specific operation: The terminal starts collecting data, and the server analyzes the data sent from the terminal. The collected data is input, and the analysis results are obtained as output.
[0913] Input and Output: Input is the collected data and output is the analysis result.
[0914] Step 4: Notification of information
[0915] Subject: Device
[0916] Specific operation: The device displays or outputs the analysis results to the user. The analysis results are input and notified to the user as output.
[0917] Input and output: The input is the analysis result, and the output is the notification to the user.
[0918] Step 5: Terminate or suspend service
[0919] Subject: User
[0920] Specific behavior: When a user wants to stop using a service, they select the option to terminate or pause the service in the device settings screen. The user's operation is the input, and the service is terminated or paused as the output.
[0921] Input and Output: Input is a user action, output is a service that has been terminated or suspended.
[0922] (Application example 1)
[0923] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0924] The purpose of this invention is to provide a means to resolve the lack of information and communication difficulties faced by hearing-impaired and visually-impaired people when shopping comfortably and safely in a brick-and-mortar store. Specifically, the objective is to provide a system that supports people with disabilities to act independently by acquiring in-store announcements, conversations with store staff, and product information in real time and conveying this information appropriately to the user.
[0925] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0926] In this invention, the server includes means for capturing ambient sounds, means for converting the captured audio data into text data, means for analyzing the type, direction, and distance of the sound, means for displaying the analyzed information, means for displaying the captured text data on a smart device display, means for capturing sign language, means for converting the captured sign language data into text data, means for converting the converted text data into audio, means for capturing ambient environmental data, means for analyzing the captured environmental data, means for generating the analyzed information as an audio guide, and means for outputting the generated audio guide. This allows hearing-impaired people to understand in-store announcements and conversations in real time through subtitles, and makes it easy to convert sign language communication into audio. Furthermore, visually impaired people can receive information about surrounding objects and products as an audio guide, enabling safe and efficient shopping.
[0927] "Means for capturing ambient sounds" is a function that collects environmental sounds using the microphone of the smart device.
[0928] "Means for converting captured voice data into text data" refers to a function that transcribes collected voice data using voice recognition technology.
[0929] "Means for analyzing the type, direction and distance of sound" is a function that analyzes collected audio data to identify the type, location and distance of the sound source.
[0930] "Means for displaying analyzed information" refers to a function that visually displays the analysis results on the display of a smart device.
[0931] The "means for displaying captured text data on the display of the smart device" is a function for displaying text obtained by speech recognition on the screen of the smart device.
[0932] A "means for capturing sign language" is a function that uses the camera of a smart device to capture sign language movements.
[0933] The "means for converting captured sign language data into text data" is a function that analyzes the video of the sign language and converts it into text information.
[0934] The "means for converting the converted text data into voice" is a function for outputting the text data as voice using voice synthesis technology.
[0935] "Means for capturing surrounding environmental data" refers to the ability to collect images and data of the environment using the smart device's camera and sensors.
[0936] "Means for analyzing captured environmental data" refers to a function that analyzes collected video and data to recognize objects and environments.
[0937] "Means for generating analyzed information as audio guide" is a function for creating audio guide based on the analysis results.
[0938] "Means for outputting the generated audio guide" is a function that transmits the generated audio guide to the user through the speaker of the smart device.
[0939] This invention is a system that helps hearing- and visually-impaired people enjoy shopping comfortably and safely in physical stores. The system consists of a user's smart device and a cloud-based server. Specifically, the system collects and analyzes surrounding information using the smart device's microphone, camera, display, speaker, and various sensors. This collected and analyzed information is then provided to the user appropriately.
[0940] System configuration
[0941] 1. Subtitling of ambient sounds
[0942] Capture method: Capture surrounding sounds using the microphone on your smart device.
[0943] Voice data conversion: The captured voice data is converted into text data using voice recognition technology. Specifically, Google Cloud's Speech-to-Text API is used.
[0944] Analysis means: Analyzes the type, direction and distance of sound and provides important information to the user.
[0945] Display method: The analyzed information is displayed in real time as subtitles on the smart device display, allowing hearing-impaired people to understand in-store announcements and conversations.
[0946] 2. Automatic Sign Language Recognition and Translation
[0947] Capture method: Sign language is captured using the smart device's camera.
[0948] Sign language data conversion: The captured sign language data is converted into text data using image processing techniques and sign language recognition algorithms.
[0949] Speech conversion method: The converted text data is converted into speech using speech synthesis technology (Google Cloud's Text-to-Speech API) and played back through the smart device's speaker, allowing hearing-impaired people to easily communicate with store staff and other customers.
[0950] 3. Audio guide to surrounding objects and situations
[0951] Capture method: Capture surrounding environmental data using the smart device's camera and various sensors.
[0952] Environmental data analysis: The captured data is analyzed using image recognition technology (Google Vision API) to identify objects and obtain their location information.
[0953] Audio guide generation: Audio guides are generated based on the analyzed information and converted into audio using Google Cloud's Text-to-Speech API.
[0954] Output means: The generated audio guide is output from the smart device's speaker to inform visually impaired people of their surroundings.
[0955] Specific examples
[0956] For example, when a user wears a smart device and enters a physical store, the system captures in-store announcements with a microphone and displays information such as "Today's special sale is XX" in real time as subtitles on the display. Also, if the user expresses "Where is it?" in sign language, the system captures the sign, converts it into text, and relays it to the store clerk as "Where is it?" Furthermore, when a visually impaired person stands in front of a shelf, the camera recognizes the product label and provides an audio guide such as "There is a fruit shelf ahead." This guide is provided through the smart device's speaker.
[0957] Prompt Sentence Examples
[0958] "Please provide real-time subtitles for in-store announcements."
[0959] "Please recognize the sign language, convert it into speech, and communicate it to the store clerk."
[0960] "Recognize surrounding objects and provide voice guidance"
[0961] As can be seen, this system provides a useful aid for the hearing and visually impaired to enjoy shopping independently.
[0962] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0963] Step 1:
[0964] Capture Processing
[0965] Using sensors such as the device's microphone and camera, surrounding audio and video data is captured in real time.
[0966] Input: Ambient audio, video such as sign language, and environmental data.
[0967] How it works: A microphone captures audio data, and a camera captures sign language and footage of the environment.
[0968] Step 2:
[0969] Pretreatment
[0970] Noise reduction and image processing are performed on the captured audio and video data, converting it into a format that is easy to analyze.
[0971] Input: Raw captured data (audio data, video data).
[0972] Data processing: Noise reduction of audio data, contour extraction of video data.
[0973] Output: Preprocessed audio data, preprocessed video data.
[0974] Operation: A noise reduction algorithm is used to remove noise from the audio data, and a contour extraction filter is applied to the video data.
[0975] Step 3:
[0976] Data transformation and analysis
[0977] The preprocessed data is analyzed using a voice recognition engine or image recognition algorithm.
[0978] Input: Preprocessed audio data, preprocessed video data.
[0979] Data calculation: Voice data is converted into text using a speech recognition algorithm (Google Cloud Speech-to-Text API), and sign language is analyzed and converted into text data using an image recognition algorithm.
[0980] Output: Text data, object recognition data.
[0981] How it works: It calls the Google Cloud Speech-to-Text API to convert speech to text, then uses a sign language recognition algorithm to convert sign language to text.
[0982] Step 4:
[0983] Displaying and Speech Output of Data
[0984] The converted text data is displayed on the smart device's display, and if necessary, a voice synthesis engine is used to output the text as speech.
[0985] Input: Converted text data, object recognition data.
[0986] Data calculation: Text data is displayed as a user interface, and speech is generated using a speech synthesis algorithm (Google Cloud Text-to-Speech API).
[0987] Output: Text displayed on the screen, audio guidance output from the speaker.
[0988] How it works: Text data is displayed as subtitles on the smart device's display, and audio data is generated by calling the Google Cloud Text-to-Speech API and output through the speaker.
[0989] Step 5:
[0990] User Feedback Processing
[0991] If there is additional input from the user, that input is passed back to the capture process for further analysis and interaction.
[0992] Input: New input from the user (speech, sign language, changes in the environment).
[0993] Data processing: Preprocess and analyze again as necessary.
[0994] Output: Updated visual and audio guides.
[0995] What it does: Captures user feedback, then analyzes it again and displays / speech outputs it.
[0996] Through these steps, hearing and visually impaired people can have a more comfortable and safer shopping experience in physical stores.
[0997] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0998] The present invention is an assistive system for the hearing and visually impaired, which has the ability to recognize a user's emotions and use that information to adjust output. Detailed description of the preferred embodiments of the present invention is provided below.
[0999] System configuration
[1000] This system consists of a user's mobile device (hereafter referred to as the device) and a cloud-based server (hereafter referred to as the server). The device is equipped with a microphone, camera, speaker, display, various sensors, and an emotion recognition engine, and uses this hardware to collect and analyze information about the user's surroundings and emotional state, and provide the necessary information.
[1001] Hearing-impaired features
[1002] Subtitling of ambient sounds
[1003] 1. Sound capture and processing
[1004] The device uses a microphone to capture ambient sounds at 0.1 second intervals.
[1005] The captured audio data undergoes pre-processing such as noise reduction inside the device.
[1006] The preprocessed voice data is converted into text in real time by a voice recognition engine.
[1007] 2. Analysis of sound type, direction, and distance
[1008] The device identifies the type of sound from the converted text information and sound characteristics (e.g., "ambulance siren," "human voice," etc.).
[1009] The device uses multiple microphone arrangements to calculate the direction and distance from which the sound is coming.
[1010] 3. Displaying Information
[1011] The analyzed information is displayed in text format on the device's display, allowing the user to understand the surrounding situation in real time.
[1012] 4. Emotion recognition and regulation
[1013] The device uses a camera and microphone to capture the user's facial expressions and tone of voice, which are then analyzed by an emotion engine.
[1014] The text displayed and notification method are adjusted based on the analyzed emotional information.
[1015] For example, when displaying "The sound of a horn is honking from 10 meters to the right," if the user appears surprised, the display will be emphasized.
[1016] Features for the visually impaired
[1017] Audio guidance of surrounding objects and situations
[1018] 1. Environmental data capture and processing
[1019] The device uses a camera and various sensors to capture data about the surrounding environment at 0.5-second intervals.
[1020] The captured data is subjected to image recognition preprocessing to extract the contours and feature points of the object.
[1021] 2. Object Recognition and Analysis
[1022] The pre-processed data is then used by AI algorithms to recognize surrounding objects.
[1023] The location and movement of recognized objects are analyzed to pick out important information.
[1024] 3. Audio guide generation and output
[1025] The analyzed information is generated as audio guide and converted into voice by a voice synthesis engine.
[1026] The terminal notifies the user of the generated audio guide through a speaker.
[1027] 4. Emotion recognition and regulation
[1028] The device captures the user's emotions and adjusts the audio guidance based on the analyzed emotion information.
[1029] As a specific example, if the user is nervous, the tone of the voice guide is softened.
[1030] Automatic sign language recognition and translation
[1031] 1. Sign Language Capture and Processing
[1032] The device uses a camera to capture the user's sign language in real time.
[1033] The captured video data is pre-processed to extract the shape and movement of the hand.
[1034] 2. Sign Language Recognition and Translation
[1035] The pre-processed video data is analyzed by a sign language recognition algorithm to generate corresponding text information.
[1036] The generated text information is converted into voice data by a voice synthesis engine.
[1037] 3. Audio Output
[1038] The device plays the translated audio through a speaker, helping people who don't understand sign language to communicate with the user.
[1039] 4. Emotion Recognition and Translation Adjustment
[1040] The device analyzes the user's emotions along with the captured sign language data.
[1041] The device adjusts the expression and tone of the text translation based on the emotional information.
[1042] As a specific example, if the user says "hello" in sign language and has a happy expression, the tone of the voice output is brightened.
[1043] Starting and shutting down the system
[1044] 1. Initial Setup
[1045] The user configures the system and selects the mode that best suits their needs (e.g., hearing-impaired mode, visually-impaired mode).
[1046] The server receives the user's configuration information and prepares to provide the required services.
[1047] 2. Real-time analysis and notifications
[1048] Based on the selected mode, the device will activate the necessary sensors (microphone, camera, etc.) and collect data.
[1049] The server analyzes the data sent from the terminal and returns the necessary information in real time.
[1050] The device notifies the user of the analysis results in an appropriate format (subtitles, audio), adjusting it according to the emotional information.
[1051] 3. Termination or Suspension of Service
[1052] If a user wants to stop using the system, they can select the option to terminate or suspend the service on their device's settings screen.
[1053] The terminal receives the service termination command, shuts down all sensors, and cuts off communication with the server.
[1054] This system provides useful support for hearing- and visually-impaired people to overcome difficulties in daily life and live safely and independently. Furthermore, by adjusting its response according to the user's emotions, it is possible to provide support that meets more individual needs.
[1055] The processing flow will be explained below.
[1056] Features for the hearing impaired: Subtitling of surrounding sounds
[1057] Step 1: Capture the sound
[1058] The device activates the microphone and captures ambient sounds at 0.1 second intervals.
[1059] Step 2: Preprocessing the audio data
[1060] The device performs noise reduction and normalization on the captured audio data.
[1061] Step 3: Speech to text
[1062] The device sends the preprocessed voice data to a voice recognition engine, which converts it into text data in real time.
[1063] Step 4: Identify the type of sound
[1064] The device extracts sound characteristics from the text data and identifies the type of sound based on them (e.g., "ambulance siren," "human voice," etc.).
[1065] Step 5: Analyze sound direction and distance
[1066] The device uses multiple microphone arrangements to calculate the direction and distance from which the sound is coming.
[1067] Step 6: Capturing Emotions
[1068] The device uses a camera and microphone to capture the user's facial expressions and tone of voice.
[1069] Step 7: Analyze the sentiment data
[1070] The device sends the captured emotion data to an emotion engine to analyze the user's emotions.
[1071] Step 8: Display subtitles
[1072] Based on the analyzed voice information and emotional data, the device displays text information on the screen, conveying the type, direction, and distance of the sound to the user in a format that is optimal for them.
[1073] For example, when displaying the message "A horn is honking from 10 meters to the right," if the user appears surprised, the warning message will be emphasized.
[1074] Features for the visually impaired: Audio guide to surrounding objects and situations
[1075] Step 1: Capture environment data
[1076] The device will activate its camera and sensors to capture data on the surrounding environment every 0.5 seconds.
[1077] Step 2: Preprocessing the data
[1078] The device performs image recognition preprocessing on the captured video data to extract the contours and feature points of objects.
[1079] Step 3: Object Recognition
[1080] The device sends the pre-processed data to AI algorithms to recognize surrounding objects.
[1081] Step 4: Obtaining information about the object
[1082] The device analyzes the location and movement of recognized objects and picks out important information (e.g., obstacles ahead, approaching people).
[1083] Step 5: Capturing emotions
[1084] The device uses a camera and microphone to capture the user's facial expressions and tone of voice.
[1085] Step 6: Analyze the sentiment data
[1086] The device sends the captured emotion data to an emotion engine to analyze the user's emotions.
[1087] Step 7: Generate audio guide
[1088] The device generates audio guidance based on the analyzed environmental information and emotional data.
[1089] For example, if the user is nervous, the tone of the voice prompt will be softened and the user will be notified, "There is an obstacle 1 meter ahead."
[1090] Step 8: Audio Notifications
[1091] The terminal notifies the user of the generated audio guide through a speaker.
[1092] Automatic sign language recognition and translation
[1093] Step 1: Capture sign language
[1094] The device uses a camera to capture the user's sign language in real time.
[1095] Step 2: Preprocessing the sign language data
[1096] The device performs pre-processing on the captured video data to extract the shape and movement of the hand.
[1097] Step 3: Sign Language Recognition
[1098] The device sends the preprocessed data to a sign language recognition algorithm to generate corresponding text information.
[1099] Step 4: Capturing emotions
[1100] The device captures the user's facial expressions with a camera along with the captured video data, and obtains emotion data.
[1101] Step 5: Analyze the sentiment data
[1102] The device analyzes the emotional data using an emotion engine and determines the user's emotions.
[1103] Step 6: Adjusting the text data
[1104] The device adjusts the text information generated from the sign language as needed based on emotional data.
[1105] Step 7: Converting text data into speech
[1106] The device sends the adjusted text data to a speech synthesis engine and converts it into voice data.
[1107] Step 8: Audio Output
[1108] The device plays audio data through a speaker, helping people who do not understand sign language to communicate with the user.
[1109] As a specific example, if a user who signs "hello" looks happy, the tone of the voice output is made brighter.
[1110] Example 2
[1111] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1112] The present invention aims to solve the problem of the lack of a function to adjust information according to the user's emotions in the conventional technology for life support systems for the hearing-impaired and visually-impaired. Specifically, the present invention aims to support a more comfortable and safe life by capturing surrounding sounds and environmental information in real time and providing it to the user, while adjusting the output content and format based on the user's emotions.
[1113] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1114] In this invention, the server includes means for capturing ambient sounds, means for preprocessing the captured audio data, means for converting the preprocessed audio data into text data, means for analyzing the type, direction, and distance of the sound, means for displaying the analyzed information, and means for recognizing the user's emotions and adjusting the display information. This allows the user to grasp the surrounding situation in real time and receive appropriate notifications according to the user's emotions. Furthermore, by including means for capturing sign language, means for preprocessing the captured sign language data, means for converting the preprocessed sign language data into text data, means for converting the converted text data into audio data and outputting it, and means for recognizing the user's emotions and adjusting the audio output, communication with sign language users can be facilitated. Furthermore, by including means for capturing ambient environmental data, means for preprocessing the captured environmental data, means for analyzing the preprocessed environmental data, means for generating the analyzed information as audio guidance, means for outputting the generated audio guidance, and means for recognizing the user's emotions and adjusting the audio guidance, it becomes easier for visually impaired people to grasp the surrounding situation and appropriate guidance according to their emotions can be provided.
[1115] The "means for capturing ambient sound" refers to a means for capturing ambient sound using an acoustic sensor such as a microphone and acquiring the audio data.
[1116] The "means for pre-processing captured audio data" refers to a means for performing processes such as noise reduction and filtering on the acquired audio data to make the audio signal clearer.
[1117] The "means for converting preprocessed voice data into text data" refers to a means for analyzing preprocessed voice data using a voice recognition technique and generating corresponding text data.
[1118] "Means for analyzing the type, direction and distance of a sound" refers to means for analyzing the characteristics of pre-processed and text-converted audio data to identify what the sound is (e.g., a car horn, a human voice, etc.), the direction from which it is coming, and how far away it is.
[1119] The "means for displaying analyzed information" refers to a means for displaying the analyzed audio information on a display or monitor so that the user can visually confirm the information.
[1120] "Means for recognizing the user's emotions and adjusting the displayed information" refers to means for analyzing the user's facial expressions and tone of voice using a camera or microphone, determining the user's emotions, and then adjusting the format and content of the displayed information.
[1121] The "means for capturing sign language" refers to a means for capturing a user's sign language actions in real time using a video capture device such as a camera, and acquiring the data.
[1122] The "means for preprocessing captured sign language data" refers to a means for processing the acquired sign language video data to make it easier to analyze, and extracting features such as hand shape and movement.
[1123] The "means for converting preprocessed sign language data into text data" refers to a means for analyzing the preprocessed sign language data using a sign language recognition algorithm and generating text data representing the meaning of the sign language.
[1124] "Means for converting the converted text data into voice data and outputting it" refers to means for generating voice data based on the text data using voice synthesis technology and notifying the user of the voice data through a speaker or the like.
[1125] The "means for capturing surrounding environmental data" refers to a means for obtaining visual and other data about surrounding objects and the environment using a camera or various sensors.
[1126] "Means for preprocessing captured environmental data" refers to means for performing image processing and data analysis on the acquired environmental data, extracting the contours and feature points of objects, and making the data easier to analyze.
[1127] "Means for analyzing preprocessed environmental data" refers to means for analyzing preprocessed environmental data using AI algorithms and recognizing objects and situations.
[1128] The "means for generating analyzed information as audio guidance" refers to a means for generating information to be provided to the user as audio guidance based on the analysis results.
[1129] The "means for outputting the generated audio guide" refers to a means for outputting the generated audio guide as audio using speech synthesis technology and notifying the user through a speaker.
[1130] The "means for recognizing the user's emotions and adjusting the audio guidance" refers to a means for analyzing the user's facial expressions and tone of voice, and adjusting the tone and content of the audio guidance based on the determined emotional information.
[1131] The present invention is an assistance system for the hearing-impaired and visually-impaired, which has the ability to recognize a user's emotions and adjust output using that information. Detailed embodiments for implementing the present invention will be described below.
[1132] System configuration
[1133] This system consists of a cloud-based server (hereafter referred to as the "server") and a mobile device (hereafter referred to as the "device") carried by the user. The device is equipped with a microphone, camera, speaker, display, various sensors, and an emotion recognition engine.
[1134] Embodiments for the hearing impaired
[1135] Subtitling of ambient sounds
[1136] The device uses a microphone to capture ambient sounds at 0.1-second intervals. The captured voice data undergoes preprocessing such as noise reduction within the device, and is then converted into text data by a voice recognition engine. This text data is then used to analyze the type, direction, and distance of the sound, and displayed on the device's display. The device also uses a camera and microphone to capture the user's emotions, which are analyzed by an emotion engine. The device adjusts the text displayed and notification method based on this emotional information.
[1137] As a specific example of how this works, if the user is surprised by the information "There is a horn honking 10 meters to the right," it will be displayed with an emphasis of "Caution!"
[1138] Embodiments for the visually impaired
[1139] Audio guidance of surrounding objects and situations
[1140] The device captures data on the surrounding environment every 0.5 seconds using a camera and various sensors. The captured data undergoes image recognition preprocessing to extract object contours and feature points. This data is analyzed using an AI algorithm to identify the location and movement of recognized objects and generate audio guidance. This audio guidance is output from the speaker via a speech synthesis engine. The device also captures the user's emotions and adjusts the audio guidance based on the analyzed emotional information.
[1141] As a specific example of how it works, if the user is nervous about an approaching car, the system will notify them in a soft tone, "There is a car approaching 10 meters ahead, please proceed slowly."
[1142] Embodiments of automatic sign language recognition and translation
[1143] The device uses a camera to capture sign language in real time and preprocesses the hand shape and movements. The preprocessed video data is then analyzed using a sign language recognition algorithm to generate corresponding text data. This text data is converted into audio data by a speech synthesis engine and output through the speaker. The device also captures the user's emotions and adjusts the tone of the voice based on the emotional information.
[1144] As a specific example of operation, if the user signs "hello" and looks happy, "hello" is output as voice in a bright tone.
[1145] Examples of prompt statements
[1146] An example prompt to be input to a generative AI model would be something like this:
[1147] 1. Prompt for the deaf:
[1148] "Capture ambient sounds and subtitle them in real time. Analyze the type, direction, and distance of the sound and display it in text format."
[1149] 2. Prompt for the visually impaired:
[1150] "Use cameras and sensors to capture data about the surrounding environment and output recognized object information as audio guidance."
[1151] 3. Prompt for sign language recognition:
[1152] "Use your camera to capture sign language in real time and translate it into text or speech."
[1153] As described above, this system adjusts the way information is displayed and notified according to the user's emotions, providing assistance that meets more individual needs.
[1154] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1155] Processing steps for the hearing impaired
[1156] Subtitling of ambient sounds
[1157] Step 1: Capture the sound
[1158] The device uses a microphone to capture ambient sounds at 0.1 second intervals.
[1159] Input: Ambient sound
[1160] Specific operation: The device picks up surrounding environmental sounds and voices using the microphone.
[1161] Output: Captured audio data
[1162] Step 2: Preprocessing the audio data
[1163] The device performs pre-processing such as noise reduction and filtering on the captured audio data.
[1164] Input: Captured audio data
[1165] What it does: The device removes background noise and emphasizes the main voice.
[1166] Output: Preprocessed audio data
[1167] Step 3: Convert audio data to text
[1168] The device sends the preprocessed voice data to a voice recognition engine, which converts it into text data in real time.
[1169] Input: Preprocessed audio data
[1170] Specific operation: The voice recognition engine converts spoken content such as "Hello, it's a nice day today" into text.
[1171] Output: Text data
[1172] Step 4: Analyze the type, direction and distance of the sound
[1173] The device uses the characteristics of text and voice data to identify the type of sound and calculates direction and distance using multiple microphone arrangements.
[1174] Input: Text data, audio data features
[1175] Specific operation: The device uses an analysis algorithm to classify sounds such as "car horns" and "human voices" and calculates direction and distance.
[1176] Output: Analyzed sound type, direction and distance
[1177] Step 5: Viewing information
[1178] The device displays the analyzed voice information on the display.
[1179] Input: Analyzed audio information
[1180] Specific operation: The device displays information such as "A horn is honking 10 meters to the right" on the screen.
[1181] Output: Information displayed on the display
[1182] Step 6: Capture and analyze emotions
[1183] The device uses a camera and microphone to capture the user's facial expressions and tone of voice, which are then analyzed by an emotion engine.
[1184] Input: User facial expressions and tone of voice
[1185] What it does: The device uses a facial expression analysis algorithm to detect emotions such as surprise or anxiety.
[1186] Output: Parsed emotion information
[1187] Step 7: Adjust notification methods
[1188] The device adjusts the text displayed and notification method based on the analyzed emotional information.
[1189] Input: Parsed emotion information
[1190] Specific behavior: If the user is startled, the message "Sound of horn 10 meters to the right -- Caution!" will be displayed in emphasis.
[1191] Output: Adjusted display information
[1192] Processing steps for the visually impaired
[1193] Audio guidance of surrounding objects and situations
[1194] Step 1: Capture environment data
[1195] The device uses a camera and various sensors to capture data on the surrounding environment every 0.5 seconds.
[1196] Input: Surrounding environment
[1197] Specific operation: The camera captures images of the road, pedestrians, cars, etc. ahead, and the sensor measures distance and movement.
[1198] Output: Captured environmental data
[1199] Step 2: Preprocessing environmental data
[1200] The device performs image processing on the captured environmental data to extract the contours and feature points of objects.
[1201] Input: Captured environmental data
[1202] How it works: Image processing algorithms extract the shape and features of objects.
[1203] Output: Preprocessed environmental data
[1204] Step 3: Object Recognition and Analysis
[1205] The device sends pre-processed environmental data to AI algorithms that recognize surrounding objects and analyze their position and movement.
[1206] Input: Preprocessed environmental data
[1207] Specific operation: The AI recognizes things like "there are two people ahead" or "a car is approaching."
[1208] Output: Position and movement of analyzed objects
[1209] Step 4: Generate audio guide
[1210] The device generates audio guidance based on the analyzed information and converts it into audio data using a speech synthesis engine.
[1211] Input: Analyzed object information
[1212] Specific operation: Generates voice guidance such as "There is a step 10 meters ahead."
[1213] Output: The generated audio guide
[1214] Step 5: Outputting voice guidance
[1215] The device will announce the generated audio guide through the speaker.
[1216] Input: Generated audio guide
[1217] Specific operation: The device will announce "There is a step 10 meters ahead" via voice.
[1218] Output: Guide output as audio
[1219] Step 6: Capture and analyze emotions
[1220] The device uses a camera and microphone to capture the user's emotions and analyzes them using an emotion engine.
[1221] Input: User facial expressions and tone of voice
[1222] Specific operation: The device analyzes facial expressions to determine whether the user is anxious or nervous.
[1223] Output: Parsed emotion information
[1224] Step 7: Adjust the audio prompts
[1225] The device adjusts the audio guidance based on the analyzed emotional information.
[1226] Input: Parsed emotion information
[1227] Specific behavior: If the user is nervous, a soft tone will be displayed saying, "There is a step ahead, please walk slowly."
[1228] Output: Adjusted voice prompts
[1229] Processing steps for automatic sign language recognition and translation
[1230] Step 1: Capture sign language
[1231] The device uses a camera to capture the user's sign language in real time.
[1232] Input: User's sign language actions
[1233] Specific operation: The camera captures the user's hand movements.
[1234] Output: Captured sign language data
[1235] Step 2: Preprocessing the sign language data
[1236] The device preprocesses the captured sign language data to make it easier to analyze, extracting the shape and movement of the hand.
[1237] Input: Captured sign language data
[1238] Specific operation: An algorithm is run to analyze the shape and movement of the hand.
[1239] Output: Preprocessed sign language data
[1240] Step 3: Recognize and translate sign language data
[1241] The terminal sends the preprocessed sign language data to a sign language recognition algorithm to generate corresponding text data.
[1242] Input: Preprocessed sign language data
[1243] Specific behavior: The algorithm converts the sign "hello" into text "hello."
[1244] Output: Generated text data
[1245] Step 4: Generate audio data
[1246] The terminal sends the generated text data to a speech synthesis engine and converts it into voice data.
[1247] Input: Generated text data
[1248] Specific operation: The speech synthesis engine converts "hello" into speech.
[1249] Output: Generated audio data
[1250] Step 5: Outputting audio data
[1251] The device will play the translated audio through the speaker.
[1252] Input: Generated audio data
[1253] Specific action: "Hello" is output from the speaker.
[1254] Output: Translation played back as audio
[1255] Step 6: Capture and analyze emotions
[1256] The device uses a camera and microphone to analyze the user's emotions.
[1257] Input: Captured sign language data, as well as the user's facial expressions and tone of voice
[1258] Specific behavior: Facial expression analysis algorithm determines emotions.
[1259] Output: Parsed emotion information
[1260] Step 7: Adjusting the audio data
[1261] The device adjusts the tone and content of the voice based on the analyzed emotional information.
[1262] Input: Parsed emotion information
[1263] Specific behavior: If the user has a happy expression, play "Hello" in a bright tone.
[1264] Output: Modified audio data
[1265] These are the specific processing steps of this system. Based on the data input at each step, it is possible to provide optimal output according to the user's situation and emotions.
[1266] (Application example 2)
[1267] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1268] In autonomous vehicles, there is a need to address the safety and information shortages faced by the hearing- and visually impaired. Another challenge is to provide optimal information in response to the user's emotions, enabling safer and more comfortable use.
[1269] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing ambient sounds, means for converting the captured voice data into text data, means for analyzing the type, direction, and distance of the sound, means for displaying the analyzed information, means for recognizing the user's emotions, and means for adjusting the display information based on the emotion information. This enables hearing-impaired and visually impaired people to use autonomous vehicles safely and comfortably.
[1270] An "ambient sound capturing means" is a device or mechanism capable of detecting sounds in the environment and recording them as digital data.
[1271] A "means for converting captured audio data into text data" is a mechanism or software that analyzes the recorded audio signal and converts it into corresponding text information.
[1272] "Means for analyzing sound type, direction and distance" means an algorithm or mechanism for calculating and identifying the source, type and distance of a sound from captured audio data.
[1273] "Means for displaying analyzed information" refers to a display or other display device for visually presenting the results of the analysis.
[1274] The "means for recognizing user emotions" is an engine or software for analyzing the user's tone of voice and facial expressions to determine their emotional state.
[1275] The "means for adjusting displayed information based on emotional information" refers to a mechanism or program that dynamically changes the content or emphasis of displayed information in response to a recognized emotional state.
[1276] A "means for capturing sign language" is a camera or sensor that captures and records the user's sign language actions.
[1277] A "means for converting captured sign language data into text data" is an algorithm or system for analyzing the recorded sign language movements and converting them into corresponding written information.
[1278] The "means for converting the converted text data into voice" refers to a voice synthesis device or software for outputting the character data as voice.
[1279] "Means for capturing surrounding environmental data" refers to devices or systems that use sensors and cameras to capture the conditions inside and around the vehicle.
[1280] The "means for analyzing captured environmental data" is an algorithm or program for extracting and analyzing useful information from the acquired environmental data.
[1281] The "means for generating the analyzed information as a voice guide" is an engine or program for generating voice guidance based on the analysis results.
[1282] The "means for outputting the generated audio guide" refers to a device for informing the user of the generated audio guide using a speaker or the like.
[1283] The "means for adjusting the tone and content of the audio description based on emotional information" refers to a system or algorithm for dynamically changing the tone and content of the audio description depending on the user's emotional state.
[1284] The present invention relates to an assistance system for the hearing impaired and the visually impaired that is applied to an autonomous vehicle. Specific embodiments of the present invention will be described below.
[1285] System configuration
[1286] The system consists of the user's mobile device (hereafter referred to as the device) and a cloud-based server (hereafter referred to as the server). The device is equipped with a microphone, camera, speaker, display, various sensors, and an emotion recognition engine. This hardware is used to collect and analyze information about the user's surroundings and emotional state, and provide the necessary information.
[1287] Hearing-impaired features
[1288] Subtitling of ambient sounds
[1289] 1. Sound capture and processing
[1290] The device uses a microphone (e.g., general hardware) to capture ambient sounds. The captured audio data undergoes pre-processing such as noise reduction within the device.
[1291] The preprocessed voice data is converted into text in real time using voice recognition software (e.g., Google Speech Recognition API).
[1292] 2. Analysis of sound type, direction, and distance
[1293] The device identifies the type of sound from the converted text information and sound characteristics, and uses the arrangement of multiple microphones to calculate the direction and distance from which the sound is coming.
[1294] 3. Displaying Information
[1295] The analyzed information is displayed in text format on the device's display, allowing the user to understand the surrounding situation in real time.
[1296] 4. Emotion recognition and regulation
[1297] The device uses a camera (e.g., general hardware) or microphone to capture the user's facial expressions and tone of voice, which are then analyzed by an emotion recognition engine (e.g., the EmotionRecognition library).
[1298] The system adjusts the displayed text and notification method based on the analyzed emotion information. For example, if a car horn is honking from the right at a distance of 10 meters, and the user shows signs of surprise, the display will be emphasized.
[1299] Features for the visually impaired
[1300] Audio guidance of surrounding objects and situations
[1301] 1. Environmental data capture and processing
[1302] The device uses a camera and various sensors to capture data on the surrounding environment, which is then subjected to image recognition preprocessing to extract object contours and feature points.
[1303] 2. Object Recognition and Analysis
[1304] The pre-processed data is then used by AI algorithms (e.g., TensorFlow) to recognize surrounding objects, analyze their location and movement, and extract important information.
[1305] 3. Audio guide generation and output
[1306] The analyzed information is generated as audio guidance and converted into speech by a speech synthesis engine (e.g., gTTS library). The device then notifies the user of the generated audio guidance through the speaker.
[1307] 4. Emotion recognition and regulation
[1308] The device captures the user's emotions and adjusts the voice guidance based on the analyzed emotion information, for example, softening the tone of the voice guidance if the user is nervous.
[1309] Automatic sign language recognition and translation
[1310] 1. Sign Language Capture and Processing
[1311] The device uses a camera to capture the user's sign language in real time, and the captured video data undergoes pre-processing to extract hand shapes and movements.
[1312] 2. Sign Language Recognition and Translation
[1313] The preprocessed video data is analyzed using a sign language recognition algorithm (e.g., OpenCV or TensorFlow) to generate corresponding text information, which is then converted into speech data by a speech synthesis engine.
[1314] 3. Audio Output
[1315] The device plays the translated audio through a speaker, helping people who don't understand sign language to communicate with the user.
[1316] 4. Emotion Recognition and Translation Adjustment
[1317] The device analyzes the user's emotions along with the captured sign language data, and adjusts the expression of the text translation and the tone of the voice based on the emotional information. For example, if the user signs "hello" and has a happy expression, the tone of the voice output will be brighter.
[1318] Prompt Sentence Examples
[1319] "Can you hear the horn now?", "There is an obstacle ahead", "If the user is nervous, generate voice prompts in a softer tone."
[1320] As a result, information and notifications are appropriately adjusted according to the user's emotional information, enabling hearing-impaired and visually-impaired people to use self-driving vehicles safely and comfortably.
[1321] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1322] Step 1:
[1323] Audio data capture and preprocessing
[1324] Input: Ambient sound
[1325] Output: Preprocessed audio data
[1326] The device uses a microphone to capture surrounding sounds. The captured audio data undergoes pre-processing such as noise reduction. This processing makes the audio data easier to analyze. The data is also updated at regular intervals.
[1327] Step 2:
[1328] Converting audio data to text
[1329] Input: Preprocessed audio data
[1330] Output: Text data
[1331] The device converts the pre-processed voice data into text in real time using speech recognition software (e.g., Google Speech Recognition API). This process captures the voice as text information.
[1332] Step 3:
[1333] Analysis of sound type, direction, and distance
[1334] Input: Text data and audio data features
[1335] Output: Information about the type, direction, and distance of the sound
[1336] The device identifies the type of sound from the converted text information and features of the voice data, and calculates the direction and distance from which the sound is coming using a multi-microphone arrangement, identifying, for example, the sound of an ambulance siren or car horn.
[1337] Step 4:
[1338] Viewing analysis information
[1339] Input: Information about the type, direction, and distance of the sound
[1340] Output: Text information displayed
[1341] The device displays the analyzed information in text format on the display, allowing the user to understand the surrounding situation in real time. For example, it may display information such as "There is a horn honking 10 meters from the right."
[1342] Step 5:
[1343] Emotion recognition
[1344] Input: Camera video, audio data (if necessary)
[1345] Output: Emotional information
[1346] The device uses a camera (and optionally a microphone) to capture the user's facial expressions and tone of voice, which are then analyzed by an emotion recognition engine (e.g., the EmotionRecognition library). In the process, the user's emotional state (such as surprise, tension, or joy) is determined.
[1347] Step 6:
[1348] Adjusting display and notification information
[1349] Input: Emotion information, analysis information
[1350] Output: Adjusted text information or audio notification
[1351] The device adjusts the displayed text and notification method based on the acquired emotional information. For example, if the user expresses surprise, the device will emphasize the text or display it in a more visually appealing way. If a voice notification is required, the device will use a speech synthesis engine to adjust the tone of the notification.
[1352] Step 7:
[1353] Sign language capture and translation
[1354] Input: Sign language video data
[1355] Output: Text data, audio data
[1356] The device uses a camera to capture the user's sign language in real time. The captured video data is preprocessed to extract hand movements and shapes. The sign language is analyzed using a sign language recognition algorithm (e.g., OpenCV, TensorFlow), and corresponding text information is generated and output as voice data using a speech synthesis engine.
[1357] Step 8:
[1358] Environmental Data Capture and Analysis
[1359] Input: Environmental data (video and sensor data)
[1360] Output: Analysis information
[1361] The device uses cameras and sensors to capture data about the surrounding environment, which is then pre-processed using image recognition to extract object contours and feature points, and AI algorithms are used to identify the object type and location, generating key analytics.
[1362] Step 9:
[1363] Audio description generation and output
[1364] Input: Analysis information and emotion information
[1365] Output: Voice prompt
[1366] The device generates audio guidance based on the analyzed information and converts it into voice data using a speech synthesis engine (e.g., gTTS). The generated audio guidance is then played back to the user through the speaker. The tone and content of the audio guidance are adjusted according to the user's emotional state.
[1367] Through the above steps, the present invention assists hearing-impaired and visually-impaired people in safely and comfortably using autonomous vehicles.
[1368] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1369] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1370] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1371] [Third embodiment]
[1372] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1373] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1374] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1375] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1376] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1377] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1378] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1379] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1380] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1381] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1382] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1383] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1384] The present invention is an assistance system for the hearing impaired and the visually impaired, and is implemented in the following manner.
[1385] System configuration
[1386] This system consists of a user's mobile device (hereafter referred to as the device) and a cloud-based server (hereafter referred to as the server). The device is equipped with a microphone, camera, speaker, display, and various sensors, and uses this hardware to collect and analyze information about the user's surroundings and provide the necessary information.
[1387] Hearing-impaired features
[1388] Subtitling of ambient sounds
[1389] 1. Sound capture and processing
[1390] The device uses a microphone to constantly capture ambient sounds.
[1391] The captured audio data undergoes pre-processing such as noise reduction inside the device.
[1392] The preprocessed voice data is converted into text in real time by a voice recognition engine.
[1393] 2. Analysis of sound type, direction, and distance
[1394] The device identifies the type of sound from the converted text information and sound characteristics, such as "ambulance siren" or "human voice."
[1395] The device uses multiple microphone arrangements to calculate the direction and distance of the sound and notify the user.
[1396] 3. Displaying Information
[1397] The analyzed information is displayed in text format on the device's display, allowing the user to understand the surrounding situation in real time.
[1398] 4. Specific Examples
[1399] For example, if a car horn sounds near the user, the device will display "The horn is sounding from 10 meters to the right."
[1400] Automatic sign language recognition and translation
[1401] 1. Sign Language Capture and Processing
[1402] The device uses a camera to capture the user's sign language in real time.
[1403] The captured video data is pre-processed to extract the shape and movement of the hand.
[1404] 2. Sign Language Recognition and Translation
[1405] The pre-processed video data is analyzed by a sign language recognition algorithm to generate corresponding text information.
[1406] The generated text information is converted into voice data by a voice synthesis engine.
[1407] 3. Audio Output
[1408] The device plays the translated audio through a speaker, helping people who don't understand sign language to communicate with the user.
[1409] 4. Specific Examples
[1410] For example, if a user says "hello" in sign language, the terminal recognizes the sign and outputs "hello" aloud.
[1411] Features for the visually impaired
[1412] Audio guidance of surrounding objects and situations
[1413] 1. Environmental data capture and processing
[1414] The device uses a camera and various sensors to capture data about the surrounding environment.
[1415] The captured data is subjected to image recognition preprocessing to extract the contours and feature points of the object.
[1416] 2. Object Recognition and Analysis
[1417] The preprocessed data is then used by AI algorithms to recognize surrounding objects and situations.
[1418] The location and movement of recognized objects are analyzed to pick out important information.
[1419] 3. Audio guide generation and output
[1420] The analyzed information is generated as audio guide and converted into voice by a voice synthesis engine.
[1421] The terminal notifies the user of the generated audio guide through a speaker.
[1422] 4. Specific Examples
[1423] If the user is walking and there is an obstacle one meter ahead, the device will announce with a voice message, "There is an obstacle one meter ahead."
[1424] Starting and shutting down the system
[1425] 1. Initial Setup
[1426] During initial setup, the user selects the system mode (hearing impaired mode, visually impaired mode) that best suits their needs.
[1427] The server receives the user's configuration information and prepares the required services.
[1428] 2. Service launch
[1429] The device will activate the necessary sensors (microphone, camera, etc.) based on the initial settings and begin collecting data.
[1430] The server analyzes the data sent from the terminal and returns the necessary information in real time.
[1431] The terminal displays or outputs the analysis results to the user.
[1432] 3. Termination or Suspension of Service
[1433] If a user wishes to stop using the service, they can select the option to terminate or suspend the service in their device's settings screen.
[1434] The terminal receives the service termination command, shuts down all sensors, and cuts off communication with the server.
[1435] This system provides useful support to the hearing-impaired and visually impaired to overcome difficulties in daily life and live safely and independently.
[1436] The processing flow will be explained below.
[1437] Features for the hearing impaired: Subtitling of surrounding sounds
[1438] Step 1: Capture the sound
[1439] The device activates the microphone and captures ambient sounds at 0.1 second intervals.
[1440] Step 2: Preprocessing the audio data
[1441] The device performs noise reduction and normalization on the captured audio data.
[1442] Step 3: Speech to text
[1443] The device sends the preprocessed voice data to a voice recognition engine, which converts it into text data in real time.
[1444] Step 4: Identify the type of sound
[1445] The device extracts sound characteristics from the text data and identifies the type of sound based on them (e.g., "ambulance siren," "human voice," etc.).
[1446] Step 5: Analyze sound direction and distance
[1447] The device uses multiple microphone arrangements to calculate the direction and distance from which the sound is coming.
[1448] Step 6: Display subtitles
[1449] The device displays text information on the display, informing the user of the type, direction, and distance of the sound.
[1450] As a specific example, it displays "The sound of a horn is coming from 10 meters to the right."
[1451] Features for the visually impaired: Audio guide to surrounding objects and situations
[1452] Step 1: Capture environment data
[1453] The device will activate its camera and sensors to capture data on the surrounding environment every 0.5 seconds.
[1454] Step 2: Preprocessing the data
[1455] The device performs image recognition preprocessing on the captured video data to extract the contours and feature points of objects.
[1456] Step 3: Object Recognition
[1457] The device sends the pre-processed data to AI algorithms to recognize surrounding objects.
[1458] Step 4: Obtaining information about the object
[1459] The device analyzes the location and movement of recognized objects and picks out important information (e.g., obstacles ahead, approaching people).
[1460] Step 5: Generate audio guide
[1461] The device generates the analyzed information as a text message, sends it to a speech synthesis engine, and generates audio guidance.
[1462] Step 6: Audio Notifications
[1463] The terminal notifies the user of the generated audio guide through a speaker.
[1464] As a specific example, a voice message will be displayed saying, "There is an obstacle one meter ahead."
[1465] Starting and shutting down the system
[1466] Step 1: Initial Setup
[1467] The user configures the system and selects the mode that best suits their needs (e.g., hearing-impaired mode, visually-impaired mode).
[1468] The server receives the user's configuration information and prepares to provide the required services.
[1469] Step 2: Real-time analytics and notifications
[1470] Based on the selected mode, the device will activate the necessary sensors (microphone, camera, etc.) and collect data.
[1471] The server analyzes the data sent from the terminal and returns the necessary information in real time.
[1472] The device notifies the user of the analysis results in an appropriate format (subtitles, audio).
[1473] Step 3: Terminate or suspend service
[1474] If a user wants to stop using the system, they can select the option to terminate or suspend the service on their device's settings screen.
[1475] The terminal receives the service termination command, shuts down all sensors, and cuts off communication with the server.
[1476] Example 1
[1477] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1478] The purpose of this invention is to provide support for the hearing-impaired and visually-impaired to overcome the difficulties they face in daily life and live safely and independently. Specifically, the invention involves the development of a system that recognizes surrounding sounds, sign language, and environmental conditions in real time and provides the necessary information in an appropriate format.
[1479] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1480] In this invention, the server includes means for capturing ambient sounds, means for preprocessing the captured audio data, means for converting the preprocessed audio data into text data, means for analyzing the type, direction, and distance of sound, means for displaying the analyzed information, means for capturing sign language, means for preprocessing the captured sign language data, means for converting the preprocessed sign language data into text data, means for converting the converted text data into audio data and outputting it, means for capturing ambient environmental data, means for preprocessing the captured environmental data, means for analyzing the preprocessed environmental data to recognize objects, means for generating audio guidance based on the recognized object information, and means for outputting the generated audio guidance. This enables hearing-impaired and visually impaired people to grasp their surroundings in real time and live safely and independently.
[1481] The "means for capturing ambient sound" refers to a device or sensor function for collecting audio data from the user's surrounding environment.
[1482] "Means for preprocessing captured audio data" refers to processes or algorithms used to remove noise from collected audio data and prepare it for speech recognition.
[1483] The "means for converting preprocessed voice data into text data" refers to software or a system for converting voice data into text information using voice recognition technology.
[1484] "Means for analyzing the type, direction and distance of sound" refers to an algorithm or device that analyzes text data and sound feature data to identify the source, type, direction and distance of a sound.
[1485] "Means for displaying analyzed information" refers to a device or system that uses a display, screen, or the like to visually convey the analysis results to the user.
[1486] A "means for capturing sign language" is a camera or sensor that records the user's hand movements and shapes in real time.
[1487] "Means for pre-processing captured sign language data" refers to a process or system that appropriately processes captured video data for sign language recognition and extracts hand shapes and movements.
[1488] A "means for converting preprocessed sign language data into text data" is a technology or system that analyzes captured sign language movements and converts them into corresponding text information.
[1489] "Means for converting the converted text data into audio data and outputting it" refers to software or a device for converting character information into audio and reproducing it via a speaker or the like.
[1490] The "means for capturing surrounding environmental data" refers to a device or function for collecting information about the user's surrounding environment using a camera or sensor.
[1491] "Means for pre-processing captured environmental data" refers to processes or algorithms that process collected environmental data into a form suitable for image recognition.
[1492] "Means for analyzing preprocessed environmental data to recognize objects" refers to AI algorithms or systems that identify and recognize the shape and position of objects from environmental data.
[1493] The "means for generating audio guidance based on recognized object information" refers to a system or process for creating audio guidance to be notified to the user based on information about the identified object.
[1494] The "means for outputting the generated audio guide" refers to a device or system for conveying the generated audio guide to the user using a speaker or the like.
[1495] This invention is an assistance system for hearing-impaired and visually-impaired people to overcome difficulties in daily life, and is implemented as a system including a user's mobile terminal (hereinafter referred to as "terminal") and a cloud-based server (hereinafter referred to as "server"). The terminal is equipped with a microphone, camera, speaker, display, and various sensors. These hardware components are used to collect and analyze information about the user's surroundings and provide the necessary information.
[1496] Hearing-impaired features
[1497] Subtitling of ambient sounds
[1498] The device uses a microphone to capture ambient sounds and preprocesses the audio data. It then uses noise reduction technology to remove background noise and converts the audio into text using a speech recognition engine. The device uses multiple microphone arrangements to analyze the type, direction, and distance of the sound and displays the analyzed information on the display. For example, if a user is walking on the sidewalk and hears a car horn honking from the right at a distance of 10 meters, the device will display the message, "A car horn is honking from the right at a distance of 10 meters."
[1499] Automatic sign language recognition and translation
[1500] The device uses a camera to capture the user's sign language in real time and preprocesses the video data. After extracting the hand shape and movement, a sign language recognition algorithm generates corresponding text information. Next, a speech synthesis engine converts the text information into audio data and outputs it from the speaker. For example, if the user signs "hello," the device will recognize the sign and output "hello" aloud.
[1501] Features for the visually impaired
[1502] Audio guidance of surrounding objects and situations
[1503] The device uses cameras and various sensors to capture and preprocess data on the surrounding environment. It uses image recognition technology to extract the contours and feature points of objects, and uses AI algorithms to recognize surrounding objects and situations. It analyzes the location and movement of recognized objects and generates important information as audio guidance. This audio guidance is converted into voice by a speech synthesis engine and notified to the user through the speaker. For example, if a user is walking and there is an obstacle one meter ahead, the device will announce, "There is an obstacle one meter ahead."
[1504] Starting and shutting down the system
[1505] During the initial setup, the user selects the system mode (deaf mode, visually impaired mode) that best suits their needs. The server receives the user's configuration information and prepares the required services. The device activates the required sensors (microphone, camera, etc.) based on the initial setup and begins collecting data. The server analyzes the data sent from the device and returns the required information in real time. If the user wants to terminate or pause the service, they select the option to terminate or pause the service on the device's settings screen. The device receives the service termination command, stops all sensors, and cuts off communication with the server.
[1506] Examples of concrete examples and prompts
[1507] Example prompt for speech recognition: "Transcribe the speech that says 'hello' to text."
[1508] Example prompt for sign language recognition: "Translate this sign language video into text."
[1509] This provides useful support to hearing-impaired and visually impaired people to overcome difficulties in daily life and live safely and independently.
[1510] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1511] Hearing-impaired features
[1512] Subtitling of ambient sounds
[1513] Step 1: Capture the sound
[1514] Subject: Device
[1515] How it works: The device uses a microphone to capture ambient sound, taking ambient audio data as input and raw captured data as output.
[1516] Input and Output: The input is the ambient audio data, and the output is the captured audio data.
[1517] Step 2: Pre-processing the sound
[1518] Subject: Device
[1519] Specific operation: The device preprocesses the captured audio data and removes noise. The input is the captured audio data, and the output is the noise-removed audio data.
[1520] Input and Output: The input is the captured audio data, and the output is the pre-processed audio data.
[1521] Step 3: Voice Recognition
[1522] Subject: Device
[1523] Specific operation: The device sends the preprocessed voice data to the voice recognition engine and converts it into text data. The preprocessed voice data is input, and the converted text data is obtained as output.
[1524] Input and Output: The input is preprocessed audio data, and the output is text data.
[1525] Step 4: Analyzing the sound information
[1526] Subject: Device
[1527] Specific operation: The device analyzes the text data and sound feature data to identify the type, direction, and distance of the sound. The text data and sound feature data are input, and the analysis results are obtained as output.
[1528] Input and output: The input is text data and sound feature data, and the output is analyzed sound information.
[1529] Step 5: Viewing information
[1530] Subject: Device
[1531] Specific operation: The terminal displays the analyzed information in text format on the display. The analyzed sound information is input and displayed on the display as output.
[1532] Input and output: The input is the analyzed sound information, and the output is the text information displayed on the screen.
[1533] Automatic sign language recognition and translation
[1534] Step 1: Capture sign language
[1535] Subject: Device
[1536] Specific operation: The device uses a camera to capture the user's sign language actions in real time. The sign language actions are input and the captured video data is obtained as output.
[1537] Input and Output: The input is sign language actions, and the output is captured video data.
[1538] Step 2: Preprocessing the video data
[1539] Subject: Device
[1540] Specific operation: The device preprocesses the captured video data and extracts the hand shape and movement. The captured video data is input and the preprocessed video data is output.
[1541] Input and Output: The input is the captured video data, and the output is the pre-processed video data.
[1542] Step 3: Sign Language Recognition and Text Conversion
[1543] Subject: Device
[1544] Specific operation: The device sends the preprocessed video data to a sign language recognition algorithm to generate corresponding text data. The preprocessed video data is input, and the generated text data is output.
[1545] Input and Output: The input is the preprocessed video data, and the output is the generated text data.
[1546] Step 4: Speech synthesis and output
[1547] Subject: Device
[1548] Specific operation: The device sends the generated text data to a speech synthesis engine, converts it into voice data, and outputs the converted voice data to the user through the speaker. Text data is input, and voice data is obtained as output.
[1549] Input and output: The input is the generated text data, and the output is the audio output from the speaker.
[1550] Features for the visually impaired
[1551] Audio guidance of surrounding objects and situations
[1552] Step 1: Capture environment data
[1553] Subject: Device
[1554] Specific operation: The device uses a camera and various sensors to capture surrounding environmental data. The surrounding environmental information is input, and the captured environmental data is output.
[1555] Input and Output: The input is the surrounding environment information, and the output is the captured environment data.
[1556] Step 2: Preprocessing the image data
[1557] Subject: Device
[1558] Specific operation: The device preprocesses the captured environmental data and extracts object contours and feature points. The captured environmental data is input, and the preprocessed environmental data is output.
[1559] Input and Output: The input is the captured environmental data, and the output is the preprocessed environmental data.
[1560] Step 3: Object Recognition
[1561] Subject: Device
[1562] How it works: The device analyzes the preprocessed data using AI algorithms to recognize surrounding objects and situations. The preprocessed data is input, and the recognized object information is output.
[1563] Input and Output: The input is the preprocessed data, and the output is the recognized object information.
[1564] Step 4: Location Analysis
[1565] Subject: Device
[1566] Specific operation: The device analyzes the location and movement of the recognized object and extracts important information. The recognized object information is input, and the location and movement analysis results are obtained as output.
[1567] Input and Output: The input is the recognized object information, and the output is the analysis result of the position and movement.
[1568] Step 5: Audio guide generation and output
[1569] Subject: Device
[1570] Specific operation: The device generates the analyzed information as an audio guide and converts it into speech using a speech synthesis engine. The converted audio guide is then notified to the user through the speaker. The analyzed information is input, and the audio guide is obtained as output.
[1571] Input and Output: The input is the parsed information, and the output is the audio guide from the speaker.
[1572] Starting and shutting down the system
[1573] Step 1: Initial Setup
[1574] Subject: User
[1575] Specific operation: The user selects the system mode (deaf mode, visually impaired mode) according to their needs in the initial setting. The user's needs are input, and the setting information is sent to the server as output.
[1576] Input and Output: Input is user needs, output is configuration information.
[1577] Step 2: Start the service
[1578] Subject: Server
[1579] Specific operation: The server receives the user's configuration information and prepares the necessary services. The device activates the necessary sensors (microphone, camera, etc.) based on the initial settings and begins collecting data. The configuration information is received as input and the service is started as output.
[1580] Input and Output: Input is the configuration information, output is the started service.
[1581] Step 3: Data collection and analysis
[1582] Subject: Device
[1583] Specific operation: The terminal starts collecting data, and the server analyzes the data sent from the terminal. The collected data is input, and the analysis results are obtained as output.
[1584] Input and Output: Input is the collected data and output is the analysis result.
[1585] Step 4: Notification of information
[1586] Subject: Device
[1587] Specific operation: The device displays or outputs the analysis results to the user. The analysis results are input and notified to the user as output.
[1588] Input and output: The input is the analysis result, and the output is the notification to the user.
[1589] Step 5: Terminate or suspend service
[1590] Subject: User
[1591] Specific behavior: When a user wants to stop using a service, they select the option to terminate or pause the service in the device settings screen. The user's operation is the input, and the service is terminated or paused as the output.
[1592] Input and Output: Input is a user action, output is a service that has been terminated or suspended.
[1593] (Application example 1)
[1594] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1595] The purpose of this invention is to provide a means to resolve the lack of information and communication difficulties faced by hearing-impaired and visually-impaired people when shopping comfortably and safely in a brick-and-mortar store. Specifically, the objective is to provide a system that supports people with disabilities to act independently by acquiring in-store announcements, conversations with store staff, and product information in real time and conveying this information appropriately to the user.
[1596] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1597] In this invention, the server includes means for capturing ambient sounds, means for converting the captured audio data into text data, means for analyzing the type, direction, and distance of the sound, means for displaying the analyzed information, means for displaying the captured text data on a smart device display, means for capturing sign language, means for converting the captured sign language data into text data, means for converting the converted text data into audio, means for capturing ambient environmental data, means for analyzing the captured environmental data, means for generating the analyzed information as an audio guide, and means for outputting the generated audio guide. This allows hearing-impaired people to understand in-store announcements and conversations in real time through subtitles, and makes it easy to convert sign language communication into audio. Furthermore, visually impaired people can receive information about surrounding objects and products as an audio guide, enabling safe and efficient shopping.
[1598] "Means for capturing ambient sounds" is a function that collects environmental sounds using the microphone of the smart device.
[1599] "Means for converting captured voice data into text data" refers to a function that transcribes collected voice data using voice recognition technology.
[1600] "Means for analyzing the type, direction and distance of sound" is a function that analyzes collected audio data to identify the type, location and distance of the sound source.
[1601] "Means for displaying analyzed information" refers to a function that visually displays the analysis results on the display of a smart device.
[1602] The "means for displaying captured text data on the display of the smart device" is a function for displaying text obtained by speech recognition on the screen of the smart device.
[1603] A "means for capturing sign language" is a function that uses the camera of a smart device to capture sign language movements.
[1604] The "means for converting captured sign language data into text data" is a function that analyzes the video of the sign language and converts it into text information.
[1605] The "means for converting the converted text data into voice" is a function for outputting the text data as voice using voice synthesis technology.
[1606] "Means for capturing surrounding environmental data" refers to the ability to collect images and data of the environment using the smart device's camera and sensors.
[1607] "Means for analyzing captured environmental data" refers to a function that analyzes collected video and data to recognize objects and environments.
[1608] "Means for generating analyzed information as audio guide" is a function for creating audio guide based on the analysis results.
[1609] "Means for outputting the generated audio guide" is a function that transmits the generated audio guide to the user through the speaker of the smart device.
[1610] This invention is a system that helps hearing- and visually-impaired people enjoy shopping comfortably and safely in physical stores. The system consists of a user's smart device and a cloud-based server. Specifically, the system collects and analyzes surrounding information using the smart device's microphone, camera, display, speaker, and various sensors. This collected and analyzed information is then provided to the user appropriately.
[1611] System configuration
[1612] 1. Subtitling of ambient sounds
[1613] Capture method: Capture surrounding sounds using the microphone on your smart device.
[1614] Voice data conversion: The captured voice data is converted into text data using voice recognition technology. Specifically, Google Cloud's Speech-to-Text API is used.
[1615] Analysis means: Analyzes the type, direction and distance of sound and provides important information to the user.
[1616] Display method: The analyzed information is displayed in real time as subtitles on the smart device display, allowing hearing-impaired people to understand in-store announcements and conversations.
[1617] 2. Automatic Sign Language Recognition and Translation
[1618] Capture method: Sign language is captured using the smart device's camera.
[1619] Sign language data conversion: The captured sign language data is converted into text data using image processing techniques and sign language recognition algorithms.
[1620] Speech conversion method: The converted text data is converted into speech using speech synthesis technology (Google Cloud's Text-to-Speech API) and played back through the smart device's speaker, allowing hearing-impaired people to easily communicate with store staff and other customers.
[1621] 3. Audio guide to surrounding objects and situations
[1622] Capture method: Capture surrounding environmental data using the smart device's camera and various sensors.
[1623] Environmental data analysis: The captured data is analyzed using image recognition technology (Google Vision API) to identify objects and obtain their location information.
[1624] Audio guide generation: Audio guides are generated based on the analyzed information and converted into audio using Google Cloud's Text-to-Speech API.
[1625] Output means: The generated audio guide is output from the smart device's speaker to inform visually impaired people of their surroundings.
[1626] Specific examples
[1627] For example, when a user wears a smart device and enters a physical store, the system captures in-store announcements with a microphone and displays information such as "Today's special sale is XX" in real time as subtitles on the display. Also, if the user expresses "Where is it?" in sign language, the system captures the sign, converts it into text, and relays it to the store clerk as "Where is it?" Furthermore, when a visually impaired person stands in front of a shelf, the camera recognizes the product label and provides an audio guide such as "There is a fruit shelf ahead." This guide is provided through the smart device's speaker.
[1628] Prompt Sentence Examples
[1629] "Please provide real-time subtitles for in-store announcements."
[1630] "Please recognize the sign language, convert it into speech, and communicate it to the store clerk."
[1631] "Recognize surrounding objects and provide voice guidance"
[1632] As can be seen, this system provides a useful aid for the hearing and visually impaired to enjoy shopping independently.
[1633] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1634] Step 1:
[1635] Capture Processing
[1636] Using sensors such as the device's microphone and camera, surrounding audio and video data is captured in real time.
[1637] Input: Ambient audio, video such as sign language, and environmental data.
[1638] How it works: A microphone captures audio data, and a camera captures sign language and footage of the environment.
[1639] Step 2:
[1640] Pretreatment
[1641] Noise reduction and image processing are performed on the captured audio and video data, converting it into a format that is easy to analyze.
[1642] Input: Raw captured data (audio data, video data).
[1643] Data processing: Noise reduction of audio data, contour extraction of video data.
[1644] Output: Preprocessed audio data, preprocessed video data.
[1645] Operation: A noise reduction algorithm is used to remove noise from the audio data, and a contour extraction filter is applied to the video data.
[1646] Step 3:
[1647] Data transformation and analysis
[1648] The preprocessed data is analyzed using a voice recognition engine or image recognition algorithm.
[1649] Input: Preprocessed audio data, preprocessed video data.
[1650] Data calculation: Voice data is converted into text using a speech recognition algorithm (Google Cloud Speech-to-Text API), and sign language is analyzed and converted into text data using an image recognition algorithm.
[1651] Output: Text data, object recognition data.
[1652] How it works: It calls the Google Cloud Speech-to-Text API to convert speech to text, then uses a sign language recognition algorithm to convert sign language to text.
[1653] Step 4:
[1654] Displaying and Speech Output of Data
[1655] The converted text data is displayed on the smart device's display, and if necessary, a voice synthesis engine is used to output the text as speech.
[1656] Input: Converted text data, object recognition data.
[1657] Data calculation: Text data is displayed as a user interface, and speech is generated using a speech synthesis algorithm (Google Cloud Text-to-Speech API).
[1658] Output: Text displayed on the screen, audio guidance output from the speaker.
[1659] How it works: Text data is displayed as subtitles on the smart device's display, and audio data is generated by calling the Google Cloud Text-to-Speech API and output through the speaker.
[1660] Step 5:
[1661] User Feedback Processing
[1662] If there is additional input from the user, that input is passed back to the capture process for further analysis and interaction.
[1663] Input: New input from the user (speech, sign language, changes in the environment).
[1664] Data processing: Preprocess and analyze again as necessary.
[1665] Output: Updated visual and audio guides.
[1666] What it does: Captures user feedback, then analyzes it again and displays / speech outputs it.
[1667] Through these steps, hearing and visually impaired people can have a more comfortable and safer shopping experience in physical stores.
[1668] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1669] The present invention is an assistive system for the hearing and visually impaired, which has the ability to recognize a user's emotions and use that information to adjust output. Detailed description of the preferred embodiments of the present invention is provided below.
[1670] System configuration
[1671] This system consists of a user's mobile device (hereafter referred to as the device) and a cloud-based server (hereafter referred to as the server). The device is equipped with a microphone, camera, speaker, display, various sensors, and an emotion recognition engine, and uses this hardware to collect and analyze information about the user's surroundings and emotional state, and provide the necessary information.
[1672] Hearing-impaired features
[1673] Subtitling of ambient sounds
[1674] 1. Sound capture and processing
[1675] The device uses a microphone to capture ambient sounds at 0.1 second intervals.
[1676] The captured audio data undergoes pre-processing such as noise reduction inside the device.
[1677] The preprocessed voice data is converted into text in real time by a voice recognition engine.
[1678] 2. Analysis of sound type, direction, and distance
[1679] The device identifies the type of sound from the converted text information and sound characteristics (e.g., "ambulance siren," "human voice," etc.).
[1680] The device uses multiple microphone arrangements to calculate the direction and distance from which the sound is coming.
[1681] 3. Displaying Information
[1682] The analyzed information is displayed in text format on the device's display, allowing the user to understand the surrounding situation in real time.
[1683] 4. Emotion recognition and regulation
[1684] The device uses a camera and microphone to capture the user's facial expressions and tone of voice, which are then analyzed by an emotion engine.
[1685] The text displayed and notification method are adjusted based on the analyzed emotional information.
[1686] For example, when displaying "The sound of a horn is honking from 10 meters to the right," if the user appears surprised, the display will be emphasized.
[1687] Features for the visually impaired
[1688] Audio guidance of surrounding objects and situations
[1689] 1. Environmental data capture and processing
[1690] The device uses a camera and various sensors to capture data about the surrounding environment at 0.5-second intervals.
[1691] The captured data is subjected to image recognition preprocessing to extract the contours and feature points of the object.
[1692] 2. Object Recognition and Analysis
[1693] The pre-processed data is then used by AI algorithms to recognize surrounding objects.
[1694] The location and movement of recognized objects are analyzed to pick out important information.
[1695] 3. Audio guide generation and output
[1696] The analyzed information is generated as audio guide and converted into voice by a voice synthesis engine.
[1697] The terminal notifies the user of the generated audio guide through a speaker.
[1698] 4. Emotion recognition and regulation
[1699] The device captures the user's emotions and adjusts the audio guidance based on the analyzed emotion information.
[1700] As a specific example, if the user is nervous, the tone of the voice guide is softened.
[1701] Automatic sign language recognition and translation
[1702] 1. Sign Language Capture and Processing
[1703] The device uses a camera to capture the user's sign language in real time.
[1704] The captured video data is pre-processed to extract the shape and movement of the hand.
[1705] 2. Sign Language Recognition and Translation
[1706] The pre-processed video data is analyzed by a sign language recognition algorithm to generate corresponding text information.
[1707] The generated text information is converted into voice data by a voice synthesis engine.
[1708] 3. Audio Output
[1709] The device plays the translated audio through a speaker, helping people who don't understand sign language to communicate with the user.
[1710] 4. Emotion Recognition and Translation Adjustment
[1711] The device analyzes the user's emotions along with the captured sign language data.
[1712] The device adjusts the expression and tone of the text translation based on the emotional information.
[1713] As a specific example, if the user says "hello" in sign language and has a happy expression, the tone of the voice output is brightened.
[1714] Starting and shutting down the system
[1715] 1. Initial Setup
[1716] The user configures the system and selects the mode that best suits their needs (e.g., hearing-impaired mode, visually-impaired mode).
[1717] The server receives the user's configuration information and prepares to provide the required services.
[1718] 2. Real-time analysis and notifications
[1719] Based on the selected mode, the device will activate the necessary sensors (microphone, camera, etc.) and collect data.
[1720] The server analyzes the data sent from the terminal and returns the necessary information in real time.
[1721] The device notifies the user of the analysis results in an appropriate format (subtitles, audio), adjusting it according to the emotional information.
[1722] 3. Termination or Suspension of Service
[1723] If a user wants to stop using the system, they can select the option to terminate or suspend the service on their device's settings screen.
[1724] The terminal receives the service termination command, shuts down all sensors, and cuts off communication with the server.
[1725] This system provides useful support for hearing- and visually-impaired people to overcome difficulties in daily life and live safely and independently. Furthermore, by adjusting its response according to the user's emotions, it is possible to provide support that meets more individual needs.
[1726] The processing flow will be explained below.
[1727] Features for the hearing impaired: Subtitling of surrounding sounds
[1728] Step 1: Capture the sound
[1729] The device activates the microphone and captures ambient sounds at 0.1 second intervals.
[1730] Step 2: Preprocessing the audio data
[1731] The device performs noise reduction and normalization on the captured audio data.
[1732] Step 3: Speech to text
[1733] The device sends the preprocessed voice data to a voice recognition engine, which converts it into text data in real time.
[1734] Step 4: Identify the type of sound
[1735] The device extracts sound characteristics from the text data and identifies the type of sound based on them (e.g., "ambulance siren," "human voice," etc.).
[1736] Step 5: Analyze sound direction and distance
[1737] The device uses multiple microphone arrangements to calculate the direction and distance from which the sound is coming.
[1738] Step 6: Capturing Emotions
[1739] The device uses a camera and microphone to capture the user's facial expressions and tone of voice.
[1740] Step 7: Analyze the sentiment data
[1741] The device sends the captured emotion data to an emotion engine to analyze the user's emotions.
[1742] Step 8: Display subtitles
[1743] Based on the analyzed voice information and emotional data, the device displays text information on the screen, conveying the type, direction, and distance of the sound to the user in a format that is optimal for them.
[1744] For example, when displaying the message "A horn is honking from 10 meters to the right," if the user appears surprised, the warning message will be emphasized.
[1745] Features for the visually impaired: Audio guide to surrounding objects and situations
[1746] Step 1: Capture environment data
[1747] The device will activate its camera and sensors to capture data on the surrounding environment every 0.5 seconds.
[1748] Step 2: Preprocessing the data
[1749] The device performs image recognition preprocessing on the captured video data to extract the contours and feature points of objects.
[1750] Step 3: Object Recognition
[1751] The device sends the pre-processed data to AI algorithms to recognize surrounding objects.
[1752] Step 4: Obtaining information about the object
[1753] The device analyzes the location and movement of recognized objects and picks out important information (e.g., obstacles ahead, approaching people).
[1754] Step 5: Capturing emotions
[1755] The device uses a camera and microphone to capture the user's facial expressions and tone of voice.
[1756] Step 6: Analyze the sentiment data
[1757] The device sends the captured emotion data to an emotion engine to analyze the user's emotions.
[1758] Step 7: Generate audio guide
[1759] The device generates audio guidance based on the analyzed environmental information and emotional data.
[1760] For example, if the user is nervous, the tone of the voice prompt will be softened and the user will be notified, "There is an obstacle 1 meter ahead."
[1761] Step 8: Audio Notifications
[1762] The terminal notifies the user of the generated audio guide through a speaker.
[1763] Automatic sign language recognition and translation
[1764] Step 1: Capture sign language
[1765] The device uses a camera to capture the user's sign language in real time.
[1766] Step 2: Preprocessing the sign language data
[1767] The device performs pre-processing on the captured video data to extract the shape and movement of the hand.
[1768] Step 3: Sign Language Recognition
[1769] The device sends the preprocessed data to a sign language recognition algorithm to generate corresponding text information.
[1770] Step 4: Capturing emotions
[1771] The device captures the user's facial expressions with a camera along with the captured video data, and obtains emotion data.
[1772] Step 5: Analyze the sentiment data
[1773] The device analyzes the emotional data using an emotion engine and determines the user's emotions.
[1774] Step 6: Adjusting the text data
[1775] The device adjusts the text information generated from the sign language as needed based on emotional data.
[1776] Step 7: Converting text data into speech
[1777] The device sends the adjusted text data to a speech synthesis engine and converts it into voice data.
[1778] Step 8: Audio Output
[1779] The device plays audio data through a speaker, helping people who do not understand sign language to communicate with the user.
[1780] As a specific example, if a user who signs "hello" looks happy, the tone of the voice output is made brighter.
[1781] Example 2
[1782] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1783] The present invention aims to solve the problem of the lack of a function to adjust information according to the user's emotions in the conventional technology for life support systems for the hearing-impaired and visually-impaired. Specifically, the present invention aims to support a more comfortable and safe life by capturing surrounding sounds and environmental information in real time and providing it to the user, while adjusting the output content and format based on the user's emotions.
[1784] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1785] In this invention, the server includes means for capturing ambient sounds, means for preprocessing the captured audio data, means for converting the preprocessed audio data into text data, means for analyzing the type, direction, and distance of the sound, means for displaying the analyzed information, and means for recognizing the user's emotions and adjusting the display information. This allows the user to grasp the surrounding situation in real time and receive appropriate notifications according to the user's emotions. Furthermore, by including means for capturing sign language, means for preprocessing the captured sign language data, means for converting the preprocessed sign language data into text data, means for converting the converted text data into audio data and outputting it, and means for recognizing the user's emotions and adjusting the audio output, communication with sign language users can be facilitated. Furthermore, by including means for capturing ambient environmental data, means for preprocessing the captured environmental data, means for analyzing the preprocessed environmental data, means for generating the analyzed information as audio guidance, means for outputting the generated audio guidance, and means for recognizing the user's emotions and adjusting the audio guidance, it becomes easier for visually impaired people to grasp the surrounding situation and appropriate guidance according to their emotions can be provided.
[1786] The "means for capturing ambient sound" refers to a means for capturing ambient sound using an acoustic sensor such as a microphone and acquiring the audio data.
[1787] The "means for pre-processing captured audio data" refers to a means for performing processes such as noise reduction and filtering on the acquired audio data to make the audio signal clearer.
[1788] The "means for converting preprocessed voice data into text data" refers to a means for analyzing preprocessed voice data using a voice recognition technique and generating corresponding text data.
[1789] "Means for analyzing the type, direction and distance of a sound" refers to means for analyzing the characteristics of pre-processed and text-converted audio data to identify what the sound is (e.g., a car horn, a human voice, etc.), the direction from which it is coming, and how far away it is.
[1790] The "means for displaying analyzed information" refers to a means for displaying the analyzed audio information on a display or monitor so that the user can visually confirm the information.
[1791] "Means for recognizing the user's emotions and adjusting the displayed information" refers to means for analyzing the user's facial expressions and tone of voice using a camera or microphone, determining the user's emotions, and then adjusting the format and content of the displayed information.
[1792] The "means for capturing sign language" refers to a means for capturing a user's sign language actions in real time using a video capture device such as a camera, and acquiring the data.
[1793] The "means for preprocessing captured sign language data" refers to a means for processing the acquired sign language video data to make it easier to analyze, and extracting features such as hand shape and movement.
[1794] The "means for converting preprocessed sign language data into text data" refers to a means for analyzing the preprocessed sign language data using a sign language recognition algorithm and generating text data representing the meaning of the sign language.
[1795] "Means for converting the converted text data into voice data and outputting it" refers to means for generating voice data based on the text data using voice synthesis technology and notifying the user of the voice data through a speaker or the like.
[1796] The "means for capturing surrounding environmental data" refers to a means for obtaining visual and other data about surrounding objects and the environment using a camera or various sensors.
[1797] "Means for preprocessing captured environmental data" refers to means for performing image processing and data analysis on the acquired environmental data, extracting the contours and feature points of objects, and making the data easier to analyze.
[1798] "Means for analyzing preprocessed environmental data" refers to means for analyzing preprocessed environmental data using AI algorithms and recognizing objects and situations.
[1799] The "means for generating analyzed information as audio guidance" refers to a means for generating information to be provided to the user as audio guidance based on the analysis results.
[1800] The "means for outputting the generated audio guide" refers to a means for outputting the generated audio guide as audio using speech synthesis technology and notifying the user through a speaker.
[1801] The "means for recognizing the user's emotions and adjusting the audio guidance" refers to a means for analyzing the user's facial expressions and tone of voice, and adjusting the tone and content of the audio guidance based on the determined emotional information.
[1802] The present invention is an assistance system for the hearing-impaired and visually-impaired, which has the ability to recognize a user's emotions and adjust output using that information. Detailed embodiments for implementing the present invention will be described below.
[1803] System configuration
[1804] This system consists of a cloud-based server (hereafter referred to as the "server") and a mobile device (hereafter referred to as the "device") carried by the user. The device is equipped with a microphone, camera, speaker, display, various sensors, and an emotion recognition engine.
[1805] Embodiments for the hearing impaired
[1806] Subtitling of ambient sounds
[1807] The device uses a microphone to capture ambient sounds at 0.1-second intervals. The captured voice data undergoes preprocessing such as noise reduction within the device, and is then converted into text data by a voice recognition engine. This text data is then used to analyze the type, direction, and distance of the sound, and displayed on the device's display. The device also uses a camera and microphone to capture the user's emotions, which are analyzed by an emotion engine. The device adjusts the text displayed and notification method based on this emotional information.
[1808] As a specific example of how this works, if the user is surprised by the information "There is a horn honking 10 meters to the right," it will be displayed with an emphasis of "Caution!"
[1809] Embodiments for the visually impaired
[1810] Audio guidance of surrounding objects and situations
[1811] The device captures data on the surrounding environment every 0.5 seconds using a camera and various sensors. The captured data undergoes image recognition preprocessing to extract object contours and feature points. This data is analyzed using an AI algorithm to identify the location and movement of recognized objects and generate audio guidance. This audio guidance is output from the speaker via a speech synthesis engine. The device also captures the user's emotions and adjusts the audio guidance based on the analyzed emotional information.
[1812] As a specific example of how it works, if the user is nervous about an approaching car, the system will notify them in a soft tone, "There is a car approaching 10 meters ahead, please proceed slowly."
[1813] Embodiments of automatic sign language recognition and translation
[1814] The device uses a camera to capture sign language in real time and preprocesses the hand shape and movements. The preprocessed video data is then analyzed using a sign language recognition algorithm to generate corresponding text data. This text data is converted into audio data by a speech synthesis engine and output through the speaker. The device also captures the user's emotions and adjusts the tone of the voice based on the emotional information.
[1815] As a specific example of operation, if the user signs "hello" and looks happy, "hello" is output as voice in a bright tone.
[1816] Examples of prompt statements
[1817] An example prompt to be input to a generative AI model would be something like this:
[1818] 1. Prompt for the deaf:
[1819] "Capture ambient sounds and subtitle them in real time. Analyze the type, direction, and distance of the sound and display it in text format."
[1820] 2. Prompt for the visually impaired:
[1821] "Use cameras and sensors to capture data about the surrounding environment and output recognized object information as audio guidance."
[1822] 3. Prompt for sign language recognition:
[1823] "Use your camera to capture sign language in real time and translate it into text or speech."
[1824] As described above, this system adjusts the way information is displayed and notified according to the user's emotions, providing assistance that meets more individual needs.
[1825] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1826] Processing steps for the hearing impaired
[1827] Subtitling of ambient sounds
[1828] Step 1: Capture the sound
[1829] The device uses a microphone to capture ambient sounds at 0.1 second intervals.
[1830] Input: Ambient sound
[1831] Specific operation: The device picks up surrounding environmental sounds and voices using the microphone.
[1832] Output: Captured audio data
[1833] Step 2: Preprocessing the audio data
[1834] The device performs pre-processing such as noise reduction and filtering on the captured audio data.
[1835] Input: Captured audio data
[1836] What it does: The device removes background noise and emphasizes the main voice.
[1837] Output: Preprocessed audio data
[1838] Step 3: Convert audio data to text
[1839] The device sends the preprocessed voice data to a voice recognition engine, which converts it into text data in real time.
[1840] Input: Preprocessed audio data
[1841] Specific operation: The voice recognition engine converts spoken content such as "Hello, it's a nice day today" into text.
[1842] Output: Text data
[1843] Step 4: Analyze the type, direction and distance of the sound
[1844] The device uses the characteristics of text and voice data to identify the type of sound and calculates direction and distance using multiple microphone arrangements.
[1845] Input: Text data, audio data features
[1846] Specific operation: The device uses an analysis algorithm to classify sounds such as "car horns" and "human voices" and calculates direction and distance.
[1847] Output: Analyzed sound type, direction and distance
[1848] Step 5: Viewing information
[1849] The device displays the analyzed voice information on the display.
[1850] Input: Analyzed audio information
[1851] Specific operation: The device displays information such as "A horn is honking 10 meters to the right" on the screen.
[1852] Output: Information displayed on the display
[1853] Step 6: Capture and analyze emotions
[1854] The device uses a camera and microphone to capture the user's facial expressions and tone of voice, which are then analyzed by an emotion engine.
[1855] Input: User facial expressions and tone of voice
[1856] What it does: The device uses a facial expression analysis algorithm to detect emotions such as surprise or anxiety.
[1857] Output: Parsed emotion information
[1858] Step 7: Adjust notification methods
[1859] The device adjusts the text displayed and notification method based on the analyzed emotional information.
[1860] Input: Parsed emotion information
[1861] Specific behavior: If the user is startled, the message "Sound of horn 10 meters to the right -- Caution!" will be displayed in emphasis.
[1862] Output: Adjusted display information
[1863] Processing steps for the visually impaired
[1864] Audio guidance of surrounding objects and situations
[1865] Step 1: Capture environment data
[1866] The device uses a camera and various sensors to capture data on the surrounding environment every 0.5 seconds.
[1867] Input: Surrounding environment
[1868] Specific operation: The camera captures images of the road, pedestrians, cars, etc. ahead, and the sensor measures distance and movement.
[1869] Output: Captured environmental data
[1870] Step 2: Preprocessing environmental data
[1871] The device performs image processing on the captured environmental data to extract the contours and feature points of objects.
[1872] Input: Captured environmental data
[1873] How it works: Image processing algorithms extract the shape and features of objects.
[1874] Output: Preprocessed environmental data
[1875] Step 3: Object Recognition and Analysis
[1876] The device sends pre-processed environmental data to AI algorithms that recognize surrounding objects and analyze their position and movement.
[1877] Input: Preprocessed environmental data
[1878] Specific operation: The AI recognizes things like "there are two people ahead" or "a car is approaching."
[1879] Output: Position and movement of analyzed objects
[1880] Step 4: Generate audio guide
[1881] The device generates audio guidance based on the analyzed information and converts it into audio data using a speech synthesis engine.
[1882] Input: Analyzed object information
[1883] Specific operation: Generates voice guidance such as "There is a step 10 meters ahead."
[1884] Output: The generated audio guide
[1885] Step 5: Outputting voice guidance
[1886] The device will announce the generated audio guide through the speaker.
[1887] Input: Generated audio guide
[1888] Specific operation: The device will announce "There is a step 10 meters ahead" via voice.
[1889] Output: Guide output as audio
[1890] Step 6: Capture and analyze emotions
[1891] The device uses a camera and microphone to capture the user's emotions and analyzes them using an emotion engine.
[1892] Input: User facial expressions and tone of voice
[1893] Specific operation: The device analyzes facial expressions to determine whether the user is anxious or nervous.
[1894] Output: Parsed emotion information
[1895] Step 7: Adjust the audio prompts
[1896] The device adjusts the audio guidance based on the analyzed emotional information.
[1897] Input: Parsed emotion information
[1898] Specific behavior: If the user is nervous, a soft tone will be displayed saying, "There is a step ahead, please walk slowly."
[1899] Output: Adjusted voice prompts
[1900] Processing steps for automatic sign language recognition and translation
[1901] Step 1: Capture sign language
[1902] The device uses a camera to capture the user's sign language in real time.
[1903] Input: User's sign language actions
[1904] Specific operation: The camera captures the user's hand movements.
[1905] Output: Captured sign language data
[1906] Step 2: Preprocessing the sign language data
[1907] The device preprocesses the captured sign language data to make it easier to analyze, extracting the shape and movement of the hand.
[1908] Input: Captured sign language data
[1909] Specific operation: An algorithm is run to analyze the shape and movement of the hand.
[1910] Output: Preprocessed sign language data
[1911] Step 3: Recognize and translate sign language data
[1912] The terminal sends the preprocessed sign language data to a sign language recognition algorithm to generate corresponding text data.
[1913] Input: Preprocessed sign language data
[1914] Specific behavior: The algorithm converts the sign "hello" into text "hello."
[1915] Output: Generated text data
[1916] Step 4: Generate audio data
[1917] The terminal sends the generated text data to a speech synthesis engine and converts it into voice data.
[1918] Input: Generated text data
[1919] Specific operation: The speech synthesis engine converts "hello" into speech.
[1920] Output: Generated audio data
[1921] Step 5: Outputting audio data
[1922] The device will play the translated audio through the speaker.
[1923] Input: Generated audio data
[1924] Specific action: "Hello" is output from the speaker.
[1925] Output: Translation played back as audio
[1926] Step 6: Capture and analyze emotions
[1927] The device uses a camera and microphone to analyze the user's emotions.
[1928] Input: Captured sign language data, as well as the user's facial expressions and tone of voice
[1929] Specific behavior: Facial expression analysis algorithm determines emotions.
[1930] Output: Parsed emotion information
[1931] Step 7: Adjusting the audio data
[1932] The device adjusts the tone and content of the voice based on the analyzed emotional information.
[1933] Input: Parsed emotion information
[1934] Specific behavior: If the user has a happy expression, play "Hello" in a bright tone.
[1935] Output: Modified audio data
[1936] These are the specific processing steps of this system. Based on the data input at each step, it is possible to provide optimal output according to the user's situation and emotions.
[1937] (Application example 2)
[1938] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1939] In autonomous vehicles, there is a need to address the safety and information shortages faced by the hearing- and visually impaired. Another challenge is to provide optimal information in response to the user's emotions, enabling safer and more comfortable use.
[1940] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing ambient sounds, means for converting the captured voice data into text data, means for analyzing the type, direction, and distance of the sound, means for displaying the analyzed information, means for recognizing the user's emotions, and means for adjusting the display information based on the emotion information. This enables hearing-impaired and visually impaired people to use autonomous vehicles safely and comfortably.
[1941] An "ambient sound capturing means" is a device or mechanism capable of detecting sounds in the environment and recording them as digital data.
[1942] A "means for converting captured audio data into text data" is a mechanism or software that analyzes the recorded audio signal and converts it into corresponding text information.
[1943] "Means for analyzing sound type, direction and distance" means an algorithm or mechanism for calculating and identifying the source, type and distance of a sound from captured audio data.
[1944] "Means for displaying analyzed information" refers to a display or other display device for visually presenting the results of the analysis.
[1945] The "means for recognizing user emotions" is an engine or software for analyzing the user's tone of voice and facial expressions to determine their emotional state.
[1946] The "means for adjusting displayed information based on emotional information" refers to a mechanism or program that dynamically changes the content or emphasis of displayed information in response to a recognized emotional state.
[1947] A "means for capturing sign language" is a camera or sensor that captures and records the user's sign language actions.
[1948] A "means for converting captured sign language data into text data" is an algorithm or system for analyzing the recorded sign language movements and converting them into corresponding written information.
[1949] The "means for converting the converted text data into voice" refers to a voice synthesis device or software for outputting the character data as voice.
[1950] "Means for capturing surrounding environmental data" refers to devices or systems that use sensors and cameras to capture the conditions inside and around the vehicle.
[1951] The "means for analyzing captured environmental data" is an algorithm or program for extracting and analyzing useful information from the acquired environmental data.
[1952] The "means for generating the analyzed information as a voice guide" is an engine or program for generating voice guidance based on the analysis results.
[1953] The "means for outputting the generated audio guide" refers to a device for informing the user of the generated audio guide using a speaker or the like.
[1954] The "means for adjusting the tone and content of the audio description based on emotional information" refers to a system or algorithm for dynamically changing the tone and content of the audio description depending on the user's emotional state.
[1955] The present invention relates to an assistance system for the hearing impaired and the visually impaired that is applied to an autonomous vehicle. Specific embodiments of the present invention will be described below.
[1956] System configuration
[1957] The system consists of the user's mobile device (hereafter referred to as the device) and a cloud-based server (hereafter referred to as the server). The device is equipped with a microphone, camera, speaker, display, various sensors, and an emotion recognition engine. This hardware is used to collect and analyze information about the user's surroundings and emotional state, and provide the necessary information.
[1958] Hearing-impaired features
[1959] Subtitling of ambient sounds
[1960] 1. Sound capture and processing
[1961] The device uses a microphone (e.g., general hardware) to capture ambient sounds. The captured audio data undergoes pre-processing such as noise reduction within the device.
[1962] The preprocessed voice data is converted into text in real time using voice recognition software (e.g., Google Speech Recognition API).
[1963] 2. Analysis of sound type, direction, and distance
[1964] The device identifies the type of sound from the converted text information and sound characteristics, and uses the arrangement of multiple microphones to calculate the direction and distance from which the sound is coming.
[1965] 3. Displaying Information
[1966] The analyzed information is displayed in text format on the device's display, allowing the user to understand the surrounding situation in real time.
[1967] 4. Emotion recognition and regulation
[1968] The device uses a camera (e.g., general hardware) or microphone to capture the user's facial expressions and tone of voice, which are then analyzed by an emotion recognition engine (e.g., the EmotionRecognition library).
[1969] The system adjusts the displayed text and notification method based on the analyzed emotion information. For example, if a car horn is honking from the right at a distance of 10 meters, and the user shows signs of surprise, the display will be emphasized.
[1970] Features for the visually impaired
[1971] Audio guidance of surrounding objects and situations
[1972] 1. Environmental data capture and processing
[1973] The device uses a camera and various sensors to capture data on the surrounding environment, which is then subjected to image recognition preprocessing to extract object contours and feature points.
[1974] 2. Object Recognition and Analysis
[1975] The pre-processed data is then used by AI algorithms (e.g., TensorFlow) to recognize surrounding objects, analyze their location and movement, and extract important information.
[1976] 3. Audio guide generation and output
[1977] The analyzed information is generated as audio guidance and converted into speech by a speech synthesis engine (e.g., gTTS library). The device then notifies the user of the generated audio guidance through the speaker.
[1978] 4. Emotion recognition and regulation
[1979] The device captures the user's emotions and adjusts the voice guidance based on the analyzed emotion information, for example, softening the tone of the voice guidance if the user is nervous.
[1980] Automatic sign language recognition and translation
[1981] 1. Sign Language Capture and Processing
[1982] The device uses a camera to capture the user's sign language in real time, and the captured video data undergoes pre-processing to extract hand shapes and movements.
[1983] 2. Sign Language Recognition and Translation
[1984] The preprocessed video data is analyzed using a sign language recognition algorithm (e.g., OpenCV or TensorFlow) to generate corresponding text information, which is then converted into speech data by a speech synthesis engine.
[1985] 3. Audio Output
[1986] The device plays the translated audio through a speaker, helping people who don't understand sign language to communicate with the user.
[1987] 4. Emotion Recognition and Translation Adjustment
[1988] The device analyzes the user's emotions along with the captured sign language data, and adjusts the expression of the text translation and the tone of the voice based on the emotional information. For example, if the user signs "hello" and has a happy expression, the tone of the voice output will be brighter.
[1989] Prompt Sentence Examples
[1990] "Can you hear the horn now?", "There is an obstacle ahead", "If the user is nervous, generate voice prompts in a softer tone."
[1991] As a result, information and notifications are appropriately adjusted according to the user's emotional information, enabling hearing-impaired and visually-impaired people to use self-driving vehicles safely and comfortably.
[1992] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1993] Step 1:
[1994] Audio data capture and preprocessing
[1995] Input: Ambient sound
[1996] Output: Preprocessed audio data
[1997] The device uses a microphone to capture surrounding sounds. The captured audio data undergoes pre-processing such as noise reduction. This processing makes the audio data easier to analyze. The data is also updated at regular intervals.
[1998] Step 2:
[1999] Converting audio data to text
[2000] Input: Preprocessed audio data
[2001] Output: Text data
[2002] The device converts the pre-processed voice data into text in real time using speech recognition software (e.g., Google Speech Recognition API). This process captures the voice as text information.
[2003] Step 3:
[2004] Analysis of sound type, direction, and distance
[2005] Input: Text data and audio data features
[2006] Output: Information about the type, direction, and distance of the sound
[2007] The device identifies the type of sound from the converted text information and features of the voice data, and calculates the direction and distance from which the sound is coming using a multi-microphone arrangement, identifying, for example, the sound of an ambulance siren or car horn.
[2008] Step 4:
[2009] Viewing analysis information
[2010] Input: Information about the type, direction, and distance of the sound
[2011] Output: Text information displayed
[2012] The device displays the analyzed information in text format on the display, allowing the user to understand the surrounding situation in real time. For example, it may display information such as "There is a horn honking 10 meters from the right."
[2013] Step 5:
[2014] Emotion recognition
[2015] Input: Camera video, audio data (if necessary)
[2016] Output: Emotional information
[2017] The device uses a camera (and optionally a microphone) to capture the user's facial expressions and tone of voice, which are then analyzed by an emotion recognition engine (e.g., the EmotionRecognition library). In the process, the user's emotional state (such as surprise, tension, or joy) is determined.
[2018] Step 6:
[2019] Adjusting display and notification information
[2020] Input: Emotion information, analysis information
[2021] Output: Adjusted text information or audio notification
[2022] The device adjusts the displayed text and notification method based on the acquired emotional information. For example, if the user expresses surprise, the device will emphasize the text or display it in a more visually appealing way. If a voice notification is required, the device will use a speech synthesis engine to adjust the tone of the notification.
[2023] Step 7:
[2024] Sign language capture and translation
[2025] Input: Sign language video data
[2026] Output: Text data, audio data
[2027] The device uses a camera to capture the user's sign language in real time. The captured video data is preprocessed to extract hand movements and shapes. The sign language is analyzed using a sign language recognition algorithm (e.g., OpenCV, TensorFlow), and corresponding text information is generated and output as voice data using a speech synthesis engine.
[2028] Step 8:
[2029] Environmental Data Capture and Analysis
[2030] Input: Environmental data (video and sensor data)
[2031] Output: Analysis information
[2032] The device uses cameras and sensors to capture data about the surrounding environment, which is then pre-processed using image recognition to extract object contours and feature points, and AI algorithms are used to identify the object type and location, generating key analytics.
[2033] Step 9:
[2034] Audio description generation and output
[2035] Input: Analysis information and emotion information
[2036] Output: Voice prompt
[2037] The device generates audio guidance based on the analyzed information and converts it into voice data using a speech synthesis engine (e.g., gTTS). The generated audio guidance is then played back to the user through the speaker. The tone and content of the audio guidance are adjusted according to the user's emotional state.
[2038] Through the above steps, the present invention assists hearing-impaired and visually-impaired people in safely and comfortably using autonomous vehicles.
[2039] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[2040] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2041] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[2042] [Fourth embodiment]
[2043] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[2044] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[2045] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[2046] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[2047] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[2048] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[2049] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[2050] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[2051] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[2052] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[2053] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[2054] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[2055] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2056] The present invention is an assistance system for the hearing impaired and the visually impaired, and is implemented in the following manner.
[2057] System configuration
[2058] This system consists of a user's mobile device (hereafter referred to as the device) and a cloud-based server (hereafter referred to as the server). The device is equipped with a microphone, camera, speaker, display, and various sensors, and uses this hardware to collect and analyze information about the user's surroundings and provide the necessary information.
[2059] Hearing-impaired features
[2060] Subtitling of ambient sounds
[2061] 1. Sound capture and processing
[2062] The device uses a microphone to constantly capture ambient sounds.
[2063] The captured audio data undergoes pre-processing such as noise reduction inside the device.
[2064] The preprocessed voice data is converted into text in real time by a voice recognition engine.
[2065] 2. Analysis of sound type, direction, and distance
[2066] The device identifies the type of sound from the converted text information and sound characteristics, such as "ambulance siren" or "human voice."
[2067] The device uses multiple microphone arrangements to calculate the direction and distance of the sound and notify the user.
[2068] 3. Displaying Information
[2069] The analyzed information is displayed in text format on the device's display, allowing the user to understand the surrounding situation in real time.
[2070] 4. Specific Examples
[2071] For example, if a car horn sounds near the user, the device will display "The horn is sounding from 10 meters to the right."
[2072] Automatic sign language recognition and translation
[2073] 1. Sign Language Capture and Processing
[2074] The device uses a camera to capture the user's sign language in real time.
[2075] The captured video data is pre-processed to extract the shape and movement of the hand.
[2076] 2. Sign Language Recognition and Translation
[2077] The pre-processed video data is analyzed by a sign language recognition algorithm to generate corresponding text information.
[2078] The generated text information is converted into voice data by a voice synthesis engine.
[2079] 3. Audio Output
[2080] The device plays the translated audio through a speaker, helping people who don't understand sign language to communicate with the user.
[2081] 4. Specific Examples
[2082] For example, if a user says "hello" in sign language, the terminal recognizes the sign and outputs "hello" aloud.
[2083] Features for the visually impaired
[2084] Audio guidance of surrounding objects and situations
[2085] 1. Environmental data capture and processing
[2086] The device uses a camera and various sensors to capture data about the surrounding environment.
[2087] The captured data is subjected to image recognition preprocessing to extract the contours and feature points of the object.
[2088] 2. Object Recognition and Analysis
[2089] The preprocessed data is then used by AI algorithms to recognize surrounding objects and situations.
[2090] The location and movement of recognized objects are analyzed to pick out important information.
[2091] 3. Audio guide generation and output
[2092] The analyzed information is generated as audio guide and converted into voice by a voice synthesis engine.
[2093] The terminal notifies the user of the generated audio guide through a speaker.
[2094] 4. Specific Examples
[2095] If the user is walking and there is an obstacle one meter ahead, the device will announce with a voice message, "There is an obstacle one meter ahead."
[2096] Starting and shutting down the system
[2097] 1. Initial Setup
[2098] During initial setup, the user selects the system mode (hearing impaired mode, visually impaired mode) that best suits their needs.
[2099] The server receives the user's configuration information and prepares the required services.
[2100] 2. Service launch
[2101] The device will activate the necessary sensors (microphone, camera, etc.) based on the initial settings and begin collecting data.
[2102] The server analyzes the data sent from the terminal and returns the necessary information in real time.
[2103] The terminal displays or outputs the analysis results to the user.
[2104] 3. Termination or Suspension of Service
[2105] If a user wishes to stop using the service, they can select the option to terminate or suspend the service in their device's settings screen.
[2106] The terminal receives the service termination command, shuts down all sensors, and cuts off communication with the server.
[2107] This system provides useful support to the hearing-impaired and visually impaired to overcome difficulties in daily life and live safely and independently.
[2108] The processing flow will be explained below.
[2109] Features for the hearing impaired: Subtitling of surrounding sounds
[2110] Step 1: Capture the sound
[2111] The device activates the microphone and captures ambient sounds at 0.1 second intervals.
[2112] Step 2: Preprocessing the audio data
[2113] The device performs noise reduction and normalization on the captured audio data.
[2114] Step 3: Speech to text
[2115] The device sends the preprocessed voice data to a voice recognition engine, which converts it into text data in real time.
[2116] Step 4: Identify the type of sound
[2117] The device extracts sound characteristics from the text data and identifies the type of sound based on them (e.g., "ambulance siren," "human voice," etc.).
[2118] Step 5: Analyze sound direction and distance
[2119] The device uses multiple microphone arrangements to calculate the direction and distance from which the sound is coming.
[2120] Step 6: Display subtitles
[2121] The device displays text information on the display, informing the user of the type, direction, and distance of the sound.
[2122] As a specific example, it displays "The sound of a horn is coming from 10 meters to the right."
[2123] Features for the visually impaired: Audio guide to surrounding objects and situations
[2124] Step 1: Capture environment data
[2125] The device will activate its camera and sensors to capture data on the surrounding environment every 0.5 seconds.
[2126] Step 2: Preprocessing the data
[2127] The device performs image recognition preprocessing on the captured video data to extract the contours and feature points of objects.
[2128] Step 3: Object Recognition
[2129] The device sends the pre-processed data to AI algorithms to recognize surrounding objects.
[2130] Step 4: Obtaining information about the object
[2131] The device analyzes the location and movement of recognized objects and picks out important information (e.g., obstacles ahead, approaching people).
[2132] Step 5: Generate audio guide
[2133] The device generates the analyzed information as a text message, sends it to a speech synthesis engine, and generates audio guidance.
[2134] Step 6: Audio Notifications
[2135] The terminal notifies the user of the generated audio guide through a speaker.
[2136] As a specific example, a voice message will be displayed saying, "There is an obstacle one meter ahead."
[2137] Starting and shutting down the system
[2138] Step 1: Initial Setup
[2139] The user configures the system and selects the mode that best suits their needs (e.g., hearing-impaired mode, visually-impaired mode).
[2140] The server receives the user's configuration information and prepares to provide the required services.
[2141] Step 2: Real-time analytics and notifications
[2142] Based on the selected mode, the device will activate the necessary sensors (microphone, camera, etc.) and collect data.
[2143] The server analyzes the data sent from the terminal and returns the necessary information in real time.
[2144] The device notifies the user of the analysis results in an appropriate format (subtitles, audio).
[2145] Step 3: Terminate or suspend service
[2146] If a user wants to stop using the system, they can select the option to terminate or suspend the service on their device's settings screen.
[2147] The terminal receives the service termination command, shuts down all sensors, and cuts off communication with the server.
[2148] Example 1
[2149] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2150] The purpose of this invention is to provide support for the hearing-impaired and visually-impaired to overcome the difficulties they face in daily life and live safely and independently. Specifically, the invention involves the development of a system that recognizes surrounding sounds, sign language, and environmental conditions in real time and provides the necessary information in an appropriate format.
[2151] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[2152] In this invention, the server includes means for capturing ambient sounds, means for preprocessing the captured audio data, means for converting the preprocessed audio data into text data, means for analyzing the type, direction, and distance of sound, means for displaying the analyzed information, means for capturing sign language, means for preprocessing the captured sign language data, means for converting the preprocessed sign language data into text data, means for converting the converted text data into audio data and outputting it, means for capturing ambient environmental data, means for preprocessing the captured environmental data, means for analyzing the preprocessed environmental data to recognize objects, means for generating audio guidance based on the recognized object information, and means for outputting the generated audio guidance. This enables hearing-impaired and visually impaired people to grasp their surroundings in real time and live safely and independently.
[2153] The "means for capturing ambient sound" refers to a device or sensor function for collecting audio data from the user's surrounding environment.
[2154] "Means for preprocessing captured audio data" refers to processes or algorithms used to remove noise from collected audio data and prepare it for speech recognition.
[2155] The "means for converting preprocessed voice data into text data" refers to software or a system for converting voice data into text information using voice recognition technology.
[2156] "Means for analyzing the type, direction and distance of sound" refers to an algorithm or device that analyzes text data and sound feature data to identify the source, type, direction and distance of a sound.
[2157] "Means for displaying analyzed information" refers to a device or system that uses a display, screen, or the like to visually convey the analysis results to the user.
[2158] A "means for capturing sign language" is a camera or sensor that records the user's hand movements and shapes in real time.
[2159] "Means for pre-processing captured sign language data" refers to a process or system that appropriately processes captured video data for sign language recognition and extracts hand shapes and movements.
[2160] A "means for converting preprocessed sign language data into text data" is a technology or system that analyzes captured sign language movements and converts them into corresponding text information.
[2161] "Means for converting the converted text data into audio data and outputting it" refers to software or a device for converting character information into audio and reproducing it via a speaker or the like.
[2162] The "means for capturing surrounding environmental data" refers to a device or function for collecting information about the user's surrounding environment using a camera or sensor.
[2163] "Means for pre-processing captured environmental data" refers to processes or algorithms that process collected environmental data into a form suitable for image recognition.
[2164] "Means for analyzing preprocessed environmental data to recognize objects" refers to AI algorithms or systems that identify and recognize the shape and position of objects from environmental data.
[2165] The "means for generating audio guidance based on recognized object information" refers to a system or process for creating audio guidance to be notified to the user based on information about the identified object.
[2166] The "means for outputting the generated audio guide" refers to a device or system for conveying the generated audio guide to the user using a speaker or the like.
[2167] This invention is an assistance system for hearing-impaired and visually-impaired people to overcome difficulties in daily life, and is implemented as a system including a user's mobile terminal (hereinafter referred to as "terminal") and a cloud-based server (hereinafter referred to as "server"). The terminal is equipped with a microphone, camera, speaker, display, and various sensors. These hardware components are used to collect and analyze information about the user's surroundings and provide the necessary information.
[2168] Hearing-impaired features
[2169] Subtitling of ambient sounds
[2170] The device uses a microphone to capture ambient sounds and preprocesses the audio data. It then uses noise reduction technology to remove background noise and converts the audio into text using a speech recognition engine. The device uses multiple microphone arrangements to analyze the type, direction, and distance of the sound and displays the analyzed information on the display. For example, if a user is walking on the sidewalk and hears a car horn honking from the right at a distance of 10 meters, the device will display the message, "A car horn is honking from the right at a distance of 10 meters."
[2171] Automatic sign language recognition and translation
[2172] The device uses a camera to capture the user's sign language in real time and preprocesses the video data. After extracting the hand shape and movement, a sign language recognition algorithm generates corresponding text information. Next, a speech synthesis engine converts the text information into audio data and outputs it from the speaker. For example, if the user signs "hello," the device will recognize the sign and output "hello" aloud.
[2173] Features for the visually impaired
[2174] Audio guidance of surrounding objects and situations
[2175] The device uses cameras and various sensors to capture and preprocess data on the surrounding environment. It uses image recognition technology to extract the contours and feature points of objects, and uses AI algorithms to recognize surrounding objects and situations. It analyzes the location and movement of recognized objects and generates important information as audio guidance. This audio guidance is converted into voice by a speech synthesis engine and notified to the user through the speaker. For example, if a user is walking and there is an obstacle one meter ahead, the device will announce, "There is an obstacle one meter ahead."
[2176] Starting and shutting down the system
[2177] During the initial setup, the user selects the system mode (deaf mode, visually impaired mode) that best suits their needs. The server receives the user's configuration information and prepares the required services. The device activates the required sensors (microphone, camera, etc.) based on the initial setup and begins collecting data. The server analyzes the data sent from the device and returns the required information in real time. If the user wants to terminate or pause the service, they select the option to terminate or pause the service on the device's settings screen. The device receives the service termination command, stops all sensors, and cuts off communication with the server.
[2178] Examples of concrete examples and prompts
[2179] Example prompt for speech recognition: "Transcribe the speech that says 'hello' to text."
[2180] Example prompt for sign language recognition: "Translate this sign language video into text."
[2181] This provides useful support to hearing-impaired and visually impaired people to overcome difficulties in daily life and live safely and independently.
[2182] The flow of the identification process in the first embodiment will be described with reference to FIG.
[2183] Hearing-impaired features
[2184] Subtitling of ambient sounds
[2185] Step 1: Capture the sound
[2186] Subject: Device
[2187] How it works: The device uses a microphone to capture ambient sound, taking ambient audio data as input and raw captured data as output.
[2188] Input and Output: The input is the ambient audio data, and the output is the captured audio data.
[2189] Step 2: Pre-processing the sound
[2190] Subject: Device
[2191] Specific operation: The device preprocesses the captured audio data and removes noise. The input is the captured audio data, and the output is the noise-removed audio data.
[2192] Input and Output: The input is the captured audio data, and the output is the pre-processed audio data.
[2193] Step 3: Voice Recognition
[2194] Subject: Device
[2195] Specific operation: The device sends the preprocessed voice data to the voice recognition engine and converts it into text data. The preprocessed voice data is input, and the converted text data is obtained as output.
[2196] Input and Output: The input is preprocessed audio data, and the output is text data.
[2197] Step 4: Analyzing the sound information
[2198] Subject: Device
[2199] Specific operation: The device analyzes the text data and sound feature data to identify the type, direction, and distance of the sound. The text data and sound feature data are input, and the analysis results are obtained as output.
[2200] Input and output: The input is text data and sound feature data, and the output is analyzed sound information.
[2201] Step 5: Viewing information
[2202] Subject: Device
[2203] Specific operation: The terminal displays the analyzed information in text format on the display. The analyzed sound information is input and displayed on the display as output.
[2204] Input and output: The input is the analyzed sound information, and the output is the text information displayed on the screen.
[2205] Automatic sign language recognition and translation
[2206] Step 1: Capture sign language
[2207] Subject: Device
[2208] Specific operation: The device uses a camera to capture the user's sign language actions in real time. The sign language actions are input and the captured video data is obtained as output.
[2209] Input and Output: The input is sign language actions, and the output is captured video data.
[2210] Step 2: Preprocessing the video data
[2211] Subject: Device
[2212] Specific operation: The device preprocesses the captured video data and extracts the hand shape and movement. The captured video data is input and the preprocessed video data is output.
[2213] Input and Output: The input is the captured video data, and the output is the pre-processed video data.
[2214] Step 3: Sign Language Recognition and Text Conversion
[2215] Subject: Device
[2216] Specific operation: The device sends the preprocessed video data to a sign language recognition algorithm to generate corresponding text data. The preprocessed video data is input, and the generated text data is output.
[2217] Input and Output: The input is the preprocessed video data, and the output is the generated text data.
[2218] Step 4: Speech synthesis and output
[2219] Subject: Device
[2220] Specific operation: The device sends the generated text data to a speech synthesis engine, converts it into voice data, and outputs the converted voice data to the user through the speaker. Text data is input, and voice data is obtained as output.
[2221] Input and output: The input is the generated text data, and the output is the audio output from the speaker.
[2222] Features for the visually impaired
[2223] Audio guidance of surrounding objects and situations
[2224] Step 1: Capture environment data
[2225] Subject: Device
[2226] Specific operation: The device uses a camera and various sensors to capture surrounding environmental data. The surrounding environmental information is input, and the captured environmental data is output.
[2227] Input and Output: The input is the surrounding environment information, and the output is the captured environment data.
[2228] Step 2: Preprocessing the image data
[2229] Subject: Device
[2230] Specific operation: The device preprocesses the captured environmental data and extracts object contours and feature points. The captured environmental data is input, and the preprocessed environmental data is output.
[2231] Input and Output: The input is the captured environmental data, and the output is the preprocessed environmental data.
[2232] Step 3: Object Recognition
[2233] Subject: Device
[2234] How it works: The device analyzes the preprocessed data using AI algorithms to recognize surrounding objects and situations. The preprocessed data is input, and the recognized object information is output.
[2235] Input and Output: The input is the preprocessed data, and the output is the recognized object information.
[2236] Step 4: Location Analysis
[2237] Subject: Device
[2238] Specific operation: The device analyzes the location and movement of the recognized object and extracts important information. The recognized object information is input, and the location and movement analysis results are obtained as output.
[2239] Input and Output: The input is the recognized object information, and the output is the analysis result of the position and movement.
[2240] Step 5: Audio guide generation and output
[2241] Subject: Device
[2242] Specific operation: The device generates the analyzed information as an audio guide and converts it into speech using a speech synthesis engine. The converted audio guide is then notified to the user through the speaker. The analyzed information is input, and the audio guide is obtained as output.
[2243] Input and Output: The input is the parsed information, and the output is the audio guide from the speaker.
[2244] Starting and shutting down the system
[2245] Step 1: Initial Setup
[2246] Subject: User
[2247] Specific operation: The user selects the system mode (deaf mode, visually impaired mode) according to their needs in the initial setting. The user's needs are input, and the setting information is sent to the server as output.
[2248] Input and Output: Input is user needs, output is configuration information.
[2249] Step 2: Start the service
[2250] Subject: Server
[2251] Specific operation: The server receives the user's configuration information and prepares the necessary services. The device activates the necessary sensors (microphone, camera, etc.) based on the initial settings and begins collecting data. The configuration information is received as input and the service is started as output.
[2252] Input and Output: Input is the configuration information, output is the started service.
[2253] Step 3: Data collection and analysis
[2254] Subject: Device
[2255] Specific operation: The terminal starts collecting data, and the server analyzes the data sent from the terminal. The collected data is input, and the analysis results are obtained as output.
[2256] Input and Output: Input is the collected data and output is the analysis result.
[2257] Step 4: Notification of information
[2258] Subject: Device
[2259] Specific operation: The device displays or outputs the analysis results to the user. The analysis results are input and notified to the user as output.
[2260] Input and output: The input is the analysis result, and the output is the notification to the user.
[2261] Step 5: Terminate or suspend service
[2262] Subject: User
[2263] Specific behavior: When a user wants to stop using a service, they select the option to terminate or pause the service in the device settings screen. The user's operation is the input, and the service is terminated or paused as the output.
[2264] Input and Output: Input is a user action, output is a service that has been terminated or suspended.
[2265] (Application example 1)
[2266] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2267] The purpose of this invention is to provide a means to resolve the lack of information and communication difficulties faced by hearing-impaired and visually-impaired people when shopping comfortably and safely in a brick-and-mortar store. Specifically, the objective is to provide a system that supports people with disabilities to act independently by acquiring in-store announcements, conversations with store staff, and product information in real time and conveying this information appropriately to the user.
[2268] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[2269] In this invention, the server includes means for capturing ambient sounds, means for converting the captured audio data into text data, means for a...
Claims
1. a means for capturing ambient sounds; means for converting the captured audio data into text data; a means for analyzing the type, direction and distance of a sound; a means for displaying the analyzed information; A system including:
2. a means for capturing sign language; means for converting the captured sign language data into text data; means for converting the converted text data into voice; The system of claim 1 , comprising:
3. means for capturing ambient environmental data; means for analyzing the captured environmental data; a means for generating the analyzed information as an audio guide; means for outputting the generated audio guide; The system of claim 1 , comprising:
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A