System

A smartphone application enhances the safety and independence of visually and hearing impaired individuals by detecting obstacles and signs from camera footage, converting information into audio, and providing environmental sound notifications, addressing the limitations of current technologies.

JP2026023482APending Publication Date: 2026-02-13SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024125417
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-31
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Visually and hearing impaired individuals face challenges in recognizing obstacles and environmental sounds, limiting their safety and independence due to the lack of effective assistance technologies and societal barriers.

Method used

A smartphone application that utilizes camera images for obstacle and sign detection, converts this information into audio, records and identifies important environmental sounds, and provides notifications through vibration and text display, enhancing user independence and social participation.

Benefits of technology

Enables visually and hearing impaired users to navigate safely and respond to important sounds, promoting their independence and social integration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026023482000001_ABST
    Figure 2026023482000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A smartphone application for supporting a daily life of a visually or aurally handicapped user, comprising: means for detecting an obstacle and a sign from a camera video; means for converting information on the detected obstacle and sign into a sound; means for recording an environmental sound and identifying a specific sound; and means for notifying the identified specific sound by vibration and character display.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] The current lack of assistance dogs for the visually and hearing impaired reduces the safety and convenience of their daily lives. In particular, the visually impaired have difficulty recognizing obstacles and signs, and the hearing impaired cannot hear environmental sounds or people calling out to them, making it difficult for them to lead independent lives and participate in society. Furthermore, the lack of funding and trainers for training assistance dogs, as well as the refusal of public facilities and stores to accept assistance dogs, exacerbate these issues. [Means for solving the problem]

[0005] This invention is a smartphone application that supports the daily lives of users with visual or hearing impairments. It provides a means for detecting obstacles and signs from camera images, a means for converting detected obstacle and sign information into audio, a means for recording environmental sounds and identifying specific sounds, and a means for notifying users of identified specific sounds by vibration and text display. This application allows visually impaired people to safely obtain information about their surroundings, and enables hearing impaired people to recognize important sounds and respond appropriately. Furthermore, the application also includes functions for recognizing and executing voice commands and generating visual alerts when important sounds are identified, comprehensively promoting users' independence and social participation.

[0006] "Visually or hearing impaired user" refers to a person who has a visual or hearing impairment and is unable to fully use those senses in a normal living environment.

[0007] A "smartphone application" is a software program that runs on a smartphone and is designed to provide specific functions or services.

[0008] "Camera footage" refers to video data captured from a smartphone or other camera device.

[0009] "Obstacle" refers to a physical object that impedes a user's movement or activity.

[0010] "Sign" refers to a display that provides visual information, and includes traffic signs, guide boards, etc.

[0011] "Convert to speech" refers to converting text information or digital data into a format that can be played back as speech using speech synthesis technology.

[0012] "Environmental sounds" refers to all sounds that occur within the user's environment, including background sounds and specific event sounds.

[0013] A "particular sound" refers to a sound that is identified as important or unique to the user among environmental sounds.

[0014] "Vibration" refers to the ability of a device to vibrate and provide haptic feedback to the user.

[0015] "Text display" refers to the display of information in text format on a display.

[0016] "Voice command" refers to an input method in which a user issues a command to a device by voice, causing the device to perform a specific action in response.

[0017] "Visual alert" refers to a notification that draws the user's attention by visual means, including a flashing screen or a pop-up message. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] This is a smartphone application that supports the daily lives of users with visual or hearing impairments. This application includes a means for detecting obstacles and signs from camera images and converting that information into audio, and a means for recording environmental sounds, identifying specific sounds, and notifying them with vibrations and text, thereby promoting the user's independence and participation in society.

[0040] Detects obstacles and signs from camera footage and provides audio notifications

[0041] 1. Capture camera footage:

[0042] The device uses a camera to capture real-time images that show the user's surroundings and are used to detect obstacles and signs.

[0043] 2. Obstacle and sign detection:

[0044] The device analyzes the captured video to detect obstacles and signs. Image processing and optical character recognition (OCR) technologies are used to analyze the video. Specifically, the video is converted to grayscale and text information is extracted.

[0045] 3. Speech synthesis and notifications:

[0046] The device receives information about detected obstacles and signs as text and converts it into speech using speech synthesis technology. This speech is then provided to the user via earphones or speakers, providing them with information on how to act safely.

[0047] Environmental sound recording and important sound notifications

[0048] 1. Environmental Sound Recording:

[0049] The device uses a microphone to record ambient sounds, and the recorded audio data is used to detect certain important sounds.

[0050] 2. Identifying specific sounds:

[0051] The device uses voice recognition technology to analyze recorded audio data and identify specific sounds such as a horn, a bicycle bell, a kettle, a baby crying, or someone calling your name in a public place.

[0052] 3. Vibration and text notification:

[0053] Based on the identified specific sound, the device will vibrate and display the notification content as text on the display, allowing hearing-impaired people to respond appropriately to important sounds.

[0054] Specific examples

[0055] 1. Support for visually impaired people walking on the road:

[0056] When a user is walking on the sidewalk, the device uses its camera to capture images of the area ahead and uses OCR to detect "STOP" signs. This information is converted into "There is a STOP sign ahead" using speech synthesis technology, and the user receives a voice notification through their earphones, providing them with information on how to act safely.

[0057] 2. Support for the hearing impaired in noisy cafes:

[0058] While a user is waiting to order at a cafe, the device's microphone records ambient sounds and uses voice recognition technology to recognize that "your name has been called." As a result, the device notifies the user by vibrating and displays "Your name has been called" on the display so that the user is aware of the call.

[0059] As a result, the present invention can support the daily lives of users with visual or hearing impairments, and promote their independence and social participation.

[0060] The processing flow will be explained below.

[0061] Detects obstacles and signs from camera footage and provides audio notifications

[0062] Step 1:

[0063] The device activates the camera and captures video in real time, specifically by using the camera device interface to obtain video frames.

[0064] Step 2:

[0065] The device preprocesses the captured video frames, specifically converting the video to grayscale to improve the efficiency of image processing.

[0066] Step 3:

[0067] The device detects obstacles and signs from the pre-processed video and uses optical character recognition (OCR) technology to extract text information from the video.

[0068] Step 4:

[0069] The device converts the extracted text information into speech, using speech synthesis technology to generate natural-sounding speech data.

[0070] Step 5:

[0071] The device plays the generated audio data, providing the user with an audio notification through a speaker or earphones.

[0072] Environmental sound recording and important sound notifications

[0073] Step 1:

[0074] The device will activate the microphone and record the surrounding environmental sounds at regular intervals.

[0075] Step 2:

[0076] The device analyzes the recorded voice data and uses voice recognition technology to extract text information from the recording.

[0077] Step 3:

[0078] The device analyzes the extracted text information and identifies important sounds (e.g., horns, bicycle bells, etc.) and prepares corresponding actions when a particular sound is detected.

[0079] Step 4:

[0080] The device will vibrate when an important sound is identified, notifying the user.

[0081] Step 5:

[0082] The device displays the notification content as text on the display, allowing the user to visually confirm the notification.

[0083] Example 1

[0084] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0085] Users with visual or hearing impairments have difficulty recognizing obstacles and signs, and distinguishing environmental sounds in their daily lives, which can limit their ability to travel safely and respond to emergencies. While technologies with visual and hearing assistance are necessary for these users to lead independent lives, currently available technologies do not adequately meet these needs. Specifically, improvements are needed in areas such as image recognition accuracy, prompt voice notifications, and real-time recognition of environmental sounds.

[0086] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0087] In this invention, the server includes means for detecting obstacles and signs from camera images, means for converting information about detected obstacles and signs into audio, means for recording environmental sounds and identifying specific sounds, means for performing grayscale conversion and optical character recognition through video analysis, means for converting text information into audio using speech synthesis technology, and means for identifying important sounds using speech recognition technology, thereby enabling users with visual or hearing impairments to travel safely and respond to emergencies.

[0088] "Camera footage" refers to real-time visual data acquired using the device's camera.

[0089] An "obstacle" is an object that exists in the user's direction of travel or around the user, and that obstructs the user's movement or actions.

[0090] A "sign" is a sign placed on a road or in a public place that displays letters or symbols to convey specific instructions or warnings.

[0091] "Means for converting into speech" refers to technology for converting text information into auditory information, and is a device or software that utilizes speech synthesis technology.

[0092] "Environmental sounds" are all sounds that exist around the user, and are audio data collected by a recording device.

[0093] A "specific sound" is a specific important sound that is identified from among environmental sounds, and is a sound that the user needs to recognize immediately.

[0094] "Vibration notification" refers to a technology that allows a device to generate specific vibrations to convey information to the user.

[0095] "Text display" is a means of visually providing information to a user by displaying text information on a terminal display.

[0096] "Grayscale conversion" is an image processing technique that converts camera images into images that contain only black and white shades.

[0097] Optical character recognition (OCR) is a technology that extracts character information from an image and converts it into text data.

[0098] "Speech synthesis technology" is a technology that converts text data into human speech and produces speech.

[0099] "Speech recognition technology" is a technology that analyzes recorded voice data and understands the meaning of language.

[0100] An "important sound" is a sound that should immediately draw the user's attention, such as a car horn, a baby crying, or someone calling your name.

[0101] The present invention relates to a smartphone application for supporting the daily lives of users with visual or hearing impairments. This application combines multiple technical means to help users live safely and independently.

[0102] Hardware and Software Configuration

[0103] 1. The device uses a smartphone as its main platform, which is equipped with a camera, microphone, speaker / earphone, vibration motor, and display.

[0104] 2. The server uses cloud services (e.g., Amazon Web Services, Google Cloud Platform) for data processing and analysis, which enables it to process large amounts of data in real time and provide users with timely information.

[0105] 3. The device includes the following major software components:

[0106] Image processing algorithms (e.g. OpenCV)

[0107] Optical character recognition (OCR) technology (e.g. Tesseract OCR)

[0108] Speech synthesis technology (e.g., Google Text-to-Speech, Amazon Polly)

[0109] Speech recognition technology (e.g., Google Cloud Speech-to-Text, Amazon Transcribe)

[0110] Data processing and calculation

[0111] 1. Capture camera footage:

[0112] The device uses the smartphone's camera to capture real-time images of the surrounding area, which are then stored in the device's internal memory and analyzed immediately.

[0113] 2. Obstacle and sign detection:

[0114] The device uses OpenCV to convert the video to grayscale and apply an edge detection algorithm, then uses Tesseract OCR to analyze the characters on the sign and extract the text information.

[0115] 3. Speech synthesis and notifications:

[0116] The device converts the extracted text information into speech using voice synthesis technologies such as Google Text-to-Speech or Amazon Polly, and the speech is then transmitted to the user through earphones or speakers.

[0117] 4. Recording environmental sounds and identifying specific sounds:

[0118] The device uses the smartphone's microphone to record surrounding sounds, which are then stored in the device's internal memory and analyzed using voice recognition technologies such as Google Cloud Speech-to-Text and Amazon Transcribe.

[0119] 5. Vibration and text notification:

[0120] The device will vibrate based on the identified specific sound and display the notification content as text on the display, allowing the user to immediately recognize important information.

[0121] Specific examples

[0122] 1. For a visually impaired person walking on the road:

[0123] As a user walks along the sidewalk, the device uses its camera to capture images of the area ahead. It uses OpenCV's image processing algorithms to recognize obstacles and "STOP" signs, converting that information into text using Tesseract OCR. It then uses Google Text-to-Speech technology to translate the text into audio, saying "There is a STOP sign ahead," and notifies the user through earphones.

[0124] 2. For a hearing impaired person in a noisy cafe:

[0125] While a user is waiting to order at a cafe, the device's microphone records surrounding sounds. Using Google Cloud Speech-to-Text technology, the device identifies the voice saying "Mr. / Ms. XX, your order is ready," notifying the user with a vibration and displaying "Your name has been called" on the display.

[0126] Prompt Sentence Examples

[0127] By inputting prompt statements such as the following into the generative AI model, an explanation of the system and specific examples can be generated.

[0128] Describe a smartphone application system that supports the daily lives of users with visual or hearing impairments. Explain in detail how it detects obstacles and signs from camera footage and converts that information into audio, and how it records environmental sounds, identifies specific sounds, and notifies the user by vibrating and displaying text. Give specific examples.

[0129] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0130] Step 1: Capture camera footage:

[0131] The device activates the smartphone camera and captures images of the surroundings in real time. The camera captures the scenery ahead as the user walks along the sidewalk. This image data is stored in the internal memory. The input is an image of the user's surroundings, and the output is real-time image data.

[0132] Step 2: Obstacle and sign detection:

[0133] The device uses OpenCV to analyze the captured video. Specifically, it converts the video to grayscale and applies an edge detection algorithm (Canny edge detection). After this, it uses Tesseract OCR to extract the text information of the sign from the video. The input is real-time video data, and the output is text information (e.g., a "STOP" sign).

[0134] Step 3: Text-to-Speech and Notifications:

[0135] The device converts the extracted text information into audio data using the Google Text-to-Speech API. For example, audio data such as "There is a STOP sign ahead" is generated. If the user is using earphones, the audio notification is transmitted through the earphones. The input is text information, and the output is audio data.

[0136] Step 4: Recording ambient sounds:

[0137] The device activates the smartphone's microphone and records the surrounding environmental sounds in real time. For example, it records the surrounding sounds when the user is in a cafe. This audio data is stored in the internal memory. The input is the user's surrounding sounds, and the output is audio data.

[0138] Step 5: Identifying specific sounds:

[0139] The device analyzes the recorded voice data using the Google Cloud Speech-to-Text API to identify specific important sounds (e.g., "Mr. / Ms. XX, your order is ready"). The input is the recorded voice data, and the output is the identified text information.

[0140] Step 6: Vibration and text notification:

[0141] The device vibrates based on the identified specific sound and displays the text "Your name has been called" on the display, thereby alerting the user to the important notification. The input is the identified text information, and the output is a notification via vibration and text display.

[0142] This allows users with visual or hearing impairments to move safely and quickly perceive important information.

[0143] (Application example 1)

[0144] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0145] Users with visual or hearing impairments face challenges in safely moving around in physical stores and receiving the information they need. In such situations, users may bump into obstacles or miss important announcements, limiting their independent movement. As a result, users with visual or hearing impairments face challenges in limiting their opportunities to participate in social life.

[0146] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0147] In this invention, the server includes means for detecting obstacles and signs from camera images, means for converting information about detected obstacles and signs into audio, means for recording environmental sounds and identifying specific sounds, means for detecting obstacles and guide signs in the physical store and notifying them by audio, and means for recording announcements in the physical store and notifying identified announcements by vibration and text display, thereby enabling users with visual or hearing impairments to move around the physical store safely and efficiently.

[0148] A "visually or hearing impaired user" is an individual who is visually or hearing impaired, or both.

[0149] An "information processing device" is a machine or device for inputting, processing, and outputting data.

[0150] "Camera video" refers to real-time video data acquired using a camera.

[0151] An "obstacle" is a physical object that blocks the user's movement.

[0152] A "sign" is a visual display that conveys information.

[0153] "Converting to voice" refers to converting text information into voice data using voice synthesis technology.

[0154] "Ambient sounds" are any sounds occurring around the user.

[0155] "Identifying specific sounds" means recognizing and extracting characteristic sounds from recorded audio data.

[0156] "Vibration" is a means of providing tactile feedback to the user by causing the device to vibrate.

[0157] "Character display" means displaying text information on a display or the like.

[0158] A "physical store" is a store located in a physical location that offers goods and services.

[0159] An "announcement" is an audio message intended to convey information in a public place or specific environment.

[0160] "Recording" means to record audio using a device such as a microphone.

[0161] "Natural language processing" is a computational technique for processing and understanding human language.

[0162] The present invention relates to an information processing device that supports users with visual or hearing impairments in moving around safely and efficiently in a physical store. This device analyzes camera images and environmental sounds and provides the detected information to the user through means such as voice, vibration, and text display.

[0163] Hardware and Software Configuration

[0164] The present invention uses the following hardware and software.

[0165] Hardware:

[0166] Camera: Used to capture video in real time.

[0167] Microphone: Used to record ambient sounds and announcements.

[0168] Smartphone or tablet: The platform for processing and notification.

[0169] Headphones or speakers: Used to provide audio notifications.

[0170] Display: Used to display text information.

[0171] software:

[0172] OpenCV (open source image processing library): Used to analyze camera footage and detect obstacles and signs.

[0173] PyTesseract (OCR library): Used to extract text information from captured video.

[0174] Pyttsx3 (speech synthesis library): Used to convert text information into speech.

[0175] SpeechRecognition (speech recognition library): Used to identify specific sounds and announcements from recorded audio data.

[0176] GTTs (Google Text-to-Speech): The identified voice information is further analyzed and used to generate the required notifications.

[0177] Vibrate Library: Used to vibrate when a specific sound is identified.

[0178] Specific processing steps

[0179] 1. Camera footage capture and analysis:

[0180] The device's camera captures video in real time and converts it to grayscale using OpenCV. PyTesseract is used to extract text information from the image and detect obstacles and signs. For example, signs such as "STOP," "toilet," and "exit" can be detected in a physical store, and speech synthesis technology (Pyttsx3) can be used to notify the user, such as "There is a STOP sign ahead."

[0181] 2. Environmental sound recording and analysis:

[0182] The device's microphone records environmental sounds, and the recorded audio data is analyzed using SpeechRecognition. For example, when an in-store announcement or a specific event (such as a tasting invitation) is detected, the vibration library is used to notify the user by vibrating, such as "Tasting invitation." At the same time, the text information "Tasting invitation" is displayed on the display.

[0183] Adding specific examples

[0184] As an example, the following prompt sentence can be used:

[0185] Example prompt sentence:

[0186] It uses cameras to detect footage as you walk through the store and notifies you of "STOP" signs via voice synthesis.

[0187] When a user is walking through a supermarket, the camera footage is analyzed to detect a "STOP" sign, and an audio notification is played from the smartphone speaker saying, "There is a STOP sign ahead."

[0188] It records ambient sounds within the store, identifies announcements from specific areas (such as the food sampling area), and notifies you with a vibration.

[0189] The smartphone's microphone is used to record in-store announcements, and voice recognition technology is used to identify specific announcements such as "Invitation to sample food" and notify the user by vibrating.

[0190] As a result, the present invention enables users with visual or hearing impairments to move around safely and efficiently within physical stores, promoting their independence and social participation.

[0191] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0192] Step 1:

[0193] Capture footage with the camera.

[0194] Input: Capture real-time video using the device's camera.

[0195] Data processing: This video data will be used as is in the next step.

[0196] Output: The video data sent to the next step.

[0197] Specific operation: The device's camera continuously captures images of the user's surroundings in real time.

[0198] Step 2:

[0199] The video data is converted to grayscale and text information is extracted.

[0200] Input: Video data obtained in step 1.

[0201] Data processing: Convert the video to grayscale using OpenCV and extract text information using PyTesseract.

[0202] Output: The extracted text data.

[0203] Specific operation: The device uses the OpenCV library to convert the video data to grayscale, and then performs OCR processing using PyTesseract.

[0204] Step 3:

[0205] Obstacles and signs are detected from the extracted text information.

[0206] Input: The text data extracted in step 2.

[0207] Data calculation: Detect specific keywords such as "STOP" and "toilet" from the extracted text data.

[0208] Output: Detected keywords and their locations.

[0209] Specific operation: The device analyzes the extracted text data and searches for specific keywords.

[0210] Step 4:

[0211] The detected information is notified to the user by voice.

[0212] Input: Keywords detected in step 3 and their location information.

[0213] Data calculation: Convert keywords into speech using Pyttsx3.

[0214] Output: Audio data.

[0215] Specific operation: The terminal uses the Pyttsx3 library to generate voice data such as "There is a STOP sign ahead" and notify the user through earphones or speakers.

[0216] Step 5:

[0217] Record the ambient sounds with a microphone.

[0218] Input: Ambient sounds around the user.

[0219] Data processing: Record the environmental sounds as they are.

[0220] Output: Recorded audio data.

[0221] What it does: The device's microphone continuously records ambient sounds.

[0222] Step 6:

[0223] Analyzes recorded audio data and identifies specific sounds.

[0224] Input: The audio data recorded in step 5.

[0225] Data calculation: Analyze audio data using SpeechRecognition to identify specific sounds (e.g., in-store announcements).

[0226] Output: The identified specific audio data.

[0227] What it does: The device uses the SpeechRecognition library to analyze the recorded audio data and identify specific sounds, such as "We're offering a tasting."

[0228] Step 7:

[0229] Based on the identified specific sound, the device notifies you with vibration and text display.

[0230] Input: The specific audio data identified in step 6.

[0231] Data calculation: Controls vibration and text display based on the identified voice data.

[0232] Output: Vibration and text information displayed on the display.

[0233] Specific operation: When the device identifies a specific sound, it uses the Vibrate Library to notify the user by vibrating, and at the same time displays text information such as "Tasting information available" on the display.

[0234] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0235] This invention is a smartphone application that supports the daily lives of users with visual or hearing impairments. This application includes a means for detecting obstacles and signs from camera images and converting that information into audio, as well as a means for recording environmental sounds, identifying specific sounds, and notifying them by vibration and displaying text. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, the notification method can be appropriately adjusted, further promoting the user's independence and social participation.

[0236] Detects obstacles and signs from camera footage and provides audio notifications

[0237] 1. Capture camera footage:

[0238] The device uses a camera to capture real-time images that show the user's surroundings and are used to detect obstacles and signs.

[0239] 2. Obstacle and sign detection:

[0240] The device analyzes the captured video to detect obstacles and signs. Image processing and optical character recognition (OCR) technologies are used to analyze the video. Specifically, the video is converted to grayscale and text information is extracted.

[0241] 3. Speech synthesis and notifications:

[0242] The device receives information about detected obstacles and signs as text and converts it into speech using speech synthesis technology. This speech is then provided to the user via earphones or speakers, providing them with information on how to act safely.

[0243] Environmental sound recording and important sound notifications

[0244] 1. Environmental Sound Recording:

[0245] The device uses a microphone to record ambient sounds, and the recorded audio data is used to detect certain important sounds.

[0246] 2. Identifying specific sounds:

[0247] The device uses voice recognition technology to analyze recorded audio data and identify specific sounds such as a horn, a bicycle bell, a kettle, a baby crying, or someone calling your name in a public place.

[0248] 3. Vibration and text notification:

[0249] Based on the identified specific sound, the device will vibrate and display the notification content as text on the display, allowing hearing-impaired people to respond appropriately to important sounds.

[0250] Incorporating and applying emotion engines

[0251] 1. Emotion recognition:

[0252] The device uses a camera to capture the user's facial expressions and analyzes them with an emotion engine to recognize the user's current emotional state (e.g., joy, anxiety, anger, etc.).

[0253] 2. Changes in notification methods:

[0254] The device can then adjust the notification content and method appropriately based on the user's emotional state as recognized by the emotion engine. For example, if the user is feeling anxious, the device can soften the tone of the notification.

[0255] Specific examples

[0256] 1. Support for visually impaired people walking on the road:

[0257] As a user walks down the sidewalk, the device uses its camera to capture images of the area ahead and uses OCR to detect "STOP" signs. This information is then converted into "There is a STOP sign ahead" using speech synthesis technology, and the user receives a voice notification through their earphones, providing them with information on how to act safely. If the user's facial expression indicates anxiety, the notification is delivered in a gentler tone.

[0258] 2. Support for the hearing impaired in noisy cafes:

[0259] While a user is waiting to order at a cafe, the device's microphone records ambient sounds and uses voice recognition technology to recognize "your name has been called." As a result, the device notifies the user by vibrating and displays "Your name has been called" on the display so that the user is aware of the call. Furthermore, if the user has a happy expression, the display method and vibration intensity can be adjusted.

[0260] As a result, the present invention can support the daily lives of users with visual or hearing impairments, and promote their independence and social participation.

[0261] The processing flow will be explained below.

[0262] Detects obstacles and signs from camera footage and provides audio notifications

[0263] Step 1:

[0264] The device activates the camera and captures video in real time, specifically by using the camera device interface to obtain video frames.

[0265] Step 2:

[0266] The device pre-processes the captured video frames, which includes converting the video to grayscale for efficient image processing.

[0267] Step 3:

[0268] The device detects obstacles and signs from the grayscale image and uses optical character recognition (OCR) technology to extract text information from the image.

[0269] Step 4:

[0270] The device converts the extracted text information into speech, using speech synthesis technology to generate natural-sounding speech data.

[0271] Step 5:

[0272] The device plays the generated audio data, providing the user with an audio notification through a speaker or earphones.

[0273] Environmental sound recording and important sound notifications

[0274] Step 1:

[0275] The device activates the microphone and records the surrounding environmental sounds at regular intervals to collect audio data that can be used to identify the environmental sounds.

[0276] Step 2:

[0277] The device analyzes the recorded voice data and uses voice recognition technology to extract text information from the recording.

[0278] Step 3:

[0279] The device analyzes the extracted text information to identify important sounds (e.g., horns, bells, etc.) and prepares notification actions based on the identified sounds.

[0280] Step 4:

[0281] The device will vibrate when important sounds are identified, allowing users to receive notifications through physical vibrations.

[0282] Step 5:

[0283] The device will display the notification content as text on the display, allowing the user to visually confirm the notification.

[0284] Incorporating and applying emotion engines

[0285] Step 1:

[0286] The device uses a camera to capture the user's facial expressions, and the emotion engine analyzes the user's emotional state based on the expressions.

[0287] Step 2:

[0288] The device adjusts the content and method of notifications based on the emotional state recognized by the emotion engine. For example, if the user is feeling anxious, the tone of the notification may be softened.

[0289] Step 3:

[0290] The device will reflect the changed notification settings and provide sound, vibration, and text notification, whichever is best for the user.

[0291] Specific examples

[0292] 1. Support for visually impaired people walking on the road:

[0293] When a user is walking down the sidewalk, the device captures images of the area ahead with its camera and uses OCR technology to detect "STOP" signs. This information is then converted into "There is a STOP sign ahead" using speech synthesis technology, which the user receives as a voice notification.

[0294] If the user looks anxious, the notification tone will be gentler, increasing the user's sense of security and ensuring safety.

[0295] 2. Support for the hearing impaired in cafes:

[0296] While the user is waiting for their order at a cafe, the device's microphone records ambient sounds and uses voice recognition technology to recognize that "your name has been called." As a result, the device vibrates and displays the message "Your name has been called" on the display.

[0297] If the user has a happy expression, the notification content will be adjusted appropriately, changing the vibration intensity and the way the display is presented.

[0298] Example 2

[0299] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0300] Users with visual or hearing impairments have difficulty obtaining information to act safely in their daily lives, and difficulty receiving appropriate notifications according to the situation. They also have difficulty distinguishing between external voice instructions and environmental sounds, which makes it difficult to respond appropriately in situations where a quick response is required.

[0301] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for detecting obstacles and signs from camera images, a means for converting information about the detected obstacles and signs into sound and notifying the user, a means for recording environmental sounds and identifying specific sounds, a means for notifying the user of the identified specific sounds by vibration and text display, a means for analyzing the user's facial expression and recognizing emotions, and a means for changing the content and method of notification based on the recognized emotions. This allows users with visual or hearing impairments to act safely and appropriately respond to various situations in daily life.

[0302] "Camera footage" refers to visual information captured by a camera installed on a device such as a smartphone.

[0303] An "obstacle" is an object or structure that impedes the movement or activity of a visually impaired person.

[0304] A "sign" is any public or private display intended to provide important information to persons with visual impairments.

[0305] The "means for converting into voice" is a technology for converting text information into voice signals, and is a system for providing the user with auditory information.

[0306] "Environmental sound" refers to all sound information obtained from the surrounding acoustic environment.

[0307] "Specific sounds" are sounds that are important in general or in specific situations (e.g., horns, bells, warning sounds, etc.).

[0308] "Vibration" is a means of providing sensory feedback to the user by physically shaking the device.

[0309] "Text display" is a means of displaying text information on a terminal display.

[0310] "Facial expressions" are movements and muscle patterns that appear on the user's face and indicate emotions or states.

[0311] "Means for recognizing emotions" refers to technology that analyzes the user's facial expressions to estimate their current mental state and emotions.

[0312] The "means for changing the notification content and method" is a technology for adjusting the information to be notified and the notification method based on the recognized emotion.

[0313] The present invention is a smartphone application that supports the daily lives of visually or hearing impaired users. This application uses the following specific means:

[0314] First, the smartphone camera is used to capture real-time video. The device activates the camera and continuously captures video. This video shows the user's surroundings and is used to detect obstacles and signs.

[0315] The device then captures the video frame by frame and converts it to grayscale using an image processing library (e.g., OpenCV), then uses an image recognition model (e.g., YOLO, SSD) to detect obstacles and signs, and then uses optical character recognition (OCR) technology to extract text information from signs.

[0316] The device then stores the identified obstacles and signs in text format and converts them into audio using a speech synthesis engine (e.g., Google Text-to-Speech), which is then transmitted to the user via earphones or speakers.

[0317] The device also uses the smartphone's built-in microphone to record ambient sounds. The recorded audio data is preprocessed and analyzed using a speech recognition model (e.g., Google Cloud Speech-to-Text) to identify specific sounds such as a horn, a bicycle bell, or a baby crying.

[0318] When a specific sound is identified, the device will activate the vibration motor and display a notification on the display, allowing users to be aware of important sounds without relying on hearing.

[0319] In addition, the device uses the camera to capture the user's face and analyzes their facial expressions using an emotion recognition model (e.g., FaceAPI). This allows the device to recognize the user's emotional state. Based on the recognized emotion, the device can adjust the content and method of notifications. For example, if the user has an anxious expression, the device can respond by softening the tone of the notification.

[0320] As a concrete example, consider a scenario in which a visually impaired person is walking down a street. As the user walks along the sidewalk, the device uses a camera to capture images of the area ahead and uses OCR to detect a "STOP" sign. This information is converted into "There is a STOP sign ahead" using speech synthesis technology, and the user receives a voice notification through their earphones, allowing them to act safely. Furthermore, if the user looks anxious, the notification is delivered in a gentler tone.

[0321] As another example, consider a scenario in which a hearing-impaired person is waiting to order in a noisy cafe. While the user is waiting, the device's microphone records the ambient sounds and uses voice recognition technology to recognize the message "your name has been called." As a result, the device notifies the user by vibrating and displays "Your name has been called" on the display so that the user is aware of the call. Furthermore, if the user has a happy expression, the display method and the intensity of the vibration can be adjusted.

[0322] An example of a prompt is as follows:

[0323] "When a visually impaired person is walking down the street, the system uses a smartphone camera to detect STOP signs and notifies them of this information via audio."

[0324] "While a hearing-impaired person is waiting to order in a noisy cafe, they will be notified by vibration and display when their name is called."

[0325] In this way, the present invention can support the daily lives of users with visual or hearing impairments, promoting their independence and social participation.

[0326] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0327] Step 1:

[0328] Camera footage capture

[0329] The device activates the smartphone camera and continuously captures video in real time.

[0330] Input: Visual information of the user's surroundings.

[0331] Output: Real-time video data.

[0332] Specific operation: The user launches the app and presses the "Launch Camera" button. The device begins capturing video from the camera.

[0333] Step 2:

[0334] Video pre-processing

[0335] The device captures the video frame by frame and converts it to grayscale using an image processing library (e.g., OpenCV).

[0336] Input: Captured real-time video data.

[0337] Output: Video data converted to grayscale.

[0338] Specific operation: The acquired RGB video data is converted to grayscale and filter processing is applied to remove noise.

[0339] Step 3:

[0340] Obstacle and sign detection

[0341] The device inputs the preprocessed video data into an image recognition model (e.g., YOLO, SSD) to detect obstacles and signs.

[0342] Input: Grayscale video data.

[0343] Output: Information of detected obstacles and signs (position and type).

[0344] How it works: Image recognition models analyze specific shapes and text to identify things like "STOP" signs and pedestrians.

[0345] Step 4:

[0346] Extracting text information from signs

[0347] The device extracts text information from parts of the detected signs using optical character recognition (OCR) technology.

[0348] Input: Image data of the sign.

[0349] Output: The extracted text information (e.g. "STOP").

[0350] Specific operation: Using an OCR engine, text information is extracted from the sign image data and saved in text format.

[0351] Step 5:

[0352] Text-to-Speech and Notifications

[0353] The device inputs the text information into a speech synthesis engine (e.g., Google Text-to-Speech) to convert it into speech, and notifies the user of the generated speech through earphones or speakers.

[0354] Input: The extracted text information.

[0355] Output: Audio data.

[0356] Specific operation: The text "There is a STOP sign ahead" is input into the speech synthesis engine, and the generated speech data is played back.

[0357] Step 6:

[0358] Environmental sound recording

[0359] The device uses the smartphone's built-in microphone to record environmental sounds.

[0360] Input: Ambient sound.

[0361] Output: Recorded audio data.

[0362] Specific operation: The user launches the app and presses the "Start Recording" button. The device begins recording ambient sounds.

[0363] Step 7:

[0364] Audio data preprocessing

[0365] The device samples the recorded audio data and applies a noise reduction filter.

[0366] Input: Recorded audio data.

[0367] Output: Preprocessed audio data.

[0368] What it does: Reduces background noise from audio data and normalizes audio clips.

[0369] Step 8:

[0370] Identifying specific sounds

[0371] The device inputs the preprocessed audio data into a speech recognition model (e.g., Google Cloud Speech-to-Text) to identify specific sounds (e.g., a horn, a bicycle bell, a baby crying, etc.).

[0372] Input: Preprocessed audio data.

[0373] Output: Identification result of specific sound.

[0374] Specific operation: Speech waveform data is converted into a spectrogram, and a speech recognition model identifies certain patterns.

[0375] Step 9:

[0376] Vibration and text notifications

[0377] When the device identifies a specific sound, it activates the vibration motor and displays a notification on the display.

[0378] Input: Information about the identified specific sound.

[0379] Output: Vibration and display.

[0380] Specific operation: If the identified sound is a "horn," the device will begin vibrating and the message "Horn has been honked" will be displayed on the screen.

[0381] Step 10:

[0382] emotion recognition

[0383] The device uses a camera to capture video of the user's face and analyzes facial expressions using an emotion recognition model (e.g., FaceAPI).

[0384] Input: Video data of the user's face.

[0385] Output: Perceived emotional state.

[0386] Specific operation: Analyzes the feature points of the user's face and determines whether they correspond to "happiness," "anxiety," or "anger."

[0387] Step 11:

[0388] Changes to notification content

[0389] The device will change the content and presentation of notifications based on the perceived emotion, for example softening the tone of notifications if the user is anxious.

[0390] Input: Perceived emotional state.

[0391] Output: Tailored notification content and method.

[0392] What it does: If anxiety is detected, it will change the settings of the speech synthesis engine to make the tone of notifications gentler.

[0393] (Application example 2)

[0394] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0395] It is necessary to provide a means to improve safety and efficiency for users with visual or hearing impairments, who have difficulty accurately understanding and responding to their surroundings in their daily lives and work environments. In addition, there is a lack of functionality to adjust notification methods to the user's emotional state and to simultaneously analyze multiple sensory information.

[0396] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0397] In this invention, the server includes means for detecting obstacles and signs from camera images, means for converting information about detected obstacles and signs into audio, means for recording environmental sounds and identifying specific sounds, means for notifying the identified specific sounds by vibration and text display, means for recognizing the user's emotions and appropriately changing the notification content based on the user's emotional state, and means for simultaneously capturing camera images and microphone audio and detecting obstacles and specific sounds in real time. This enables users with visual or hearing impairments to act safely and efficiently in real time, improving their awareness and response to their surroundings.

[0398] "Visually impaired persons" refers to people who have visual impairments and have difficulty obtaining information through their eyesight in their daily lives or working environments.

[0399] "Hearing impaired" refers to people who have hearing impairments and have difficulty obtaining information through sound in their daily lives or working environments.

[0400] "Smart devices" refers to portable information terminals with advanced functions such as smartphones, tablets, and smartwatches.

[0401] "Camera footage" refers to video data captured by a camera to obtain visual information.

[0402] "Obstacle" refers to a physical object that impedes a user's movement or activity.

[0403] "Sign" refers to an object on a road or building that has figures or letters on it to provide information or instructions.

[0404] "Speech synthesis" refers to the technology of analyzing text data and converting it into voice data.

[0405] "Environmental sounds" refers to natural and artificial sounds that exist in the surrounding environment.

[0406] "Specific sounds" refer to sounds that need to be specifically recognized, such as horns, bells, and alarms.

[0407] "Vibration" refers to the technology in which a device vibrates to convey information to the user.

[0408] "Character display" refers to the technology of displaying text information on a display.

[0409] An "emotion engine" refers to an algorithm or system that analyzes a user's facial expressions and behavior to recognize their emotional state.

[0410] "Simultaneous capture" refers to the process of capturing camera video and microphone audio at the same time.

[0411] "Real-time" refers to processing and reaction occurring immediately, without delay.

[0412] To implement this invention, it is important to understand the system configuration and its specific operation shown below. The system is composed of a smart device, an industrial camera, a high-sensitivity microphone, a built-in vibration motor, and an LCD display. Various processes are also realized using open source libraries and cloud APIs.

[0413] First, an industrial camera (e.g., industrial camera) is used to capture video in real time. The video data is processed using the OpenCV library to detect obstacles and signs. The TensorFlow library is then used to perform image analysis using a deep learning model. An OCR engine (e.g., Tesseract) is also used to extract text information from signs.

[0414] Next, environmental sounds are recorded using a high-sensitivity microphone (e.g., a high-sensitivity microphone). The recorded audio data is analyzed using the DeepSpeech library to identify certain important sounds (e.g., horns, warning sounds). The identified sounds are notified to the user through the built-in vibration motor and LCD display (e.g., a TFT tactile display).

[0415] Furthermore, to recognize the user's emotional state, the system captures the user's facial expressions using camera footage and performs emotion analysis using the Microsoft Azure Emotion API. Based on the user's emotional state, the system appropriately adjusts the content and method of notifications (audio tone and notification frequency).

[0416] Particularly in a factory environment, it is necessary to simultaneously capture camera images and microphone audio, and detect obstacles and specific sounds in real time. The specific program for this is as follows:

[0417] As a concrete example, consider a scenario in which, when an obstacle is detected, a voice notification is given saying "Warning: Obstacle detected," and when a warning sound is detected in the ambient sound, the user is notified by vibration and a display. This system enables users with visual or hearing impairments to act safely and efficiently in real time.

[0418] Prompt Sentence Examples

[0419] Create a program that simultaneously captures camera images and microphone audio, detects obstacles and specific sounds in the factory in real time, and notifies the user by sound, vibration, and display using technologies such as OpenCV, TensorFlow, DeepSpeech, and Microsoft Azure Emotion API.

[0420] This concludes the "Mode for Carrying Out the Invention." Using this system, visually and hearing impaired users can significantly improve safety and efficiency in their living and working environments.

[0421] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0422] Step 1:

[0423] The terminal uses an industrial camera to capture images in real time.

[0424] Input: Real-time video from inside the factory

[0425] Output: Captured video data

[0426] How it works: The camera constantly captures images of the factory and generates video data, which is then passed on to the next processing step.

[0427] Step 2:

[0428] The device uses the OpenCV library to process the captured images and detect obstacles and signs.

[0429] Input: Video data acquired in step 1

[0430] Output: Location information of obstacles and signs

[0431] Specific operation: The OpenCV library analyzes video data and uses an object recognition algorithm to detect obstacles and signs. The detection results are generated as location information.

[0432] Step 3:

[0433] The device uses the TensorFlow library and an OCR engine (e.g., Tesseract) to parse and extract text information from signs.

[0434] Input: Location of signs detected in step 2

[0435] Output: Sign text information

[0436] Specific operation: The TensorFlow library is used to identify signs in the video, and the OCR engine is used to extract the text information written on the signs.

[0437] Step 4:

[0438] The device uses the Google Cloud Text-to-Speech API to convert the extracted text information into audio.

[0439] Input: Text information of signs extracted in step 3

[0440] Output: Audio data

[0441] Specific operation: Using the Google Cloud Text-to-Speech API, text information is converted into audio data and notified to the user.

[0442] Step 5:

[0443] The device uses a highly sensitive microphone to record ambient sounds.

[0444] Input: Environmental sounds inside the factory

[0445] Output: Recorded audio data

[0446] Specific operation: The microphone constantly records the surrounding environmental sounds and generates audio data.

[0447] Step 6:

[0448] The device uses the DeepSpeech library to analyze the recorded audio data and identify specific sounds.

[0449] Input: Audio data recorded in step 5

[0450] Output: Specific sound identification result

[0451] Specific operation: Performs voice recognition using the DeepSpeech library and identifies specific sounds such as warning sounds and horns.

[0452] Step 7:

[0453] The terminal notifies the user based on the identified specific sound using the built-in vibration motor and LCD display.

[0454] Input: The specific sound results identified in step 6

[0455] Output: Vibration and text notification

[0456] Specific operation: The device vibrates in response to a specific sound and displays a notification on the display to provide information to the user.

[0457] Step 8:

[0458] The device uses a camera to capture the user's facial expressions and analyzes their emotional state using the Microsoft Azure Emotion API.

[0459] Input: Video data of the user's face

[0460] Output: User's emotional state

[0461] Specific operation: The camera captures the user's face, passes the video data to the Emotion API to analyze emotions, and obtains the results.

[0462] Step 9:

[0463] The terminal appropriately changes the notification content and method based on the user's emotional state.

[0464] Input: The user's emotional state obtained in step 8

[0465] Output: Properly adjusted notification content and notification method

[0466] Specific operation: The tone and method of notification will change depending on the analyzed emotional state, notifying in a gentle tone if the user is anxious, and increasing the vibration if the user is impatient.

[0467] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0468] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0469] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0470] [Second embodiment]

[0471] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0472] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0473] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0474] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0475] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0476] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0477] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0478] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0479] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0480] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0481] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0482] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0483] This is a smartphone application that supports the daily lives of users with visual or hearing impairments. This application includes a means for detecting obstacles and signs from camera images and converting that information into audio, and a means for recording environmental sounds, identifying specific sounds, and notifying them with vibrations and text, thereby promoting the user's independence and participation in society.

[0484] Detects obstacles and signs from camera footage and provides audio notifications

[0485] 1. Capture camera footage:

[0486] The device uses a camera to capture real-time images that show the user's surroundings and are used to detect obstacles and signs.

[0487] 2. Obstacle and sign detection:

[0488] The device analyzes the captured video to detect obstacles and signs. Image processing and optical character recognition (OCR) technologies are used to analyze the video. Specifically, the video is converted to grayscale and text information is extracted.

[0489] 3. Speech synthesis and notifications:

[0490] The device receives information about detected obstacles and signs as text and converts it into speech using speech synthesis technology. This speech is then provided to the user via earphones or speakers, providing them with information on how to act safely.

[0491] Environmental sound recording and important sound notifications

[0492] 1. Environmental Sound Recording:

[0493] The device uses a microphone to record ambient sounds, and the recorded audio data is used to detect certain important sounds.

[0494] 2. Identifying specific sounds:

[0495] The device uses voice recognition technology to analyze recorded audio data and identify specific sounds such as a horn, a bicycle bell, a kettle, a baby crying, or someone calling your name in a public place.

[0496] 3. Vibration and text notification:

[0497] Based on the identified specific sound, the device will vibrate and display the notification content as text on the display, allowing hearing-impaired people to respond appropriately to important sounds.

[0498] Specific examples

[0499] 1. Support for visually impaired people walking on the road:

[0500] When a user is walking on the sidewalk, the device uses its camera to capture images of the area ahead and uses OCR to detect "STOP" signs. This information is converted into "There is a STOP sign ahead" using speech synthesis technology, and the user receives a voice notification through their earphones, providing them with information on how to act safely.

[0501] 2. Support for the hearing impaired in noisy cafes:

[0502] While a user is waiting to order at a cafe, the device's microphone records ambient sounds and uses voice recognition technology to recognize that "your name has been called." As a result, the device notifies the user by vibrating and displays "Your name has been called" on the display so that the user is aware of the call.

[0503] As a result, the present invention can support the daily lives of users with visual or hearing impairments, and promote their independence and social participation.

[0504] The processing flow will be explained below.

[0505] Detects obstacles and signs from camera footage and provides audio notifications

[0506] Step 1:

[0507] The device activates the camera and captures video in real time, specifically by using the camera device interface to obtain video frames.

[0508] Step 2:

[0509] The device preprocesses the captured video frames, specifically converting the video to grayscale to improve the efficiency of image processing.

[0510] Step 3:

[0511] The device detects obstacles and signs from the pre-processed video and uses optical character recognition (OCR) technology to extract text information from the video.

[0512] Step 4:

[0513] The device converts the extracted text information into speech, using speech synthesis technology to generate natural-sounding speech data.

[0514] Step 5:

[0515] The device plays the generated audio data, providing the user with an audio notification through a speaker or earphones.

[0516] Environmental sound recording and important sound notifications

[0517] Step 1:

[0518] The device will activate the microphone and record the surrounding environmental sounds at regular intervals.

[0519] Step 2:

[0520] The device analyzes the recorded voice data and uses voice recognition technology to extract text information from the recording.

[0521] Step 3:

[0522] The device analyzes the extracted text information and identifies important sounds (e.g., horns, bicycle bells, etc.) and prepares corresponding actions when a particular sound is detected.

[0523] Step 4:

[0524] The device will vibrate when an important sound is identified, notifying the user.

[0525] Step 5:

[0526] The device displays the notification content as text on the display, allowing the user to visually confirm the notification.

[0527] Example 1

[0528] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0529] Users with visual or hearing impairments have difficulty recognizing obstacles and signs, and distinguishing environmental sounds in their daily lives, which can limit their ability to travel safely and respond to emergencies. While technologies with visual and hearing assistance are necessary for these users to lead independent lives, currently available technologies do not adequately meet these needs. Specifically, improvements are needed in areas such as image recognition accuracy, prompt voice notifications, and real-time recognition of environmental sounds.

[0530] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0531] In this invention, the server includes means for detecting obstacles and signs from camera images, means for converting information about detected obstacles and signs into audio, means for recording environmental sounds and identifying specific sounds, means for performing grayscale conversion and optical character recognition through video analysis, means for converting text information into audio using speech synthesis technology, and means for identifying important sounds using speech recognition technology, thereby enabling users with visual or hearing impairments to travel safely and respond to emergencies.

[0532] "Camera footage" refers to real-time visual data acquired using the device's camera.

[0533] An "obstacle" is an object that exists in the user's direction of travel or around the user, and that obstructs the user's movement or actions.

[0534] A "sign" is a sign placed on a road or in a public place that displays letters or symbols to convey specific instructions or warnings.

[0535] "Means for converting into speech" refers to technology for converting text information into auditory information, and is a device or software that utilizes speech synthesis technology.

[0536] "Environmental sounds" are all sounds that exist around the user, and are audio data collected by a recording device.

[0537] A "specific sound" is a specific important sound that is identified from among environmental sounds, and is a sound that the user needs to recognize immediately.

[0538] "Vibration notification" refers to a technology that allows a device to generate specific vibrations to convey information to the user.

[0539] "Text display" is a means of visually providing information to a user by displaying text information on a terminal display.

[0540] "Grayscale conversion" is an image processing technique that converts camera images into images that contain only black and white shades.

[0541] Optical character recognition (OCR) is a technology that extracts character information from an image and converts it into text data.

[0542] "Speech synthesis technology" is a technology that converts text data into human speech and produces speech.

[0543] "Speech recognition technology" is a technology that analyzes recorded voice data and understands the meaning of language.

[0544] An "important sound" is a sound that should immediately draw the user's attention, such as a car horn, a baby crying, or someone calling your name.

[0545] The present invention relates to a smartphone application for supporting the daily lives of users with visual or hearing impairments. This application combines multiple technical means to help users live safely and independently.

[0546] Hardware and Software Configuration

[0547] 1. The device uses a smartphone as its main platform, which is equipped with a camera, microphone, speaker / earphone, vibration motor, and display.

[0548] 2. The server uses cloud services (e.g., Amazon Web Services, Google Cloud Platform) for data processing and analysis, which enables it to process large amounts of data in real time and provide users with timely information.

[0549] 3. The device includes the following major software components:

[0550] Image processing algorithms (e.g. OpenCV)

[0551] Optical character recognition (OCR) technology (e.g. Tesseract OCR)

[0552] Speech synthesis technology (e.g., Google Text-to-Speech, Amazon Polly)

[0553] Speech recognition technology (e.g., Google Cloud Speech-to-Text, Amazon Transcribe)

[0554] Data processing and calculation

[0555] 1. Capture camera footage:

[0556] The device uses the smartphone's camera to capture real-time images of the surrounding area, which are then stored in the device's internal memory and analyzed immediately.

[0557] 2. Obstacle and sign detection:

[0558] The device uses OpenCV to convert the video to grayscale and apply an edge detection algorithm, then uses Tesseract OCR to analyze the characters on the sign and extract the text information.

[0559] 3. Speech synthesis and notifications:

[0560] The device converts the extracted text information into speech using voice synthesis technologies such as Google Text-to-Speech or Amazon Polly, and the speech is then transmitted to the user through earphones or speakers.

[0561] 4. Recording environmental sounds and identifying specific sounds:

[0562] The device uses the smartphone's microphone to record surrounding sounds, which are then stored in the device's internal memory and analyzed using voice recognition technologies such as Google Cloud Speech-to-Text and Amazon Transcribe.

[0563] 5. Vibration and text notification:

[0564] The device will vibrate based on the identified specific sound and display the notification content as text on the display, allowing the user to immediately recognize important information.

[0565] Specific examples

[0566] 1. For a visually impaired person walking on the road:

[0567] As a user walks along the sidewalk, the device uses its camera to capture images of the area ahead. It uses OpenCV's image processing algorithms to recognize obstacles and "STOP" signs, converting that information into text using Tesseract OCR. It then uses Google Text-to-Speech technology to translate the text into audio, saying "There is a STOP sign ahead," and notifies the user through earphones.

[0568] 2. For a hearing impaired person in a noisy cafe:

[0569] While a user is waiting to order at a cafe, the device's microphone records surrounding sounds. Using Google Cloud Speech-to-Text technology, the device identifies the voice saying "Mr. / Ms. XX, your order is ready," notifying the user with a vibration and displaying "Your name has been called" on the display.

[0570] Prompt Sentence Examples

[0571] By inputting prompt statements such as the following into the generative AI model, an explanation of the system and specific examples can be generated.

[0572] Describe a smartphone application system that supports the daily lives of users with visual or hearing impairments. Explain in detail how it detects obstacles and signs from camera footage and converts that information into audio, and how it records environmental sounds, identifies specific sounds, and notifies the user by vibrating and displaying text. Give specific examples.

[0573] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0574] Step 1: Capture camera footage:

[0575] The device activates the smartphone camera and captures images of the surroundings in real time. The camera captures the scenery ahead as the user walks along the sidewalk. This image data is stored in the internal memory. The input is an image of the user's surroundings, and the output is real-time image data.

[0576] Step 2: Obstacle and sign detection:

[0577] The device uses OpenCV to analyze the captured video. Specifically, it converts the video to grayscale and applies an edge detection algorithm (Canny edge detection). After this, it uses Tesseract OCR to extract the text information of the sign from the video. The input is real-time video data, and the output is text information (e.g., a "STOP" sign).

[0578] Step 3: Text-to-Speech and Notifications:

[0579] The device converts the extracted text information into audio data using the Google Text-to-Speech API. For example, audio data such as "There is a STOP sign ahead" is generated. If the user is using earphones, the audio notification is transmitted through the earphones. The input is text information, and the output is audio data.

[0580] Step 4: Recording ambient sounds:

[0581] The device activates the smartphone's microphone and records the surrounding environmental sounds in real time. For example, it records the surrounding sounds when the user is in a cafe. This audio data is stored in the internal memory. The input is the user's surrounding sounds, and the output is audio data.

[0582] Step 5: Identifying specific sounds:

[0583] The device analyzes the recorded voice data using the Google Cloud Speech-to-Text API to identify specific important sounds (e.g., "Mr. / Ms. XX, your order is ready"). The input is the recorded voice data, and the output is the identified text information.

[0584] Step 6: Vibration and text notification:

[0585] The device vibrates based on the identified specific sound and displays the text "Your name has been called" on the display, thereby alerting the user to the important notification. The input is the identified text information, and the output is a notification via vibration and text display.

[0586] This allows users with visual or hearing impairments to move safely and quickly perceive important information.

[0587] (Application example 1)

[0588] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0589] Users with visual or hearing impairments face challenges in safely moving around in physical stores and receiving the information they need. In such situations, users may bump into obstacles or miss important announcements, limiting their independent movement. As a result, users with visual or hearing impairments face challenges in limiting their opportunities to participate in social life.

[0590] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0591] In this invention, the server includes means for detecting obstacles and signs from camera images, means for converting information about detected obstacles and signs into audio, means for recording environmental sounds and identifying specific sounds, means for detecting obstacles and guide signs in the physical store and notifying them by audio, and means for recording announcements in the physical store and notifying identified announcements by vibration and text display, thereby enabling users with visual or hearing impairments to move around the physical store safely and efficiently.

[0592] A "visually or hearing impaired user" is an individual who is visually or hearing impaired, or both.

[0593] An "information processing device" is a machine or device for inputting, processing, and outputting data.

[0594] "Camera video" refers to real-time video data acquired using a camera.

[0595] An "obstacle" is a physical object that blocks the user's movement.

[0596] A "sign" is a visual display that conveys information.

[0597] "Converting to voice" refers to converting text information into voice data using voice synthesis technology.

[0598] "Ambient sounds" are any sounds occurring around the user.

[0599] "Identifying specific sounds" means recognizing and extracting characteristic sounds from recorded audio data.

[0600] "Vibration" is a means of providing tactile feedback to the user by causing the device to vibrate.

[0601] "Character display" means displaying text information on a display or the like.

[0602] A "physical store" is a store located in a physical location that offers goods and services.

[0603] An "announcement" is an audio message intended to convey information in a public place or specific environment.

[0604] "Recording" means to record audio using a device such as a microphone.

[0605] "Natural language processing" is a computational technique for processing and understanding human language.

[0606] The present invention relates to an information processing device that supports users with visual or hearing impairments in moving around safely and efficiently in a physical store. This device analyzes camera images and environmental sounds and provides the detected information to the user through means such as voice, vibration, and text display.

[0607] Hardware and Software Configuration

[0608] The present invention uses the following hardware and software.

[0609] Hardware:

[0610] Camera: Used to capture video in real time.

[0611] Microphone: Used to record ambient sounds and announcements.

[0612] Smartphone or tablet: The platform for processing and notification.

[0613] Headphones or speakers: Used to provide audio notifications.

[0614] Display: Used to display text information.

[0615] software:

[0616] OpenCV (open source image processing library): Used to analyze camera footage and detect obstacles and signs.

[0617] PyTesseract (OCR library): Used to extract text information from captured video.

[0618] Pyttsx3 (speech synthesis library): Used to convert text information into speech.

[0619] SpeechRecognition (speech recognition library): Used to identify specific sounds and announcements from recorded audio data.

[0620] GTTs (Google Text-to-Speech): The identified voice information is further analyzed and used to generate the required notifications.

[0621] Vibrate Library: Used to vibrate when a specific sound is identified.

[0622] Specific processing steps

[0623] 1. Camera footage capture and analysis:

[0624] The device's camera captures video in real time and converts it to grayscale using OpenCV. PyTesseract is used to extract text information from the image and detect obstacles and signs. For example, signs such as "STOP," "toilet," and "exit" can be detected in a physical store, and speech synthesis technology (Pyttsx3) can be used to notify the user, such as "There is a STOP sign ahead."

[0625] 2. Environmental sound recording and analysis:

[0626] The device's microphone records environmental sounds, and the recorded audio data is analyzed using SpeechRecognition. For example, when an in-store announcement or a specific event (such as a tasting invitation) is detected, the vibration library is used to notify the user by vibrating, such as "Tasting invitation." At the same time, the text information "Tasting invitation" is displayed on the display.

[0627] Adding specific examples

[0628] As an example, the following prompt sentence can be used:

[0629] Example prompt sentence:

[0630] It uses cameras to detect footage as you walk through the store and notifies you of "STOP" signs via voice synthesis.

[0631] When a user is walking through a supermarket, the camera footage is analyzed to detect a "STOP" sign, and an audio notification is played from the smartphone speaker saying, "There is a STOP sign ahead."

[0632] It records ambient sounds within the store, identifies announcements from specific areas (such as the food sampling area), and notifies you with a vibration.

[0633] The smartphone's microphone is used to record in-store announcements, and voice recognition technology is used to identify specific announcements such as "Invitation to sample food" and notify the user by vibrating.

[0634] As a result, the present invention enables users with visual or hearing impairments to move around safely and efficiently within physical stores, promoting their independence and social participation.

[0635] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0636] Step 1:

[0637] Capture footage with the camera.

[0638] Input: Capture real-time video using the device's camera.

[0639] Data processing: This video data will be used as is in the next step.

[0640] Output: The video data sent to the next step.

[0641] Specific operation: The device's camera continuously captures images of the user's surroundings in real time.

[0642] Step 2:

[0643] The video data is converted to grayscale and text information is extracted.

[0644] Input: Video data obtained in step 1.

[0645] Data processing: Convert the video to grayscale using OpenCV and extract text information using PyTesseract.

[0646] Output: The extracted text data.

[0647] Specific operation: The device uses the OpenCV library to convert the video data to grayscale, and then performs OCR processing using PyTesseract.

[0648] Step 3:

[0649] Obstacles and signs are detected from the extracted text information.

[0650] Input: The text data extracted in step 2.

[0651] Data calculation: Detect specific keywords such as "STOP" and "toilet" from the extracted text data.

[0652] Output: Detected keywords and their locations.

[0653] Specific operation: The device analyzes the extracted text data and searches for specific keywords.

[0654] Step 4:

[0655] The detected information is notified to the user by voice.

[0656] Input: Keywords detected in step 3 and their location information.

[0657] Data calculation: Convert keywords into speech using Pyttsx3.

[0658] Output: Audio data.

[0659] Specific operation: The terminal uses the Pyttsx3 library to generate voice data such as "There is a STOP sign ahead" and notify the user through earphones or speakers.

[0660] Step 5:

[0661] Record the ambient sounds with a microphone.

[0662] Input: Ambient sounds around the user.

[0663] Data processing: Record the environmental sounds as they are.

[0664] Output: Recorded audio data.

[0665] What it does: The device's microphone continuously records ambient sounds.

[0666] Step 6:

[0667] Analyzes recorded audio data and identifies specific sounds.

[0668] Input: The audio data recorded in step 5.

[0669] Data calculation: Analyze audio data using SpeechRecognition to identify specific sounds (e.g., in-store announcements).

[0670] Output: The identified specific audio data.

[0671] What it does: The device uses the SpeechRecognition library to analyze the recorded audio data and identify specific sounds, such as "We're offering a tasting."

[0672] Step 7:

[0673] Based on the identified specific sound, the device notifies you with vibration and text display.

[0674] Input: The specific audio data identified in step 6.

[0675] Data calculation: Controls vibration and text display based on the identified voice data.

[0676] Output: Vibration and text information displayed on the display.

[0677] Specific operation: When the device identifies a specific sound, it uses the Vibrate Library to notify the user by vibrating, and at the same time displays text information such as "Tasting information available" on the display.

[0678] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0679] This invention is a smartphone application that supports the daily lives of users with visual or hearing impairments. This application includes a means for detecting obstacles and signs from camera images and converting that information into audio, as well as a means for recording environmental sounds, identifying specific sounds, and notifying them by vibration and displaying text. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, the notification method can be appropriately adjusted, further promoting the user's independence and social participation.

[0680] Detects obstacles and signs from camera footage and provides audio notifications

[0681] 1. Capture camera footage:

[0682] The device uses a camera to capture real-time images that show the user's surroundings and are used to detect obstacles and signs.

[0683] 2. Obstacle and sign detection:

[0684] The device analyzes the captured video to detect obstacles and signs. Image processing and optical character recognition (OCR) technologies are used to analyze the video. Specifically, the video is converted to grayscale and text information is extracted.

[0685] 3. Speech synthesis and notifications:

[0686] The device receives information about detected obstacles and signs as text and converts it into speech using speech synthesis technology. This speech is then provided to the user via earphones or speakers, providing them with information on how to act safely.

[0687] Environmental sound recording and important sound notifications

[0688] 1. Environmental Sound Recording:

[0689] The device uses a microphone to record ambient sounds, and the recorded audio data is used to detect certain important sounds.

[0690] 2. Identifying specific sounds:

[0691] The device uses voice recognition technology to analyze recorded audio data and identify specific sounds such as a horn, a bicycle bell, a kettle, a baby crying, or someone calling your name in a public place.

[0692] 3. Vibration and text notification:

[0693] Based on the identified specific sound, the device will vibrate and display the notification content as text on the display, allowing hearing-impaired people to respond appropriately to important sounds.

[0694] Incorporating and applying emotion engines

[0695] 1. Emotion recognition:

[0696] The device uses a camera to capture the user's facial expressions and analyzes them with an emotion engine to recognize the user's current emotional state (e.g., joy, anxiety, anger, etc.).

[0697] 2. Changes in notification methods:

[0698] The device can then adjust the notification content and method appropriately based on the user's emotional state as recognized by the emotion engine. For example, if the user is feeling anxious, the device can soften the tone of the notification.

[0699] Specific examples

[0700] 1. Support for visually impaired people walking on the road:

[0701] As a user walks down the sidewalk, the device uses its camera to capture images of the area ahead and uses OCR to detect "STOP" signs. This information is then converted into "There is a STOP sign ahead" using speech synthesis technology, and the user receives a voice notification through their earphones, providing them with information on how to act safely. If the user's facial expression indicates anxiety, the notification is delivered in a gentler tone.

[0702] 2. Support for the hearing impaired in noisy cafes:

[0703] While a user is waiting to order at a cafe, the device's microphone records ambient sounds and uses voice recognition technology to recognize "your name has been called." As a result, the device notifies the user by vibrating and displays "Your name has been called" on the display so that the user is aware of the call. Furthermore, if the user has a happy expression, the display method and vibration intensity can be adjusted.

[0704] As a result, the present invention can support the daily lives of users with visual or hearing impairments, and promote their independence and social participation.

[0705] The processing flow will be explained below.

[0706] Detects obstacles and signs from camera footage and provides audio notifications

[0707] Step 1:

[0708] The device activates the camera and captures video in real time, specifically by using the camera device interface to obtain video frames.

[0709] Step 2:

[0710] The device pre-processes the captured video frames, which includes converting the video to grayscale for efficient image processing.

[0711] Step 3:

[0712] The device detects obstacles and signs from the grayscale image and uses optical character recognition (OCR) technology to extract text information from the image.

[0713] Step 4:

[0714] The device converts the extracted text information into speech, using speech synthesis technology to generate natural-sounding speech data.

[0715] Step 5:

[0716] The device plays the generated audio data, providing the user with an audio notification through a speaker or earphones.

[0717] Environmental sound recording and important sound notifications

[0718] Step 1:

[0719] The device activates the microphone and records the surrounding environmental sounds at regular intervals to collect audio data that can be used to identify the environmental sounds.

[0720] Step 2:

[0721] The device analyzes the recorded voice data and uses voice recognition technology to extract text information from the recording.

[0722] Step 3:

[0723] The device analyzes the extracted text information to identify important sounds (e.g., horns, bells, etc.) and prepares notification actions based on the identified sounds.

[0724] Step 4:

[0725] The device will vibrate when important sounds are identified, allowing users to receive notifications through physical vibrations.

[0726] Step 5:

[0727] The device will display the notification content as text on the display, allowing the user to visually confirm the notification.

[0728] Incorporating and applying emotion engines

[0729] Step 1:

[0730] The device uses a camera to capture the user's facial expressions, and the emotion engine analyzes the user's emotional state based on the expressions.

[0731] Step 2:

[0732] The device adjusts the content and method of notifications based on the emotional state recognized by the emotion engine. For example, if the user is feeling anxious, the tone of the notification may be softened.

[0733] Step 3:

[0734] The device will reflect the changed notification settings and provide sound, vibration, and text notification, whichever is best for the user.

[0735] Specific examples

[0736] 1. Support for visually impaired people walking on the road:

[0737] When a user is walking down the sidewalk, the device captures images of the area ahead with its camera and uses OCR technology to detect "STOP" signs. This information is then converted into "There is a STOP sign ahead" using speech synthesis technology, which the user receives as a voice notification.

[0738] If the user looks anxious, the notification tone will be gentler, increasing the user's sense of security and ensuring safety.

[0739] 2. Support for the hearing impaired in cafes:

[0740] While the user is waiting for their order at a cafe, the device's microphone records ambient sounds and uses voice recognition technology to recognize that "your name has been called." As a result, the device vibrates and displays the message "Your name has been called" on the display.

[0741] If the user has a happy expression, the notification content will be adjusted appropriately, changing the vibration intensity and the way the display is presented.

[0742] Example 2

[0743] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0744] Users with visual or hearing impairments have difficulty obtaining information to act safely in their daily lives, and difficulty receiving appropriate notifications according to the situation. They also have difficulty distinguishing between external voice instructions and environmental sounds, which makes it difficult to respond appropriately in situations where a quick response is required.

[0745] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for detecting obstacles and signs from camera images, a means for converting information about the detected obstacles and signs into sound and notifying the user, a means for recording environmental sounds and identifying specific sounds, a means for notifying the user of the identified specific sounds by vibration and text display, a means for analyzing the user's facial expression and recognizing emotions, and a means for changing the content and method of notification based on the recognized emotions. This allows users with visual or hearing impairments to act safely and appropriately respond to various situations in daily life.

[0746] "Camera footage" refers to visual information captured by a camera installed on a device such as a smartphone.

[0747] An "obstacle" is an object or structure that impedes the movement or activity of a visually impaired person.

[0748] A "sign" is any public or private display intended to provide important information to persons with visual impairments.

[0749] The "means for converting into voice" is a technology for converting text information into voice signals, and is a system for providing the user with auditory information.

[0750] "Environmental sound" refers to all sound information obtained from the surrounding acoustic environment.

[0751] "Specific sounds" are sounds that are important in general or in specific situations (e.g., horns, bells, warning sounds, etc.).

[0752] "Vibration" is a means of providing sensory feedback to the user by physically shaking the device.

[0753] "Text display" is a means of displaying text information on a terminal display.

[0754] "Facial expressions" are movements and muscle patterns that appear on the user's face and indicate emotions or states.

[0755] "Means for recognizing emotions" refers to technology that analyzes the user's facial expressions to estimate their current mental state and emotions.

[0756] The "means for changing the notification content and method" is a technology for adjusting the information to be notified and the notification method based on the recognized emotion.

[0757] The present invention is a smartphone application that supports the daily lives of visually or hearing impaired users. This application uses the following specific means:

[0758] First, the smartphone camera is used to capture real-time video. The device activates the camera and continuously captures video. This video shows the user's surroundings and is used to detect obstacles and signs.

[0759] The device then captures the video frame by frame and converts it to grayscale using an image processing library (e.g., OpenCV), then uses an image recognition model (e.g., YOLO, SSD) to detect obstacles and signs, and then uses optical character recognition (OCR) technology to extract text information from signs.

[0760] The device then stores the identified obstacles and signs in text format and converts them into audio using a speech synthesis engine (e.g., Google Text-to-Speech), which is then transmitted to the user via earphones or speakers.

[0761] The device also uses the smartphone's built-in microphone to record ambient sounds. The recorded audio data is preprocessed and analyzed using a speech recognition model (e.g., Google Cloud Speech-to-Text) to identify specific sounds such as a horn, a bicycle bell, or a baby crying.

[0762] When a specific sound is identified, the device will activate the vibration motor and display a notification on the display, allowing users to be aware of important sounds without relying on hearing.

[0763] In addition, the device uses the camera to capture the user's face and analyzes their facial expressions using an emotion recognition model (e.g., FaceAPI). This allows the device to recognize the user's emotional state. Based on the recognized emotion, the device can adjust the content and method of notifications. For example, if the user has an anxious expression, the device can respond by softening the tone of the notification.

[0764] As a concrete example, consider a scenario in which a visually impaired person is walking down a street. As the user walks along the sidewalk, the device uses a camera to capture images of the area ahead and uses OCR to detect a "STOP" sign. This information is converted into "There is a STOP sign ahead" using speech synthesis technology, and the user receives a voice notification through their earphones, allowing them to act safely. Furthermore, if the user looks anxious, the notification is delivered in a gentler tone.

[0765] As another example, consider a scenario in which a hearing-impaired person is waiting to order in a noisy cafe. While the user is waiting, the device's microphone records the ambient sounds and uses voice recognition technology to recognize the message "your name has been called." As a result, the device notifies the user by vibrating and displays "Your name has been called" on the display so that the user is aware of the call. Furthermore, if the user has a happy expression, the display method and the intensity of the vibration can be adjusted.

[0766] An example of a prompt is as follows:

[0767] "When a visually impaired person is walking down the street, the system uses a smartphone camera to detect STOP signs and notifies them of this information via audio."

[0768] "While a hearing-impaired person is waiting to order in a noisy cafe, they will be notified by vibration and display when their name is called."

[0769] In this way, the present invention can support the daily lives of users with visual or hearing impairments, promoting their independence and social participation.

[0770] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0771] Step 1:

[0772] Camera footage capture

[0773] The device activates the smartphone camera and continuously captures video in real time.

[0774] Input: Visual information of the user's surroundings.

[0775] Output: Real-time video data.

[0776] Specific operation: The user launches the app and presses the "Launch Camera" button. The device begins capturing video from the camera.

[0777] Step 2:

[0778] Video pre-processing

[0779] The device captures the video frame by frame and converts it to grayscale using an image processing library (e.g., OpenCV).

[0780] Input: Captured real-time video data.

[0781] Output: Video data converted to grayscale.

[0782] Specific operation: The acquired RGB video data is converted to grayscale and filter processing is applied to remove noise.

[0783] Step 3:

[0784] Obstacle and sign detection

[0785] The device inputs the preprocessed video data into an image recognition model (e.g., YOLO, SSD) to detect obstacles and signs.

[0786] Input: Grayscale video data.

[0787] Output: Information of detected obstacles and signs (position and type).

[0788] How it works: Image recognition models analyze specific shapes and text to identify things like "STOP" signs and pedestrians.

[0789] Step 4:

[0790] Extracting text information from signs

[0791] The device extracts text information from parts of the detected signs using optical character recognition (OCR) technology.

[0792] Input: Image data of the sign.

[0793] Output: The extracted text information (e.g. "STOP").

[0794] Specific operation: Using an OCR engine, text information is extracted from the sign image data and saved in text format.

[0795] Step 5:

[0796] Text-to-Speech and Notifications

[0797] The device inputs the text information into a speech synthesis engine (e.g., Google Text-to-Speech) to convert it into speech, and notifies the user of the generated speech through earphones or speakers.

[0798] Input: The extracted text information.

[0799] Output: Audio data.

[0800] Specific operation: The text "There is a STOP sign ahead" is input into the speech synthesis engine, and the generated speech data is played back.

[0801] Step 6:

[0802] Environmental sound recording

[0803] The device uses the smartphone's built-in microphone to record environmental sounds.

[0804] Input: Ambient sound.

[0805] Output: Recorded audio data.

[0806] Specific operation: The user launches the app and presses the "Start Recording" button. The device begins recording ambient sounds.

[0807] Step 7:

[0808] Audio data preprocessing

[0809] The device samples the recorded audio data and applies a noise reduction filter.

[0810] Input: Recorded audio data.

[0811] Output: Preprocessed audio data.

[0812] What it does: Reduces background noise from audio data and normalizes audio clips.

[0813] Step 8:

[0814] Identifying specific sounds

[0815] The device inputs the preprocessed audio data into a speech recognition model (e.g., Google Cloud Speech-to-Text) to identify specific sounds (e.g., a horn, a bicycle bell, a baby crying, etc.).

[0816] Input: Preprocessed audio data.

[0817] Output: Identification result of specific sound.

[0818] Specific operation: Speech waveform data is converted into a spectrogram, and a speech recognition model identifies certain patterns.

[0819] Step 9:

[0820] Vibration and text notifications

[0821] When the device identifies a specific sound, it activates the vibration motor and displays a notification on the display.

[0822] Input: Information about the identified specific sound.

[0823] Output: Vibration and display.

[0824] Specific operation: If the identified sound is a "horn," the device will begin vibrating and the message "Horn has been honked" will be displayed on the screen.

[0825] Step 10:

[0826] emotion recognition

[0827] The device uses a camera to capture video of the user's face and analyzes facial expressions using an emotion recognition model (e.g., FaceAPI).

[0828] Input: Video data of the user's face.

[0829] Output: Perceived emotional state.

[0830] Specific operation: Analyzes the feature points of the user's face and determines whether they correspond to "happiness," "anxiety," or "anger."

[0831] Step 11:

[0832] Changes to notification content

[0833] The device will change the content and presentation of notifications based on the perceived emotion, for example softening the tone of notifications if the user is anxious.

[0834] Input: Perceived emotional state.

[0835] Output: Tailored notification content and method.

[0836] What it does: If anxiety is detected, it will change the settings of the speech synthesis engine to make the tone of notifications gentler.

[0837] (Application example 2)

[0838] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0839] It is necessary to provide a means to improve safety and efficiency for users with visual or hearing impairments, who have difficulty accurately understanding and responding to their surroundings in their daily lives and work environments. In addition, there is a lack of functionality to adjust notification methods to the user's emotional state and to simultaneously analyze multiple sensory information.

[0840] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0841] In this invention, the server includes means for detecting obstacles and signs from camera images, means for converting information about detected obstacles and signs into audio, means for recording environmental sounds and identifying specific sounds, means for notifying the identified specific sounds by vibration and text display, means for recognizing the user's emotions and appropriately changing the notification content based on the user's emotional state, and means for simultaneously capturing camera images and microphone audio and detecting obstacles and specific sounds in real time. This enables users with visual or hearing impairments to act safely and efficiently in real time, improving their awareness and response to their surroundings.

[0842] "Visually impaired persons" refers to people who have visual impairments and have difficulty obtaining information through their eyesight in their daily lives or working environments.

[0843] "Hearing impaired" refers to people who have hearing impairments and have difficulty obtaining information through sound in their daily lives or working environments.

[0844] "Smart devices" refers to portable information terminals with advanced functions such as smartphones, tablets, and smartwatches.

[0845] "Camera footage" refers to video data captured by a camera to obtain visual information.

[0846] "Obstacle" refers to a physical object that impedes a user's movement or activity.

[0847] "Sign" refers to an object on a road or building that has figures or letters on it to provide information or instructions.

[0848] "Speech synthesis" refers to the technology of analyzing text data and converting it into voice data.

[0849] "Environmental sounds" refers to natural and artificial sounds that exist in the surrounding environment.

[0850] "Specific sounds" refer to sounds that need to be specifically recognized, such as horns, bells, and alarms.

[0851] "Vibration" refers to the technology in which a device vibrates to convey information to the user.

[0852] "Character display" refers to the technology of displaying text information on a display.

[0853] An "emotion engine" refers to an algorithm or system that analyzes a user's facial expressions and behavior to recognize their emotional state.

[0854] "Simultaneous capture" refers to the process of capturing camera video and microphone audio at the same time.

[0855] "Real-time" refers to processing and reaction occurring immediately, without delay.

[0856] To implement this invention, it is important to understand the system configuration and its specific operation shown below. The system is composed of a smart device, an industrial camera, a high-sensitivity microphone, a built-in vibration motor, and an LCD display. Various processes are also realized using open source libraries and cloud APIs.

[0857] First, an industrial camera (e.g., industrial camera) is used to capture video in real time. The video data is processed using the OpenCV library to detect obstacles and signs. The TensorFlow library is then used to perform image analysis using a deep learning model. An OCR engine (e.g., Tesseract) is also used to extract text information from signs.

[0858] Next, environmental sounds are recorded using a high-sensitivity microphone (e.g., a high-sensitivity microphone). The recorded audio data is analyzed using the DeepSpeech library to identify certain important sounds (e.g., horns, warning sounds). The identified sounds are notified to the user through the built-in vibration motor and LCD display (e.g., a TFT tactile display).

[0859] Furthermore, to recognize the user's emotional state, the system captures the user's facial expressions using camera footage and performs emotion analysis using the Microsoft Azure Emotion API. Based on the user's emotional state, the system appropriately adjusts the content and method of notifications (audio tone and notification frequency).

[0860] Particularly in a factory environment, it is necessary to simultaneously capture camera images and microphone audio, and detect obstacles and specific sounds in real time. The specific program for this is as follows:

[0861] As a concrete example, consider a scenario in which, when an obstacle is detected, a voice notification is given saying "Warning: Obstacle detected," and when a warning sound is detected in the ambient sound, the user is notified by vibration and a display. This system enables users with visual or hearing impairments to act safely and efficiently in real time.

[0862] Prompt Sentence Examples

[0863] Create a program that simultaneously captures camera images and microphone audio, detects obstacles and specific sounds in the factory in real time, and notifies the user by sound, vibration, and display using technologies such as OpenCV, TensorFlow, DeepSpeech, and Microsoft Azure Emotion API.

[0864] This concludes the "Mode for Carrying Out the Invention." Using this system, visually and hearing impaired users can significantly improve safety and efficiency in their living and working environments.

[0865] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0866] Step 1:

[0867] The terminal uses an industrial camera to capture images in real time.

[0868] Input: Real-time video from inside the factory

[0869] Output: Captured video data

[0870] How it works: The camera constantly captures images of the factory and generates video data, which is then passed on to the next processing step.

[0871] Step 2:

[0872] The device uses the OpenCV library to process the captured images and detect obstacles and signs.

[0873] Input: Video data acquired in step 1

[0874] Output: Location information of obstacles and signs

[0875] Specific operation: The OpenCV library analyzes video data and uses an object recognition algorithm to detect obstacles and signs. The detection results are generated as location information.

[0876] Step 3:

[0877] The device uses the TensorFlow library and an OCR engine (e.g., Tesseract) to parse and extract text information from signs.

[0878] Input: Location of signs detected in step 2

[0879] Output: Sign text information

[0880] Specific operation: The TensorFlow library is used to identify signs in the video, and the OCR engine is used to extract the text information written on the signs.

[0881] Step 4:

[0882] The device uses the Google Cloud Text-to-Speech API to convert the extracted text information into audio.

[0883] Input: Text information of signs extracted in step 3

[0884] Output: Audio data

[0885] Specific operation: Using the Google Cloud Text-to-Speech API, text information is converted into audio data and notified to the user.

[0886] Step 5:

[0887] The device uses a highly sensitive microphone to record ambient sounds.

[0888] Input: Environmental sounds inside the factory

[0889] Output: Recorded audio data

[0890] Specific operation: The microphone constantly records the surrounding environmental sounds and generates audio data.

[0891] Step 6:

[0892] The device uses the DeepSpeech library to analyze the recorded audio data and identify specific sounds.

[0893] Input: Audio data recorded in step 5

[0894] Output: Specific sound identification result

[0895] Specific operation: Performs voice recognition using the DeepSpeech library and identifies specific sounds such as warning sounds and horns.

[0896] Step 7:

[0897] The terminal notifies the user based on the identified specific sound using the built-in vibration motor and LCD display.

[0898] Input: The specific sound results identified in step 6

[0899] Output: Vibration and text notification

[0900] Specific operation: The device vibrates in response to a specific sound and displays a notification on the display to provide information to the user.

[0901] Step 8:

[0902] The device uses a camera to capture the user's facial expressions and analyzes their emotional state using the Microsoft Azure Emotion API.

[0903] Input: Video data of the user's face

[0904] Output: User's emotional state

[0905] Specific operation: The camera captures the user's face, passes the video data to the Emotion API to analyze emotions, and obtains the results.

[0906] Step 9:

[0907] The terminal appropriately changes the notification content and method based on the user's emotional state.

[0908] Input: The user's emotional state obtained in step 8

[0909] Output: Properly adjusted notification content and notification method

[0910] Specific operation: The tone and method of notification will change depending on the analyzed emotional state, notifying in a gentle tone if the user is anxious, and increasing the vibration if the user is impatient.

[0911] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0912] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0913] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0914] [Third embodiment]

[0915] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0916] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0917] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0918] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0919] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0920] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0921] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0922] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0923] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0924] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0925] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0926] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0927] This is a smartphone application that supports the daily lives of users with visual or hearing impairments. This application includes a means for detecting obstacles and signs from camera images and converting that information into audio, and a means for recording environmental sounds, identifying specific sounds, and notifying them with vibrations and text, thereby promoting the user's independence and participation in society.

[0928] Detects obstacles and signs from camera footage and provides audio notifications

[0929] 1. Capture camera footage:

[0930] The device uses a camera to capture real-time images that show the user's surroundings and are used to detect obstacles and signs.

[0931] 2. Obstacle and sign detection:

[0932] The device analyzes the captured video to detect obstacles and signs. Image processing and optical character recognition (OCR) technologies are used to analyze the video. Specifically, the video is converted to grayscale and text information is extracted.

[0933] 3. Speech synthesis and notifications:

[0934] The device receives information about detected obstacles and signs as text and converts it into speech using speech synthesis technology. This speech is then provided to the user via earphones or speakers, providing them with information on how to act safely.

[0935] Environmental sound recording and important sound notifications

[0936] 1. Environmental Sound Recording:

[0937] The device uses a microphone to record ambient sounds, and the recorded audio data is used to detect certain important sounds.

[0938] 2. Identifying specific sounds:

[0939] The device uses voice recognition technology to analyze recorded audio data and identify specific sounds such as a horn, a bicycle bell, a kettle, a baby crying, or someone calling your name in a public place.

[0940] 3. Vibration and text notification:

[0941] Based on the identified specific sound, the device will vibrate and display the notification content as text on the display, allowing hearing-impaired people to respond appropriately to important sounds.

[0942] Specific examples

[0943] 1. Support for visually impaired people walking on the road:

[0944] When a user is walking on the sidewalk, the device uses its camera to capture images of the area ahead and uses OCR to detect "STOP" signs. This information is converted into "There is a STOP sign ahead" using speech synthesis technology, and the user receives a voice notification through their earphones, providing them with information on how to act safely.

[0945] 2. Support for the hearing impaired in noisy cafes:

[0946] While a user is waiting to order at a cafe, the device's microphone records ambient sounds and uses voice recognition technology to recognize that "your name has been called." As a result, the device notifies the user by vibrating and displays "Your name has been called" on the display so that the user is aware of the call.

[0947] As a result, the present invention can support the daily lives of users with visual or hearing impairments, and promote their independence and social participation.

[0948] The processing flow will be explained below.

[0949] Detects obstacles and signs from camera footage and provides audio notifications

[0950] Step 1:

[0951] The device activates the camera and captures video in real time, specifically by using the camera device interface to obtain video frames.

[0952] Step 2:

[0953] The device preprocesses the captured video frames, specifically converting the video to grayscale to improve the efficiency of image processing.

[0954] Step 3:

[0955] The device detects obstacles and signs from the pre-processed video and uses optical character recognition (OCR) technology to extract text information from the video.

[0956] Step 4:

[0957] The device converts the extracted text information into speech, using speech synthesis technology to generate natural-sounding speech data.

[0958] Step 5:

[0959] The device plays the generated audio data, providing the user with an audio notification through a speaker or earphones.

[0960] Environmental sound recording and important sound notifications

[0961] Step 1:

[0962] The device will activate the microphone and record the surrounding environmental sounds at regular intervals.

[0963] Step 2:

[0964] The device analyzes the recorded voice data and uses voice recognition technology to extract text information from the recording.

[0965] Step 3:

[0966] The device analyzes the extracted text information and identifies important sounds (e.g., horns, bicycle bells, etc.) and prepares corresponding actions when a particular sound is detected.

[0967] Step 4:

[0968] The device will vibrate when an important sound is identified, notifying the user.

[0969] Step 5:

[0970] The device displays the notification content as text on the display, allowing the user to visually confirm the notification.

[0971] Example 1

[0972] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0973] Users with visual or hearing impairments have difficulty recognizing obstacles and signs, and distinguishing environmental sounds in their daily lives, which can limit their ability to travel safely and respond to emergencies. While technologies with visual and hearing assistance are necessary for these users to lead independent lives, currently available technologies do not adequately meet these needs. Specifically, improvements are needed in areas such as image recognition accuracy, prompt voice notifications, and real-time recognition of environmental sounds.

[0974] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0975] In this invention, the server includes means for detecting obstacles and signs from camera images, means for converting information about detected obstacles and signs into audio, means for recording environmental sounds and identifying specific sounds, means for performing grayscale conversion and optical character recognition through video analysis, means for converting text information into audio using speech synthesis technology, and means for identifying important sounds using speech recognition technology, thereby enabling users with visual or hearing impairments to travel safely and respond to emergencies.

[0976] "Camera footage" refers to real-time visual data acquired using the device's camera.

[0977] An "obstacle" is an object that exists in the user's direction of travel or around the user, and that obstructs the user's movement or actions.

[0978] A "sign" is a sign placed on a road or in a public place that displays letters or symbols to convey specific instructions or warnings.

[0979] "Means for converting into speech" refers to technology for converting text information into auditory information, and is a device or software that utilizes speech synthesis technology.

[0980] "Environmental sounds" are all sounds that exist around the user, and are audio data collected by a recording device.

[0981] A "specific sound" is a specific important sound that is identified from among environmental sounds, and is a sound that the user needs to recognize immediately.

[0982] "Vibration notification" refers to a technology that allows a device to generate specific vibrations to convey information to the user.

[0983] "Text display" is a means of visually providing information to a user by displaying text information on a terminal display.

[0984] "Grayscale conversion" is an image processing technique that converts camera images into images that contain only black and white shades.

[0985] Optical character recognition (OCR) is a technology that extracts character information from an image and converts it into text data.

[0986] "Speech synthesis technology" is a technology that converts text data into human speech and produces speech.

[0987] "Speech recognition technology" is a technology that analyzes recorded voice data and understands the meaning of language.

[0988] An "important sound" is a sound that should immediately draw the user's attention, such as a car horn, a baby crying, or someone calling your name.

[0989] The present invention relates to a smartphone application for supporting the daily lives of users with visual or hearing impairments. This application combines multiple technical means to help users live safely and independently.

[0990] Hardware and Software Configuration

[0991] 1. The device uses a smartphone as its main platform, which is equipped with a camera, microphone, speaker / earphone, vibration motor, and display.

[0992] 2. The server uses cloud services (e.g., Amazon Web Services, Google Cloud Platform) for data processing and analysis, which enables it to process large amounts of data in real time and provide users with timely information.

[0993] 3. The device includes the following major software components:

[0994] Image processing algorithms (e.g. OpenCV)

[0995] Optical character recognition (OCR) technology (e.g. Tesseract OCR)

[0996] Speech synthesis technology (e.g., Google Text-to-Speech, Amazon Polly)

[0997] Speech recognition technology (e.g., Google Cloud Speech-to-Text, Amazon Transcribe)

[0998] Data processing and calculation

[0999] 1. Capture camera footage:

[1000] The device uses the smartphone's camera to capture real-time images of the surrounding area, which are then stored in the device's internal memory and analyzed immediately.

[1001] 2. Obstacle and sign detection:

[1002] The device uses OpenCV to convert the video to grayscale and apply an edge detection algorithm, then uses Tesseract OCR to analyze the characters on the sign and extract the text information.

[1003] 3. Speech synthesis and notifications:

[1004] The device converts the extracted text information into speech using voice synthesis technologies such as Google Text-to-Speech or Amazon Polly, and the speech is then transmitted to the user through earphones or speakers.

[1005] 4. Recording environmental sounds and identifying specific sounds:

[1006] The device uses the smartphone's microphone to record surrounding sounds, which are then stored in the device's internal memory and analyzed using voice recognition technologies such as Google Cloud Speech-to-Text and Amazon Transcribe.

[1007] 5. Vibration and text notification:

[1008] The device will vibrate based on the identified specific sound and display the notification content as text on the display, allowing the user to immediately recognize important information.

[1009] Specific examples

[1010] 1. For a visually impaired person walking on the road:

[1011] As a user walks along the sidewalk, the device uses its camera to capture images of the area ahead. It uses OpenCV's image processing algorithms to recognize obstacles and "STOP" signs, converting that information into text using Tesseract OCR. It then uses Google Text-to-Speech technology to translate the text into audio, saying "There is a STOP sign ahead," and notifies the user through earphones.

[1012] 2. For a hearing impaired person in a noisy cafe:

[1013] While a user is waiting to order at a cafe, the device's microphone records surrounding sounds. Using Google Cloud Speech-to-Text technology, the device identifies the voice saying "Mr. / Ms. XX, your order is ready," notifying the user with a vibration and displaying "Your name has been called" on the display.

[1014] Prompt Sentence Examples

[1015] By inputting prompt statements such as the following into the generative AI model, an explanation of the system and specific examples can be generated.

[1016] Describe a smartphone application system that supports the daily lives of users with visual or hearing impairments. Explain in detail how it detects obstacles and signs from camera footage and converts that information into audio, and how it records environmental sounds, identifies specific sounds, and notifies the user by vibrating and displaying text. Give specific examples.

[1017] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1018] Step 1: Capture camera footage:

[1019] The device activates the smartphone camera and captures images of the surroundings in real time. The camera captures the scenery ahead as the user walks along the sidewalk. This image data is stored in the internal memory. The input is an image of the user's surroundings, and the output is real-time image data.

[1020] Step 2: Obstacle and sign detection:

[1021] The device uses OpenCV to analyze the captured video. Specifically, it converts the video to grayscale and applies an edge detection algorithm (Canny edge detection). After this, it uses Tesseract OCR to extract the text information of the sign from the video. The input is real-time video data, and the output is text information (e.g., a "STOP" sign).

[1022] Step 3: Text-to-Speech and Notifications:

[1023] The device converts the extracted text information into audio data using the Google Text-to-Speech API. For example, audio data such as "There is a STOP sign ahead" is generated. If the user is using earphones, the audio notification is transmitted through the earphones. The input is text information, and the output is audio data.

[1024] Step 4: Recording ambient sounds:

[1025] The device activates the smartphone's microphone and records the surrounding environmental sounds in real time. For example, it records the surrounding sounds when the user is in a cafe. This audio data is stored in the internal memory. The input is the user's surrounding sounds, and the output is audio data.

[1026] Step 5: Identifying specific sounds:

[1027] The device analyzes the recorded voice data using the Google Cloud Speech-to-Text API to identify specific important sounds (e.g., "Mr. / Ms. XX, your order is ready"). The input is the recorded voice data, and the output is the identified text information.

[1028] Step 6: Vibration and text notification:

[1029] The device vibrates based on the identified specific sound and displays the text "Your name has been called" on the display, thereby alerting the user to the important notification. The input is the identified text information, and the output is a notification via vibration and text display.

[1030] This allows users with visual or hearing impairments to move safely and quickly perceive important information.

[1031] (Application example 1)

[1032] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1033] Users with visual or hearing impairments face challenges in safely moving around in physical stores and receiving the information they need. In such situations, users may bump into obstacles or miss important announcements, limiting their independent movement. As a result, users with visual or hearing impairments face challenges in limiting their opportunities to participate in social life.

[1034] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1035] In this invention, the server includes means for detecting obstacles and signs from camera images, means for converting information about detected obstacles and signs into audio, means for recording environmental sounds and identifying specific sounds, means for detecting obstacles and guide signs in the physical store and notifying them by audio, and means for recording announcements in the physical store and notifying identified announcements by vibration and text display, thereby enabling users with visual or hearing impairments to move around the physical store safely and efficiently.

[1036] A "visually or hearing impaired user" is an individual who is visually or hearing impaired, or both.

[1037] An "information processing device" is a machine or device for inputting, processing, and outputting data.

[1038] "Camera video" refers to real-time video data acquired using a camera.

[1039] An "obstacle" is a physical object that blocks the user's movement.

[1040] A "sign" is a visual display that conveys information.

[1041] "Converting to voice" refers to converting text information into voice data using voice synthesis technology.

[1042] "Ambient sounds" are any sounds occurring around the user.

[1043] "Identifying specific sounds" means recognizing and extracting characteristic sounds from recorded audio data.

[1044] "Vibration" is a means of providing tactile feedback to the user by causing the device to vibrate.

[1045] "Character display" means displaying text information on a display or the like.

[1046] A "physical store" is a store located in a physical location that offers goods and services.

[1047] An "announcement" is an audio message intended to convey information in a public place or specific environment.

[1048] "Recording" means to record audio using a device such as a microphone.

[1049] "Natural language processing" is a computational technique for processing and understanding human language.

[1050] The present invention relates to an information processing device that supports users with visual or hearing impairments in moving around safely and efficiently in a physical store. This device analyzes camera images and environmental sounds and provides the detected information to the user through means such as voice, vibration, and text display.

[1051] Hardware and Software Configuration

[1052] The present invention uses the following hardware and software.

[1053] Hardware:

[1054] Camera: Used to capture video in real time.

[1055] Microphone: Used to record ambient sounds and announcements.

[1056] Smartphone or tablet: The platform for processing and notification.

[1057] Headphones or speakers: Used to provide audio notifications.

[1058] Display: Used to display text information.

[1059] software:

[1060] OpenCV (open source image processing library): Used to analyze camera footage and detect obstacles and signs.

[1061] PyTesseract (OCR library): Used to extract text information from captured video.

[1062] Pyttsx3 (speech synthesis library): Used to convert text information into speech.

[1063] SpeechRecognition (speech recognition library): Used to identify specific sounds and announcements from recorded audio data.

[1064] GTTs (Google Text-to-Speech): The identified voice information is further analyzed and used to generate the required notifications.

[1065] Vibrate Library: Used to vibrate when a specific sound is identified.

[1066] Specific processing steps

[1067] 1. Camera footage capture and analysis:

[1068] The device's camera captures video in real time and converts it to grayscale using OpenCV. PyTesseract is used to extract text information from the image and detect obstacles and signs. For example, signs such as "STOP," "toilet," and "exit" can be detected in a physical store, and speech synthesis technology (Pyttsx3) can be used to notify the user, such as "There is a STOP sign ahead."

[1069] 2. Environmental sound recording and analysis:

[1070] The device's microphone records environmental sounds, and the recorded audio data is analyzed using SpeechRecognition. For example, when an in-store announcement or a specific event (such as a tasting invitation) is detected, the vibration library is used to notify the user by vibrating, such as "Tasting invitation." At the same time, the text information "Tasting invitation" is displayed on the display.

[1071] Adding specific examples

[1072] As an example, the following prompt sentence can be used:

[1073] Example prompt sentence:

[1074] It uses cameras to detect footage as you walk through the store and notifies you of "STOP" signs via voice synthesis.

[1075] When a user is walking through a supermarket, the camera footage is analyzed to detect a "STOP" sign, and an audio notification is played from the smartphone speaker saying, "There is a STOP sign ahead."

[1076] It records ambient sounds within the store, identifies announcements from specific areas (such as the food sampling area), and notifies you with a vibration.

[1077] The smartphone's microphone is used to record in-store announcements, and voice recognition technology is used to identify specific announcements such as "Invitation to sample food" and notify the user by vibrating.

[1078] As a result, the present invention enables users with visual or hearing impairments to move around safely and efficiently within physical stores, promoting their independence and social participation.

[1079] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1080] Step 1:

[1081] Capture footage with the camera.

[1082] Input: Capture real-time video using the device's camera.

[1083] Data processing: This video data will be used as is in the next step.

[1084] Output: The video data sent to the next step.

[1085] Specific operation: The device's camera continuously captures images of the user's surroundings in real time.

[1086] Step 2:

[1087] The video data is converted to grayscale and text information is extracted.

[1088] Input: Video data obtained in step 1.

[1089] Data processing: Convert the video to grayscale using OpenCV and extract text information using PyTesseract.

[1090] Output: The extracted text data.

[1091] Specific operation: The device uses the OpenCV library to convert the video data to grayscale, and then performs OCR processing using PyTesseract.

[1092] Step 3:

[1093] Obstacles and signs are detected from the extracted text information.

[1094] Input: The text data extracted in step 2.

[1095] Data calculation: Detect specific keywords such as "STOP" and "toilet" from the extracted text data.

[1096] Output: Detected keywords and their locations.

[1097] Specific operation: The device analyzes the extracted text data and searches for specific keywords.

[1098] Step 4:

[1099] The detected information is notified to the user by voice.

[1100] Input: Keywords detected in step 3 and their location information.

[1101] Data calculation: Convert keywords into speech using Pyttsx3.

[1102] Output: Audio data.

[1103] Specific operation: The terminal uses the Pyttsx3 library to generate voice data such as "There is a STOP sign ahead" and notify the user through earphones or speakers.

[1104] Step 5:

[1105] Record the ambient sounds with a microphone.

[1106] Input: Ambient sounds around the user.

[1107] Data processing: Record the environmental sounds as they are.

[1108] Output: Recorded audio data.

[1109] What it does: The device's microphone continuously records ambient sounds.

[1110] Step 6:

[1111] Analyzes recorded audio data and identifies specific sounds.

[1112] Input: The audio data recorded in step 5.

[1113] Data calculation: Analyze audio data using SpeechRecognition to identify specific sounds (e.g., in-store announcements).

[1114] Output: The identified specific audio data.

[1115] What it does: The device uses the SpeechRecognition library to analyze the recorded audio data and identify specific sounds, such as "We're offering a tasting."

[1116] Step 7:

[1117] Based on the identified specific sound, the device notifies you with vibration and text display.

[1118] Input: The specific audio data identified in step 6.

[1119] Data calculation: Controls vibration and text display based on the identified voice data.

[1120] Output: Vibration and text information displayed on the display.

[1121] Specific operation: When the device identifies a specific sound, it uses the Vibrate Library to notify the user by vibrating, and at the same time displays text information such as "Tasting information available" on the display.

[1122] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1123] This invention is a smartphone application that supports the daily lives of users with visual or hearing impairments. This application includes a means for detecting obstacles and signs from camera images and converting that information into audio, as well as a means for recording environmental sounds, identifying specific sounds, and notifying them by vibration and displaying text. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, the notification method can be appropriately adjusted, further promoting the user's independence and social participation.

[1124] Detects obstacles and signs from camera footage and provides audio notifications

[1125] 1. Capture camera footage:

[1126] The device uses a camera to capture real-time images that show the user's surroundings and are used to detect obstacles and signs.

[1127] 2. Obstacle and sign detection:

[1128] The device analyzes the captured video to detect obstacles and signs. Image processing and optical character recognition (OCR) technologies are used to analyze the video. Specifically, the video is converted to grayscale and text information is extracted.

[1129] 3. Speech synthesis and notifications:

[1130] The device receives information about detected obstacles and signs as text and converts it into speech using speech synthesis technology. This speech is then provided to the user via earphones or speakers, providing them with information on how to act safely.

[1131] Environmental sound recording and important sound notifications

[1132] 1. Environmental Sound Recording:

[1133] The device uses a microphone to record ambient sounds, and the recorded audio data is used to detect certain important sounds.

[1134] 2. Identifying specific sounds:

[1135] The device uses voice recognition technology to analyze recorded audio data and identify specific sounds such as a horn, a bicycle bell, a kettle, a baby crying, or someone calling your name in a public place.

[1136] 3. Vibration and text notification:

[1137] Based on the identified specific sound, the device will vibrate and display the notification content as text on the display, allowing hearing-impaired people to respond appropriately to important sounds.

[1138] Incorporating and applying emotion engines

[1139] 1. Emotion recognition:

[1140] The device uses a camera to capture the user's facial expressions and analyzes them with an emotion engine to recognize the user's current emotional state (e.g., joy, anxiety, anger, etc.).

[1141] 2. Changes in notification methods:

[1142] The device can then adjust the notification content and method appropriately based on the user's emotional state as recognized by the emotion engine. For example, if the user is feeling anxious, the device can soften the tone of the notification.

[1143] Specific examples

[1144] 1. Support for visually impaired people walking on the road:

[1145] As a user walks down the sidewalk, the device uses its camera to capture images of the area ahead and uses OCR to detect "STOP" signs. This information is then converted into "There is a STOP sign ahead" using speech synthesis technology, and the user receives a voice notification through their earphones, providing them with information on how to act safely. If the user's facial expression indicates anxiety, the notification is delivered in a gentler tone.

[1146] 2. Support for the hearing impaired in noisy cafes:

[1147] While a user is waiting to order at a cafe, the device's microphone records ambient sounds and uses voice recognition technology to recognize "your name has been called." As a result, the device notifies the user by vibrating and displays "Your name has been called" on the display so that the user is aware of the call. Furthermore, if the user has a happy expression, the display method and vibration intensity can be adjusted.

[1148] As a result, the present invention can support the daily lives of users with visual or hearing impairments, and promote their independence and social participation.

[1149] The processing flow will be explained below.

[1150] Detects obstacles and signs from camera footage and provides audio notifications

[1151] Step 1:

[1152] The device activates the camera and captures video in real time, specifically by using the camera device interface to obtain video frames.

[1153] Step 2:

[1154] The device pre-processes the captured video frames, which includes converting the video to grayscale for efficient image processing.

[1155] Step 3:

[1156] The device detects obstacles and signs from the grayscale image and uses optical character recognition (OCR) technology to extract text information from the image.

[1157] Step 4:

[1158] The device converts the extracted text information into speech, using speech synthesis technology to generate natural-sounding speech data.

[1159] Step 5:

[1160] The device plays the generated audio data, providing the user with an audio notification through a speaker or earphones.

[1161] Environmental sound recording and important sound notifications

[1162] Step 1:

[1163] The device activates the microphone and records the surrounding environmental sounds at regular intervals to collect audio data that can be used to identify the environmental sounds.

[1164] Step 2:

[1165] The device analyzes the recorded voice data and uses voice recognition technology to extract text information from the recording.

[1166] Step 3:

[1167] The device analyzes the extracted text information to identify important sounds (e.g., horns, bells, etc.) and prepares notification actions based on the identified sounds.

[1168] Step 4:

[1169] The device will vibrate when important sounds are identified, allowing users to receive notifications through physical vibrations.

[1170] Step 5:

[1171] The device will display the notification content as text on the display, allowing the user to visually confirm the notification.

[1172] Incorporating and applying emotion engines

[1173] Step 1:

[1174] The device uses a camera to capture the user's facial expressions, and the emotion engine analyzes the user's emotional state based on the expressions.

[1175] Step 2:

[1176] The device adjusts the content and method of notifications based on the emotional state recognized by the emotion engine. For example, if the user is feeling anxious, the tone of the notification may be softened.

[1177] Step 3:

[1178] The device will reflect the changed notification settings and provide sound, vibration, and text notification, whichever is best for the user.

[1179] Specific examples

[1180] 1. Support for visually impaired people walking on the road:

[1181] When a user is walking down the sidewalk, the device captures images of the area ahead with its camera and uses OCR technology to detect "STOP" signs. This information is then converted into "There is a STOP sign ahead" using speech synthesis technology, which the user receives as a voice notification.

[1182] If the user looks anxious, the notification tone will be gentler, increasing the user's sense of security and ensuring safety.

[1183] 2. Support for the hearing impaired in cafes:

[1184] While the user is waiting for their order at a cafe, the device's microphone records ambient sounds and uses voice recognition technology to recognize that "your name has been called." As a result, the device vibrates and displays the message "Your name has been called" on the display.

[1185] If the user has a happy expression, the notification content will be adjusted appropriately, changing the vibration intensity and the way the display is presented.

[1186] Example 2

[1187] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1188] Users with visual or hearing impairments have difficulty obtaining information to act safely in their daily lives, and difficulty receiving appropriate notifications according to the situation. They also have difficulty distinguishing between external voice instructions and environmental sounds, which makes it difficult to respond appropriately in situations where a quick response is required.

[1189] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for detecting obstacles and signs from camera images, a means for converting information about the detected obstacles and signs into sound and notifying the user, a means for recording environmental sounds and identifying specific sounds, a means for notifying the user of the identified specific sounds by vibration and text display, a means for analyzing the user's facial expression and recognizing emotions, and a means for changing the content and method of notification based on the recognized emotions. This allows users with visual or hearing impairments to act safely and appropriately respond to various situations in daily life.

[1190] "Camera footage" refers to visual information captured by a camera installed on a device such as a smartphone.

[1191] An "obstacle" is an object or structure that impedes the movement or activity of a visually impaired person.

[1192] A "sign" is any public or private display intended to provide important information to persons with visual impairments.

[1193] The "means for converting into voice" is a technology for converting text information into voice signals, and is a system for providing the user with auditory information.

[1194] "Environmental sound" refers to all sound information obtained from the surrounding acoustic environment.

[1195] "Specific sounds" are sounds that are important in general or in specific situations (e.g., horns, bells, warning sounds, etc.).

[1196] "Vibration" is a means of providing sensory feedback to the user by physically shaking the device.

[1197] "Text display" is a means of displaying text information on a terminal display.

[1198] "Facial expressions" are movements and muscle patterns that appear on the user's face and indicate emotions or states.

[1199] "Means for recognizing emotions" refers to technology that analyzes the user's facial expressions to estimate their current mental state and emotions.

[1200] The "means for changing the notification content and method" is a technology for adjusting the information to be notified and the notification method based on the recognized emotion.

[1201] The present invention is a smartphone application that supports the daily lives of visually or hearing impaired users. This application uses the following specific means:

[1202] First, the smartphone camera is used to capture real-time video. The device activates the camera and continuously captures video. This video shows the user's surroundings and is used to detect obstacles and signs.

[1203] The device then captures the video frame by frame and converts it to grayscale using an image processing library (e.g., OpenCV), then uses an image recognition model (e.g., YOLO, SSD) to detect obstacles and signs, and then uses optical character recognition (OCR) technology to extract text information from signs.

[1204] The device then stores the identified obstacles and signs in text format and converts them into audio using a speech synthesis engine (e.g., Google Text-to-Speech), which is then transmitted to the user via earphones or speakers.

[1205] The device also uses the smartphone's built-in microphone to record ambient sounds. The recorded audio data is preprocessed and analyzed using a speech recognition model (e.g., Google Cloud Speech-to-Text) to identify specific sounds such as a horn, a bicycle bell, or a baby crying.

[1206] When a specific sound is identified, the device will activate the vibration motor and display a notification on the display, allowing users to be aware of important sounds without relying on hearing.

[1207] In addition, the device uses the camera to capture the user's face and analyzes their facial expressions using an emotion recognition model (e.g., FaceAPI). This allows the device to recognize the user's emotional state. Based on the recognized emotion, the device can adjust the content and method of notifications. For example, if the user has an anxious expression, the device can respond by softening the tone of the notification.

[1208] As a concrete example, consider a scenario in which a visually impaired person is walking down a street. As the user walks along the sidewalk, the device uses a camera to capture images of the area ahead and uses OCR to detect a "STOP" sign. This information is converted into "There is a STOP sign ahead" using speech synthesis technology, and the user receives a voice notification through their earphones, allowing them to act safely. Furthermore, if the user looks anxious, the notification is delivered in a gentler tone.

[1209] As another example, consider a scenario in which a hearing-impaired person is waiting to order in a noisy cafe. While the user is waiting, the device's microphone records the ambient sounds and uses voice recognition technology to recognize the message "your name has been called." As a result, the device notifies the user by vibrating and displays "Your name has been called" on the display so that the user is aware of the call. Furthermore, if the user has a happy expression, the display method and the intensity of the vibration can be adjusted.

[1210] An example of a prompt is as follows:

[1211] "When a visually impaired person is walking down the street, the system uses a smartphone camera to detect STOP signs and notifies them of this information via audio."

[1212] "While a hearing-impaired person is waiting to order in a noisy cafe, they will be notified by vibration and display when their name is called."

[1213] In this way, the present invention can support the daily lives of users with visual or hearing impairments, promoting their independence and social participation.

[1214] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1215] Step 1:

[1216] Camera footage capture

[1217] The device activates the smartphone camera and continuously captures video in real time.

[1218] Input: Visual information of the user's surroundings.

[1219] Output: Real-time video data.

[1220] Specific operation: The user launches the app and presses the "Launch Camera" button. The device begins capturing video from the camera.

[1221] Step 2:

[1222] Video pre-processing

[1223] The device captures the video frame by frame and converts it to grayscale using an image processing library (e.g., OpenCV).

[1224] Input: Captured real-time video data.

[1225] Output: Video data converted to grayscale.

[1226] Specific operation: The acquired RGB video data is converted to grayscale and filter processing is applied to remove noise.

[1227] Step 3:

[1228] Obstacle and sign detection

[1229] The device inputs the preprocessed video data into an image recognition model (e.g., YOLO, SSD) to detect obstacles and signs.

[1230] Input: Grayscale video data.

[1231] Output: Information of detected obstacles and signs (position and type).

[1232] How it works: Image recognition models analyze specific shapes and text to identify things like "STOP" signs and pedestrians.

[1233] Step 4:

[1234] Extracting text information from signs

[1235] The device extracts text information from parts of the detected signs using optical character recognition (OCR) technology.

[1236] Input: Image data of the sign.

[1237] Output: The extracted text information (e.g. "STOP").

[1238] Specific operation: Using an OCR engine, text information is extracted from the sign image data and saved in text format.

[1239] Step 5:

[1240] Text-to-Speech and Notifications

[1241] The device inputs the text information into a speech synthesis engine (e.g., Google Text-to-Speech) to convert it into speech, and notifies the user of the generated speech through earphones or speakers.

[1242] Input: The extracted text information.

[1243] Output: Audio data.

[1244] Specific operation: The text "There is a STOP sign ahead" is input into the speech synthesis engine, and the generated speech data is played back.

[1245] Step 6:

[1246] Environmental sound recording

[1247] The device uses the smartphone's built-in microphone to record environmental sounds.

[1248] Input: Ambient sound.

[1249] Output: Recorded audio data.

[1250] Specific operation: The user launches the app and presses the "Start Recording" button. The device begins recording ambient sounds.

[1251] Step 7:

[1252] Audio data preprocessing

[1253] The device samples the recorded audio data and applies a noise reduction filter.

[1254] Input: Recorded audio data.

[1255] Output: Preprocessed audio data.

[1256] What it does: Reduces background noise from audio data and normalizes audio clips.

[1257] Step 8:

[1258] Identifying specific sounds

[1259] The device inputs the preprocessed audio data into a speech recognition model (e.g., Google Cloud Speech-to-Text) to identify specific sounds (e.g., a horn, a bicycle bell, a baby crying, etc.).

[1260] Input: Preprocessed audio data.

[1261] Output: Identification result of specific sound.

[1262] Specific operation: Speech waveform data is converted into a spectrogram, and a speech recognition model identifies certain patterns.

[1263] Step 9:

[1264] Vibration and text notifications

[1265] When the device identifies a specific sound, it activates the vibration motor and displays a notification on the display.

[1266] Input: Information about the identified specific sound.

[1267] Output: Vibration and display.

[1268] Specific operation: If the identified sound is a "horn," the device will begin vibrating and the message "Horn has been honked" will be displayed on the screen.

[1269] Step 10:

[1270] emotion recognition

[1271] The device uses a camera to capture video of the user's face and analyzes facial expressions using an emotion recognition model (e.g., FaceAPI).

[1272] Input: Video data of the user's face.

[1273] Output: Perceived emotional state.

[1274] Specific operation: Analyzes the feature points of the user's face and determines whether they correspond to "happiness," "anxiety," or "anger."

[1275] Step 11:

[1276] Changes to notification content

[1277] The device will change the content and presentation of notifications based on the perceived emotion, for example softening the tone of notifications if the user is anxious.

[1278] Input: Perceived emotional state.

[1279] Output: Tailored notification content and method.

[1280] What it does: If anxiety is detected, it will change the settings of the speech synthesis engine to make the tone of notifications gentler.

[1281] (Application example 2)

[1282] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1283] It is necessary to provide a means to improve safety and efficiency for users with visual or hearing impairments, who have difficulty accurately understanding and responding to their surroundings in their daily lives and work environments. In addition, there is a lack of functionality to adjust notification methods to the user's emotional state and to simultaneously analyze multiple sensory information.

[1284] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1285] In this invention, the server includes means for detecting obstacles and signs from camera images, means for converting information about detected obstacles and signs into audio, means for recording environmental sounds and identifying specific sounds, means for notifying the identified specific sounds by vibration and text display, means for recognizing the user's emotions and appropriately changing the notification content based on the user's emotional state, and means for simultaneously capturing camera images and microphone audio and detecting obstacles and specific sounds in real time. This enables users with visual or hearing impairments to act safely and efficiently in real time, improving their awareness and response to their surroundings.

[1286] "Visually impaired persons" refers to people who have visual impairments and have difficulty obtaining information through their eyesight in their daily lives or working environments.

[1287] "Hearing impaired" refers to people who have hearing impairments and have difficulty obtaining information through sound in their daily lives or working environments.

[1288] "Smart devices" refers to portable information terminals with advanced functions such as smartphones, tablets, and smartwatches.

[1289] "Camera footage" refers to video data captured by a camera to obtain visual information.

[1290] "Obstacle" refers to a physical object that impedes a user's movement or activity.

[1291] "Sign" refers to an object on a road or building that has figures or letters on it to provide information or instructions.

[1292] "Speech synthesis" refers to the technology of analyzing text data and converting it into voice data.

[1293] "Environmental sounds" refers to natural and artificial sounds that exist in the surrounding environment.

[1294] "Specific sounds" refer to sounds that need to be specifically recognized, such as horns, bells, and alarms.

[1295] "Vibration" refers to the technology in which a device vibrates to convey information to the user.

[1296] "Character display" refers to the technology of displaying text information on a display.

[1297] An "emotion engine" refers to an algorithm or system that analyzes a user's facial expressions and behavior to recognize their emotional state.

[1298] "Simultaneous capture" refers to the process of capturing camera video and microphone audio at the same time.

[1299] "Real-time" refers to processing and reaction occurring immediately, without delay.

[1300] To implement this invention, it is important to understand the system configuration and its specific operation shown below. The system is composed of a smart device, an industrial camera, a high-sensitivity microphone, a built-in vibration motor, and an LCD display. Various processes are also realized using open source libraries and cloud APIs.

[1301] First, an industrial camera (e.g., industrial camera) is used to capture video in real time. The video data is processed using the OpenCV library to detect obstacles and signs. The TensorFlow library is then used to perform image analysis using a deep learning model. An OCR engine (e.g., Tesseract) is also used to extract text information from signs.

[1302] Next, environmental sounds are recorded using a high-sensitivity microphone (e.g., a high-sensitivity microphone). The recorded audio data is analyzed using the DeepSpeech library to identify certain important sounds (e.g., horns, warning sounds). The identified sounds are notified to the user through the built-in vibration motor and LCD display (e.g., a TFT tactile display).

[1303] Furthermore, to recognize the user's emotional state, the system captures the user's facial expressions using camera footage and performs emotion analysis using the Microsoft Azure Emotion API. Based on the user's emotional state, the system appropriately adjusts the content and method of notifications (audio tone and notification frequency).

[1304] Particularly in a factory environment, it is necessary to simultaneously capture camera images and microphone audio, and detect obstacles and specific sounds in real time. The specific program for this is as follows:

[1305] As a concrete example, consider a scenario in which, when an obstacle is detected, a voice notification is given saying "Warning: Obstacle detected," and when a warning sound is detected in the ambient sound, the user is notified by vibration and a display. This system enables users with visual or hearing impairments to act safely and efficiently in real time.

[1306] Prompt Sentence Examples

[1307] Create a program that simultaneously captures camera images and microphone audio, detects obstacles and specific sounds in the factory in real time, and notifies the user by sound, vibration, and display using technologies such as OpenCV, TensorFlow, DeepSpeech, and Microsoft Azure Emotion API.

[1308] This concludes the "Mode for Carrying Out the Invention." Using this system, visually and hearing impaired users can significantly improve safety and efficiency in their living and working environments.

[1309] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1310] Step 1:

[1311] The terminal uses an industrial camera to capture images in real time.

[1312] Input: Real-time video from inside the factory

[1313] Output: Captured video data

[1314] How it works: The camera constantly captures images of the factory and generates video data, which is then passed on to the next processing step.

[1315] Step 2:

[1316] The device uses the OpenCV library to process the captured images and detect obstacles and signs.

[1317] Input: Video data acquired in step 1

[1318] Output: Location information of obstacles and signs

[1319] Specific operation: The OpenCV library analyzes video data and uses an object recognition algorithm to detect obstacles and signs. The detection results are generated as location information.

[1320] Step 3:

[1321] The device uses the TensorFlow library and an OCR engine (e.g., Tesseract) to parse and extract text information from signs.

[1322] Input: Location of signs detected in step 2

[1323] Output: Sign text information

[1324] Specific operation: The TensorFlow library is used to identify signs in the video, and the OCR engine is used to extract the text information written on the signs.

[1325] Step 4:

[1326] The device uses the Google Cloud Text-to-Speech API to convert the extracted text information into audio.

[1327] Input: Text information of signs extracted in step 3

[1328] Output: Audio data

[1329] Specific operation: Using the Google Cloud Text-to-Speech API, text information is converted into audio data and notified to the user.

[1330] Step 5:

[1331] The device uses a highly sensitive microphone to record ambient sounds.

[1332] Input: Environmental sounds inside the factory

[1333] Output: Recorded audio data

[1334] Specific operation: The microphone constantly records the surrounding environmental sounds and generates audio data.

[1335] Step 6:

[1336] The device uses the DeepSpeech library to analyze the recorded audio data and identify specific sounds.

[1337] Input: Audio data recorded in step 5

[1338] Output: Specific sound identification result

[1339] Specific operation: Performs voice recognition using the DeepSpeech library and identifies specific sounds such as warning sounds and horns.

[1340] Step 7:

[1341] The terminal notifies the user based on the identified specific sound using the built-in vibration motor and LCD display.

[1342] Input: The specific sound results identified in step 6

[1343] Output: Vibration and text notification

[1344] Specific operation: The device vibrates in response to a specific sound and displays a notification on the display to provide information to the user.

[1345] Step 8:

[1346] The device uses a camera to capture the user's facial expressions and analyzes their emotional state using the Microsoft Azure Emotion API.

[1347] Input: Video data of the user's face

[1348] Output: User's emotional state

[1349] Specific operation: The camera captures the user's face, passes the video data to the Emotion API to analyze emotions, and obtains the results.

[1350] Step 9:

[1351] The terminal appropriately changes the notification content and method based on the user's emotional state.

[1352] Input: The user's emotional state obtained in step 8

[1353] Output: Properly adjusted notification content and notification method

[1354] Specific operation: The tone and method of notification will change depending on the analyzed emotional state, notifying in a gentle tone if the user is anxious, and increasing the vibration if the user is impatient.

[1355] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1356] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1357] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1358] [Fourth embodiment]

[1359] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1360] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1361] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1362] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1363] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1364] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1365] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1366] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1367] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1368] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1369] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1370] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1371] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1372] This is a smartphone application that supports the daily lives of users with visual or hearing impairments. This application includes a means for detecting obstacles and signs from camera images and converting that information into audio, and a means for recording environmental sounds, identifying specific sounds, and notifying them with vibrations and text, thereby promoting the user's independence and participation in society.

[1373] Detects obstacles and signs from camera footage and provides audio notifications

[1374] 1. Capture camera footage:

[1375] The device uses a camera to capture real-time images that show the user's surroundings and are used to detect obstacles and signs.

[1376] 2. Obstacle and sign detection:

[1377] The device analyzes the captured video to detect obstacles and signs. Image processing and optical character recognition (OCR) technologies are used to analyze the video. Specifically, the video is converted to grayscale and text information is extracted.

[1378] 3. Speech synthesis and notifications:

[1379] The device receives information about detected obstacles and signs as text and converts it into speech using speech synthesis technology. This speech is then provided to the user via earphones or speakers, providing them with information on how to act safely.

[1380] Environmental sound recording and important sound notifications

[1381] 1. Environmental Sound Recording:

[1382] The device uses a microphone to record ambient sounds, and the recorded audio data is used to detect certain important sounds.

[1383] 2. Identifying specific sounds:

[1384] The device uses voice recognition technology to analyze recorded audio data and identify specific sounds such as a horn, a bicycle bell, a kettle, a baby crying, or someone calling your name in a public place.

[1385] 3. Vibration and text notification:

[1386] Based on the identified specific sound, the device will vibrate and display the notification content as text on the display, allowing hearing-impaired people to respond appropriately to important sounds.

[1387] Specific examples

[1388] 1. Support for visually impaired people walking on the road:

[1389] When a user is walking on the sidewalk, the device uses its camera to capture images of the area ahead and uses OCR to detect "STOP" signs. This information is converted into "There is a STOP sign ahead" using speech synthesis technology, and the user receives a voice notification through their earphones, providing them with information on how to act safely.

[1390] 2. Support for the hearing impaired in noisy cafes:

[1391] While a user is waiting to order at a cafe, the device's microphone records ambient sounds and uses voice recognition technology to recognize that "your name has been called." As a result, the device notifies the user by vibrating and displays "Your name has been called" on the display so that the user is aware of the call.

[1392] As a result, the present invention can support the daily lives of users with visual or hearing impairments, and promote their independence and social participation.

[1393] The processing flow will be explained below.

[1394] Detects obstacles and signs from camera footage and provides audio notifications

[1395] Step 1:

[1396] The device activates the camera and captures video in real time, specifically by using the camera device interface to obtain video frames.

[1397] Step 2:

[1398] The device preprocesses the captured video frames, specifically converting the video to grayscale to improve the efficiency of image processing.

[1399] Step 3:

[1400] The device detects obstacles and signs from the pre-processed video and uses optical character recognition (OCR) technology to extract text information from the video.

[1401] Step 4:

[1402] The device converts the extracted text information into speech, using speech synthesis technology to generate natural-sounding speech data.

[1403] Step 5:

[1404] The device plays the generated audio data, providing the user with an audio notification through a speaker or earphones.

[1405] Environmental sound recording and important sound notifications

[1406] Step 1:

[1407] The device will activate the microphone and record the surrounding environmental sounds at regular intervals.

[1408] Step 2:

[1409] The device analyzes the recorded voice data and uses voice recognition technology to extract text information from the recording.

[1410] Step 3:

[1411] The device analyzes the extracted text information and identifies important sounds (e.g., horns, bicycle bells, etc.) and prepares corresponding actions when a particular sound is detected.

[1412] Step 4:

[1413] The device will vibrate when an important sound is identified, notifying the user.

[1414] Step 5:

[1415] The device displays the notification content as text on the display, allowing the user to visually confirm the notification.

[1416] Example 1

[1417] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1418] Users with visual or hearing impairments have difficulty recognizing obstacles and signs, and distinguishing environmental sounds in their daily lives, which can limit their ability to travel safely and respond to emergencies. While technologies with visual and hearing assistance are necessary for these users to lead independent lives, currently available technologies do not adequately meet these needs. Specifically, improvements are needed in areas such as image recognition accuracy, prompt voice notifications, and real-time recognition of environmental sounds.

[1419] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1420] In this invention, the server includes means for detecting obstacles and signs from camera images, means for converting information about detected obstacles and signs into audio, means for recording environmental sounds and identifying specific sounds, means for performing grayscale conversion and optical character recognition through video analysis, means for converting text information into audio using speech synthesis technology, and means for identifying important sounds using speech recognition technology, thereby enabling users with visual or hearing impairments to travel safely and respond to emergencies.

[1421] "Camera footage" refers to real-time visual data acquired using the device's camera.

[1422] An "obstacle" is an object that exists in the user's direction of travel or around the user, and that obstructs the user's movement or actions.

[1423] A "sign" is a sign placed on a road or in a public place that displays letters or symbols to convey specific instructions or warnings.

[1424] "Means for converting into speech" refers to technology for converting text information into auditory information, and is a device or software that utilizes speech synthesis technology.

[1425] "Environmental sounds" are all sounds that exist around the user, and are audio data collected by a recording device.

[1426] A "specific sound" is a specific important sound that is identified from among environmental sounds, and is a sound that the user needs to recognize immediately.

[1427] "Vibration notification" refers to a technology that allows a device to generate specific vibrations to convey information to the user.

[1428] "Text display" is a means of visually providing information to a user by displaying text information on a terminal display.

[1429] "Grayscale conversion" is an image processing technique that converts camera images into images that contain only black and white shades.

[1430] Optical character recognition (OCR) is a technology that extracts character information from an image and converts it into text data.

[1431] "Speech synthesis technology" is a technology that converts text data into human speech and produces speech.

[1432] "Speech recognition technology" is a technology that analyzes recorded voice data and understands the meaning of language.

[1433] An "important sound" is a sound that should immediately draw the user's attention, such as a car horn, a baby crying, or someone calling your name.

[1434] The present invention relates to a smartphone application for supporting the daily lives of users with visual or hearing impairments. This application combines multiple technical means to help users live safely and independently.

[1435] Hardware and Software Configuration

[1436] 1. The device uses a smartphone as its main platform, which is equipped with a camera, microphone, speaker / earphone, vibration motor, and display.

[1437] 2. The server uses cloud services (e.g., Amazon Web Services, Google Cloud Platform) for data processing and analysis, which enables it to process large amounts of data in real time and provide users with timely information.

[1438] 3. The device includes the following major software components:

[1439] Image processing algorithms (e.g. OpenCV)

[1440] Optical character recognition (OCR) technology (e.g. Tesseract OCR)

[1441] Speech synthesis technology (e.g., Google Text-to-Speech, Amazon Polly)

[1442] Speech recognition technology (e.g., Google Cloud Speech-to-Text, Amazon Transcribe)

[1443] Data processing and calculation

[1444] 1. Capture camera footage:

[1445] The device uses the smartphone's camera to capture real-time images of the surrounding area, which are then stored in the device's internal memory and analyzed immediately.

[1446] 2. Obstacle and sign detection:

[1447] The device uses OpenCV to convert the video to grayscale and apply an edge detection algorithm, then uses Tesseract OCR to analyze the characters on the sign and extract the text information.

[1448] 3. Speech synthesis and notifications:

[1449] The device converts the extracted text information into speech using voice synthesis technologies such as Google Text-to-Speech or Amazon Polly, and the speech is then transmitted to the user through earphones or speakers.

[1450] 4. Recording environmental sounds and identifying specific sounds:

[1451] The device uses the smartphone's microphone to record surrounding sounds, which are then stored in the device's internal memory and analyzed using voice recognition technologies such as Google Cloud Speech-to-Text and Amazon Transcribe.

[1452] 5. Vibration and text notification:

[1453] The device will vibrate based on the identified specific sound and display the notification content as text on the display, allowing the user to immediately recognize important information.

[1454] Specific examples

[1455] 1. For a visually impaired person walking on the road:

[1456] As a user walks along the sidewalk, the device uses its camera to capture images of the area ahead. It uses OpenCV's image processing algorithms to recognize obstacles and "STOP" signs, converting that information into text using Tesseract OCR. It then uses Google Text-to-Speech technology to translate the text into audio, saying "There is a STOP sign ahead," and notifies the user through earphones.

[1457] 2. For a hearing impaired person in a noisy cafe:

[1458] While a user is waiting to order at a cafe, the device's microphone records surrounding sounds. Using Google Cloud Speech-to-Text technology, the device identifies the voice saying "Mr. / Ms. XX, your order is ready," notifying the user with a vibration and displaying "Your name has been called" on the display.

[1459] Prompt Sentence Examples

[1460] By inputting prompt statements such as the following into the generative AI model, an explanation of the system and specific examples can be generated.

[1461] Describe a smartphone application system that supports the daily lives of users with visual or hearing impairments. Explain in detail how it detects obstacles and signs from camera footage and converts that information into audio, and how it records environmental sounds, identifies specific sounds, and notifies the user by vibrating and displaying text. Give specific examples.

[1462] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1463] Step 1: Capture camera footage:

[1464] The device activates the smartphone camera and captures images of the surroundings in real time. The camera captures the scenery ahead as the user walks along the sidewalk. This image data is stored in the internal memory. The input is an image of the user's surroundings, and the output is real-time image data.

[1465] Step 2: Obstacle and sign detection:

[1466] The device uses OpenCV to analyze the captured video. Specifically, it converts the video to grayscale and applies an edge detection algorithm (Canny edge detection). After this, it uses Tesseract OCR to extract the text information of the sign from the video. The input is real-time video data, and the output is text information (e.g., a "STOP" sign).

[1467] Step 3: Text-to-Speech and Notifications:

[1468] The device converts the extracted text information into audio data using the Google Text-to-Speech API. For example, audio data such as "There is a STOP sign ahead" is generated. If the user is using earphones, the audio notification is transmitted through the earphones. The input is text information, and the output is audio data.

[1469] Step 4: Recording ambient sounds:

[1470] The device activates the smartphone's microphone and records the surrounding environmental sounds in real time. For example, it records the surrounding sounds when the user is in a cafe. This audio data is stored in the internal memory. The input is the user's surrounding sounds, and the output is audio data.

[1471] Step 5: Identifying specific sounds:

[1472] The device analyzes the recorded voice data using the Google Cloud Speech-to-Text API to identify specific important sounds (e.g., "Mr. / Ms. XX, your order is ready"). The input is the recorded voice data, and the output is the identified text information.

[1473] Step 6: Vibration and text notification:

[1474] The device vibrates based on the identified specific sound and displays the text "Your name has been called" on the display, thereby alerting the user to the important notification. The input is the identified text information, and the output is a notification via vibration and text display.

[1475] This allows users with visual or hearing impairments to move safely and quickly perceive important information.

[1476] (Application example 1)

[1477] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1478] Users with visual or hearing impairments face challenges in safely moving around in physical stores and receiving the information they need. In such situations, users may bump into obstacles or miss important announcements, limiting their independent movement. As a result, users with visual or hearing impairments face challenges in limiting their opportunities to participate in social life.

[1479] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1480] In this invention, the server includes means for detecting obstacles and signs from camera images, means for converting information about detected obstacles and signs into audio, means for recording environmental sounds and identifying specific sounds, means for detecting obstacles and guide signs in the physical store and notifying them by audio, and means for recording announcements in the physical store and notifying identified announcements by vibration and text display, thereby enabling users with visual or hearing impairments to move around the physical store safely and efficiently.

[1481] A "visually or hearing impaired user" is an individual who is visually or hearing impaired, or both.

[1482] An "information processing device" is a machine or device for inputting, processing, and outputting data.

[1483] "Camera video" refers to real-time video data acquired using a camera.

[1484] An "obstacle" is a physical object that blocks the user's movement.

[1485] A "sign" is a visual display that conveys information.

[1486] "Converting to voice" refers to converting text information into voice data using voice synthesis technology.

[1487] "Ambient sounds" are any sounds occurring around the user.

[1488] "Identifying specific sounds" means recognizing and extracting characteristic sounds from recorded audio data.

[1489] "Vibration" is a means of providing tactile feedback to the user by causing the device to vibrate.

[1490] "Character display" means displaying text information on a display or the like.

[1491] A "physical store" is a store located in a physical location that offers goods and services.

[1492] An "announcement" is an audio message intended to convey information in a public place or specific environment.

[1493] "Recording" means to record audio using a device such as a microphone.

[1494] "Natural language processing" is a computational technique for processing and understanding human language.

[1495] The present invention relates to an information processing device that supports users with visual or hearing impairments in moving around safely and efficiently in a physical store. This device analyzes camera images and environmental sounds and provides the detected information to the user through means such as voice, vibration, and text display.

[1496] Hardware and Software Configuration

[1497] The present invention uses the following hardware and software.

[1498] Hardware:

[1499] Camera: Used to capture video in real time.

[1500] Microphone: Used to record ambient sounds and announcements.

[1501] Smartphone or tablet: The platform for processing and notification.

[1502] Headphones or speakers: Used to provide audio notifications.

[1503] Display: Used to display text information.

[1504] software:

[1505] OpenCV (open source image processing library): Used to analyze camera footage and detect obstacles and signs.

[1506] PyTesseract (OCR library): Used to extract text information from captured video.

[1507] Pyttsx3 (speech synthesis library): Used to convert text information into speech.

[1508] SpeechRecognition (speech recognition library): Used to identify specific sounds and announcements from recorded audio data.

[1509] GTTs (Google Text-to-Speech): The identified voice information is further analyzed and used to generate the required notifications.

[1510] Vibrate Library: Used to vibrate when a specific sound is identified.

[1511] Specific processing steps

[1512] 1. Camera footage capture and analysis:

[1513] The device's camera captures video in real time and converts it to grayscale using OpenCV. PyTesseract is used to extract text information from the image and detect obstacles and signs. For example, signs such as "STOP," "toilet," and "exit" can be detected in a physical store, and speech synthesis technology (Pyttsx3) can be used to notify the user, such as "There is a STOP sign ahead."

[1514] 2. Environmental sound recording and analysis:

[1515] The device's microphone records environmental sounds, and the recorded audio data is analyzed using SpeechRecognition. For example, when an in-store announcement or a specific event (such as a tasting invitation) is detected, the vibration library is used to notify the user by vibrating, such as "Tasting invitation." At the same time, the text information "Tasting invitation" is displayed on the display.

[1516] Adding specific examples

[1517] As an example, the following prompt sentence can be used:

[1518] Example prompt sentence:

[1519] It uses cameras to detect footage as you walk through the store and notifies you of "STOP" signs via voice synthesis.

[1520] When a user is walking through a supermarket, the camera footage is analyzed to detect a "STOP" sign, and an audio notification is played from the smartphone speaker saying, "There is a STOP sign ahead."

[1521] It records ambient sounds within the store, identifies announcements from specific areas (such as the food sampling area), and notifies you with a vibration.

[1522] The smartphone's microphone is used to record in-store announcements, and voice recognition technology is used to identify specific announcements such as "Invitation to sample food" and notify the user by vibrating.

[1523] As a result, the present invention enables users with visual or hearing impairments to move around safely and efficiently within physical stores, promoting their independence and social participation.

[1524] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1525] Step 1:

[1526] Capture footage with the camera.

[1527] Input: Capture real-time video using the device's camera.

[1528] Data processing: This video data will be used as is in the next step.

[1529] Output: The video data sent to the next step.

[1530] Specific operation: The device's camera continuously captures images of the user's surroundings in real time.

[1531] Step 2:

[1532] The video data is converted to grayscale and text information is extracted.

[1533] Input: Video data obtained in step 1.

[1534] Data processing: Convert the video to grayscale using OpenCV and extract text information using PyTesseract.

[1535] Output: The extracted text data.

[1536] Specific operation: The device uses the OpenCV library to convert the video data to grayscale, and then performs OCR processing using PyTesseract.

[1537] Step 3:

[1538] Obstacles and signs are detected from the extracted text information.

[1539] Input: The text data extracted in step 2.

[1540] Data calculation: Detect specific keywords such as "STOP" and "toilet" from the extracted text data.

[1541] Output: Detected keywords and their locations.

[1542] Specific operation: The device analyzes the extracted text data and searches for specific keywords.

[1543] Step 4:

[1544] The detected information is notified to the user by voice.

[1545] Input: Keywords detected in step 3 and their location information.

[1546] Data calculation: Convert keywords into speech using Pyttsx3.

[1547] Output: Audio data.

[1548] Specific operation: The terminal uses the Pyttsx3 library to generate voice data such as "There is a STOP sign ahead" and notify the user through earphones or speakers.

[1549] Step 5:

[1550] Record the ambient sounds with a microphone.

[1551] Input: Ambient sounds around the user.

[1552] Data processing: Record the environmental sounds as they are.

[1553] Output: Recorded audio data.

[1554] What it does: The device's microphone continuously records ambient sounds.

[1555] Step 6:

[1556] Analyzes recorded audio data and identifies specific sounds.

[1557] Input: The audio data recorded in step 5.

[1558] Data calculation: Analyze audio data using SpeechRecognition to identify specific sounds (e.g., in-store announcements).

[1559] Output: The identified specific audio data.

[1560] What it does: The device uses the SpeechRecognition library to analyze the recorded audio data and identify specific sounds, such as "We're offering a tasting."

[1561] Step 7:

[1562] Based on the identified specific sound, the device notifies you with vibration and text display.

[1563] Input: The specific audio data identified in step 6.

[1564] Data calculation: Controls vibration and text display based on the identified voice data.

[1565] Output: Vibration and text information displayed on the display.

[1566] Specific operation: When the device identifies a specific sound, it uses the Vibrate Library to notify the user by vibrating, and at the same time displays text information such as "Tasting information available" on the display.

[1567] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1568] This invention is a smartphone application that supports the daily lives of users with visual or hearing impairments. This application includes a means for detecting obstacles and signs from camera images and converting that information into audio, as well as a means for recording environmental sounds, identifying specific sounds, and notifying them by vibration and displaying text. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, the notification method can be appropriately adjusted, further promoting the user's independence and social participation.

[1569] Detects obstacles and signs from camera footage and provides audio notifications

[1570] 1. Capture camera footage:

[1571] The device uses a camera to capture real-time images that show the user's surroundings and are used to detect obstacles and signs.

[1572] 2. Obstacle and sign detection:

[1573] The device analyzes the captured video to detect obstacles and signs. Image processing and optical character recognition (OCR) technologies are used to analyze the video. Specifically, the video is converted to grayscale and text information is extracted.

[1574] 3. Speech synthesis and notifications:

[1575] The device receives information about detected obstacles and signs as text and converts it into speech using speech synthesis technology. This speech is then provided to the user via earphones or speakers, providing them with information on how to act safely.

[1576] Environmental sound recording and important sound notifications

[1577] 1. Environmental Sound Recording:

[1578] The device uses a microphone to record ambient sounds, and the recorded audio data is used to detect certain important sounds.

[1579] 2. Identifying specific sounds:

[1580] The device uses voice recognition technology to analyze recorded audio data and identify specific sounds such as a horn, a bicycle bell, a kettle, a baby crying, or someone calling your name in a public place.

[1581] 3. Vibration and text notification:

[1582] Based on the identified specific sound, the device will vibrate and display the notification content as text on the display, allowing hearing-impaired people to respond appropriately to important sounds.

[1583] Incorporating and applying emotion engines

[1584] 1. Emotion recognition:

[1585] The device uses a camera to capture the user's facial expressions and analyzes them with an emotion engine to recognize the user's current emotional state (e.g., joy, anxiety, anger, etc.).

[1586] 2. Changes in notification methods:

[1587] The device can then adjust the notification content and method appropriately based on the user's emotional state as recognized by the emotion engine. For example, if the user is feeling anxious, the device can soften the tone of the notification.

[1588] Specific examples

[1589] 1. Support for visually impaired people walking on the road:

[1590] As a user walks down the sidewalk, the device uses its camera to capture images of the area ahead and uses OCR to detect "STOP" signs. This information is then converted into "There is a STOP sign ahead" using speech synthesis technology, and the user receives a voice notification through their earphones, providing them with information on how to act safely. If the user's facial expression indicates anxiety, the notification is delivered in a gentler tone.

[1591] 2. Support for the hearing impaired in noisy cafes:

[1592] While a user is waiting to order at a cafe, the device's microphone records ambient sounds and uses voice recognition technology to recognize "your name has been called." As a result, the device notifies the user by vibrating and displays "Your name has been called" on the display so that the user is aware of the call. Furthermore, if the user has a happy expression, the display method and vibration intensity can be adjusted.

[1593] As a result, the present invention can support the daily lives of users with visual or hearing impairments, and promote their independence and social participation.

[1594] The processing flow will be explained below.

[1595] Detects obstacles and signs from camera footage and provides audio notifications

[1596] Step 1:

[1597] The device activates the camera and captures video in real time, specifically by using the camera device interface to obtain video frames.

[1598] Step 2:

[1599] The device pre-processes the captured video frames, which includes converting the video to grayscale for efficient image processing.

[1600] Step 3:

[1601] The device detects obstacles and signs from the grayscale image and uses optical character recognition (OCR) technology to extract text information from the image.

[1602] Step 4:

[1603] The device converts the extracted text information into speech, using speech synthesis technology to generate natural-sounding speech data.

[1604] Step 5:

[1605] The device plays the generated audio data, providing the user with an audio notification through a speaker or earphones.

[1606] Environmental sound recording and important sound notifications

[1607] Step 1:

[1608] The device activates the microphone and records the surrounding environmental sounds at regular intervals to collect audio data that can be used to identify the environmental sounds.

[1609] Step 2:

[1610] The device analyzes the recorded voice data and uses voice recognition technology to extract text information from the recording.

[1611] Step 3:

[1612] The device analyzes the extracted text information to identify important sounds (e.g., horns, bells, etc.) and prepares notification actions based on the identified sounds.

[1613] Step 4:

[1614] The device will vibrate when important sounds are identified, allowing users to receive notifications through physical vibrations.

[1615] Step 5:

[1616] The device will display the notification content as text on the display, allowing the user to visually confirm the notification.

[1617] Incorporating and applying emotion engines

[1618] Step 1:

[1619] The device uses a camera to capture the user's facial expressions, and the emotion engine analyzes the user's emotional state based on the expressions.

[1620] Step 2:

[1621] The device adjusts the content and method of notifications based on the emotional state recognized by the emotion engine. For example, if the user is feeling anxious, the tone of the notification may be softened.

[1622] Step 3:

[1623] The device will reflect the changed notification settings and provide sound, vibration, and text notification, whichever is best for the user.

[1624] Specific examples

[1625] 1. Support for visually impaired people walking on the road:

[1626] When a user is walking down the sidewalk, the device captures images of the area ahead with its camera and uses OCR technology to detect "STOP" signs. This information is then converted into "There is a STOP sign ahead" using speech synthesis technology, which the user receives as a voice notification.

[1627] If the user looks anxious, the notification tone will be gentler, increasing the user's sense of security and ensuring safety.

[1628] 2. Support for the hearing impaired in cafes:

[1629] While the user is waiting for their order at a cafe, the device's microphone records ambient sounds and uses voice recognition technology to recognize that "your name has been called." As a result, the device vibrates and displays the message "Your name has been called" on the display.

[1630] If the user has a happy expression, the notification content will be adjusted appropriately, changing the vibration intensity and the way the display is presented.

[1631] Example 2

[1632] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1633] Users with visual or hearing impairments have difficulty obtaining information to act safely in their daily lives, and difficulty receiving appropriate notifications according to the situation. They also have difficulty distinguishing between external voice instructions and environmental sounds, which makes it difficult to respond appropriately in situations where a quick response is required.

[1634] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for detecting obstacles and signs from camera images, a means for converting information about the detected obstacles and signs into sound and notifying the user, a means for recording environmental sounds and identifying specific sounds, a means for notifying the user of the identified specific sounds by vibration and text display, a means for analyzing the user's facial expression and recognizing emotions, and a means for changing the content and method of notification based on the recognized emotions. This allows users with visual or hearing impairments to act safely and appropriately respond to various situations in daily life.

[1635] "Camera footage" refers to visual information captured by a camera installed on a device such as a smartphone.

[1636] An "obstacle" is an object or structure that impedes the movement or activity of a visually impaired person.

[1637] A "sign" is any public or private display intended to provide important information to persons with visual impairments.

[1638] The "means for converting into voice" is a technology for converting text information into voice signals, and is a system for providing the user with auditory information.

[1639] "Environmental sound" refers to all sound information obtained from the surrounding acoustic environment.

[1640] "Specific sounds" are sounds that are important in general or in specific situations (e.g., horns, bells, warning sounds, etc.).

[1641] "Vibration" is a means of providing sensory feedback to the user by physically shaking the device.

[1642] "Text display" is a means of displaying text information on a terminal display.

[1643] "Facial expressions" are movements and muscle patterns that appear on the user's face and indicate emotions or states.

[1644] "Means for recognizing emotions" refers to technology that analyzes the user's facial expressions to estimate their current mental state and emotions.

[1645] The "means for changing the notification content and method" is a technology for adjusting the information to be notified and the notification method based on the recognized emotion.

[1646] The present invention is a smartphone application that supports the daily lives of visually or hearing impaired users. This application uses the following specific means:

[1647] First, the smartphone camera is used to capture real-time video. The device activates the camera and continuously captures video. This video shows the user's surroundings and is used to detect obstacles and signs.

[1648] The device then captures the video frame by frame and converts it to grayscale using an image processing library (e.g., OpenCV), then uses an image recognition model (e.g., YOLO, SSD) to detect obstacles and signs, and then uses optical character recognition (OCR) technology to extract text information from signs.

[1649] The device then stores the identified obstacles and signs in text format and converts them into audio using a speech synthesis engine (e.g., Google Text-to-Speech), which is then transmitted to the user via earphones or speakers.

[1650] The device also uses the smartphone's built-in microphone to record ambient sounds. The recorded audio data is preprocessed and analyzed using a speech recognition model (e.g., Google Cloud Speech-to-Text) to identify specific sounds such as a horn, a bicycle bell, or a baby crying.

[1651] When a specific sound is identified, the device will activate the vibration motor and display a notification on the display, allowing users to be aware of important sounds without relying on hearing.

[1652] In addition, the device uses the camera to capture the user's face and analyzes their facial expressions using an emotion recognition model (e.g., FaceAPI). This allows the device to recognize the user's emotional state. Based on the recognized emotion, the device can adjust the content and method of notifications. For example, if the user has an anxious expression, the device can respond by softening the tone of the notification.

[1653] As a concrete example, consider a scenario in which a visually impaired person is walking down a street. As the user walks along the sidewalk, the device uses a camera to capture images of the area ahead and uses OCR to detect a "STOP" sign. This information is converted into "There is a STOP sign ahead" using speech synthesis technology, and the user receives a voice notification through their earphones, allowing them to act safely. Furthermore, if the user looks anxious, the notification is delivered in a gentler tone.

[1654] As another example, consider a scenario in which a hearing-impaired person is waiting to order in a noisy cafe. While the user is waiting, the device's microphone records the ambient sounds and uses voice recognition technology to recognize the message "your name has been called." As a result, the device notifies the user by vibrating and displays "Your name has been called" on the display so that the user is aware of the call. Furthermore, if the user has a happy expression, the display method and the intensity of the vibration can be adjusted.

[1655] An example of a prompt is as follows:

[1656] "When a visually impaired person is walking down the street, the system uses a smartphone camera to detect STOP signs and notifies them of this information via audio."

[1657] "While a hearing-impaired person is waiting to order in a noisy cafe, they will be notified by vibration and display when their name is called."

[1658] In this way, the present invention can support the daily lives of users with visual or hearing impairments, promoting their independence and social participation.

[1659] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1660] Step 1:

[1661] Camera footage capture

[1662] The device activates the smartphone camera and continuously captures video in real time.

[1663] Input: Visual information of the user's surroundings.

[1664] Output: Real-time video data.

[1665] Specific operation: The user launches the app and presses the "Launch Camera" button. The device begins capturing video from the camera.

[1666] Step 2:

[1667] Video pre-processing

[1668] The device captures the video frame by frame and converts it to grayscale using an image processing library (e.g., OpenCV).

[1669] Input: Captured real-time video data.

[1670] Output: Video data converted to grayscale.

[1671] Specific operation: The acquired RGB video data is converted to grayscale and filter processing is applied to remove noise.

[1672] Step 3:

[1673] Obstacle and sign detection

[1674] The device inputs the preprocessed video data into an image recognition model (e.g., YOLO, SSD) to detect obstacles and signs.

[1675] Input: Grayscale video data.

[1676] Output: Information of detected obstacles and signs (position and type).

[1677] How it works: Image recognition models analyze specific shapes and text to identify things like "STOP" signs and pedestrians.

[1678] Step 4:

[1679] Extracting text information from signs

[1680] The device extracts text information from parts of the detected signs using optical character recognition (OCR) technology.

[1681] Input: Image data of the sign.

[1682] Output: The extracted text information (e.g. "STOP").

[1683] Specific operation: Using an OCR engine, text information is extracted from the sign image data and saved in text format.

[1684] Step 5:

[1685] Text-to-Speech and Notifications

[1686] The device inputs the text information into a speech synthesis engine (e.g., Google Text-to-Speech) to convert it into speech, and notifies the user of the generated speech through earphones or speakers.

[1687] Input: The extracted text information.

[1688] Output: Audio data.

[1689] Specific operation: The text "There is a STOP sign ahead" is input into the speech synthesis engine, and the generated speech data is played back.

[1690] Step 6:

[1691] Environmental sound recording

[1692] The device uses the smartphone's built-in microphone to record environmental sounds.

[1693] Input: Ambient sound.

[1694] Output: Recorded audio data.

[1695] Specific operation: The user launches the app and presses the "Start Recording" button. The device begins recording ambient sounds.

[1696] Step 7:

[1697] Audio data preprocessing

[1698] The device samples the recorded audio data and applies a noise reduction filter.

[1699] Input: Recorded audio data.

[1700] Output: Preprocessed audio data.

[1701] What it does: Reduces background noise from audio data and normalizes audio clips.

[1702] Step 8:

[1703] Identifying specific sounds

[1704] The device inputs the preprocessed audio data into a speech recognition model (e.g., Google Cloud Speech-to-Text) to identify specific sounds (e.g., a horn, a bicycle bell, a baby crying, etc.).

[1705] Input: Preprocessed audio data.

[1706] Output: Identification result of specific sound.

[1707] Specific operation: Speech waveform data is converted into a spectrogram, and a speech recognition model identifies certain patterns.

[1708] Step 9:

[1709] Vibration and text notifications

[1710] When the device identifies a specific sound, it activates the vibration motor and displays a notification on the display.

[1711] Input: Information about the identified specific sound.

[1712] Output: Vibration and display.

[1713] Specific operation: If the identified sound is a "horn," the device will begin vibrating and the message "Horn has been honked" will be displayed on the screen.

[1714] Step 10:

[1715] emotion recognition

[1716] The device uses a camera to capture video of the user's face and analyzes facial expressions using an emotion recognition model (e.g., FaceAPI).

[1717] Input: Video data of the user's face.

[1718] Output: Perceived emotional state.

[1719] Specific operation: Analyzes the feature points of the user's face and determines whether they correspond to "happiness," "anxiety," or "anger."

[1720] Step 11:

[1721] Changes to notification content

[1722] The device will change the content and presentation of notifications based on the perceived emotion, for example softening the tone of notifications if the user is anxious.

[1723] Input: Perceived emotional state.

[1724] Output: Tailored notification content and method.

[1725] What it does: If anxiety is detected, it will change the settings of the speech synthesis engine to make the tone of notifications gentler.

[1726] (Application example 2)

[1727] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1728] It is necessary to provide a means to improve safety and efficiency for users with visual or hearing impairments, who have difficulty accurately understanding and responding to their surroundings in their daily lives and work environments. In addition, there is a lack of functionality to adjust notification methods to the user's emotional state and to simultaneously analyze multiple sensory information.

[1729] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1730] In this invention, the server includes means for detecting obstacles and signs from camera images, means for converting information about detected obstacles and signs into audio, means for recording environmental sounds and identifying specific sounds, means for notifying the identified specific sounds by vibration and text display, means for recognizing the user's emotions and appropriately changing the notification content based on the user's emotional state, and means for simultaneously capturing camera images and microphone audio and detecting obstacles and specific sounds in real time. This enables users with visual or hearing impairments to act safely and efficiently in real time, improving their awareness and response to their surroundings.

[1731] "Visually impaired persons" refers to people who have visual impairments and have difficulty obtaining information through their eyesight in their daily lives or working environments.

[1732] "Hearing impaired" refers to people who have hearing impairments and have difficulty obtaining information through sound in their daily lives or working environments.

[1733] "Smart devices" refers to portable information terminals with advanced functions such as smartphones, tablets, and smartwatches.

[1734] "Camera footage" refers to video data captured by a camera to obtain visual information.

[1735] "Obstacle" refers to a physical object that impedes a user's movement or activity.

[1736] "Sign" refers to an object on a road or building that has figures or letters on it to provide information or instructions.

[1737] "Speech synthesis" refers to the technology of analyzing text data and converting it into voice data.

[1738] "Environmental sounds" refers to natural and artificial sounds that exist in the surrounding environment.

[1739] "Specific sounds" refer to sounds that need to be specifically recognized, such as horns, bells, and alarms.

[1740] "Vibration" refers to the technology in which a device vibrates to convey information to the user.

[1741] "Character display" refers to the technology of displaying text information on a display.

[1742] An "emotion engine" refers to an algorithm or system that analyzes a user's facial expressions and behavior to recognize their emotional state.

[1743] "Simultaneous capture" refers to the process of capturing camera video and microphone audio at the same time.

[1744] "Real-time" refers to processing and reaction occurring immediately, without delay.

[1745] To implement this invention, it is important to understand the system configuration and its specific operation shown below. The system is composed of a smart device, an industrial camera, a high-sensitivity microphone, a built-in vibration motor, and an LCD display. Various processes are also realized using open source libraries and cloud APIs.

[1746] First, an industrial camera (e.g., industrial camera) is used to capture video in real time. The video data is processed using the OpenCV library to detect obstacles and signs. The TensorFlow library is then used to perform image analysis using a deep learning model. An OCR engine (e.g., Tesseract) is also used to extract text information from signs.

[1747] Next, environmental sounds are recorded using a high-sensitivity microphone (e.g., a high-sensitivity microphone). The recorded audio data is analyzed using the DeepSpeech library to identify certain important sounds (e.g., horns, warning sounds). The identified sounds are notified to the user through the built-in vibration motor and LCD display (e.g., a TFT tactile display).

[1748] Furthermore, to recognize the user's emotional state, the system captures the user's facial expressions using camera footage and performs emotion analysis using the Microsoft Azure Emotion API. Based on the user's emotional state, the system appropriately adjusts the content and method of notifications (audio tone and notification frequency).

[1749] Particularly in a factory environment, it is necessary to simultaneously capture camera images and microphone audio, and detect obstacles and specific sounds in real time. The specific program for this is as follows:

[1750] As a concrete example, consider a scenario in which, when an obstacle is detected, a voice notification is given saying "Warning: Obstacle detected," and when a warning sound is detected in the ambient sound, the user is notified by vibration and a display. This system enables users with visual or hearing impairments to act safely and efficiently in real time.

[1751] Prompt Sentence Examples

[1752] Create a program that simultaneously captures camera images and microphone audio, detects obstacles and specific sounds in the factory in real time, and notifies the user by sound, vibration, and display using technologies such as OpenCV, TensorFlow, DeepSpeech, and Microsoft Azure Emotion API.

[1753] This concludes the "Mode for Carrying Out the Invention." Using this system, visually and hearing impaired users can significantly improve safety and efficiency in their living and working environments.

[1754] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1755] Step 1:

[1756] The terminal uses an industrial camera to capture images in real time.

[1757] Input: Real-time video from inside the factory

[1758] Output: Captured video data

[1759] How it works: The camera constantly captures images of the factory and generates video data, which is then passed on to the next processing step.

[1760] Step 2:

[1761] The device uses the OpenCV library to process the captured images and detect obstacles and signs.

[1762] Input: Video data acquired in step 1

[1763] Output: Location information of obstacles and signs

[1764] Specific operation: The OpenCV library analyzes video data and uses an object recognition algorithm to detect obstacles and signs. The detection results are generated as location information.

[1765] Step 3:

[1766] The device uses the TensorFlow library and an OCR engine (e.g., Tesseract) to parse and extract text information from signs.

[1767] Input: Location of signs detected in step 2

[1768] Output: Sign text information

[1769] Specific operation: The TensorFlow library is used to identify signs in the video, and the OCR engine is used to extract the text information written on the signs.

[1770] Step 4:

[1771] The device uses the Google Cloud Text-to-Speech API to convert the extracted text information into audio.

[1772] Input: Text information of signs extracted in step 3

[1773] Output: Audio data

[1774] Specific operation: Using the Google Cloud Text-to-Speech API, text information is converted into audio data and notified to the user.

[1775] Step 5:

[1776] The device uses a highly sensitive microphone to record ambient sounds.

[1777] Input: Environmental sounds inside the factory

[1778] Output: Recorded audio data

[1779] Specific operation: The microphone constantly records the surrounding environmental sounds and generates audio data.

[1780] Step 6:

[1781] The device uses the DeepSpeech library to analyze the recorded audio data and identify specific sounds.

[1782] Input: Audio data recorded in step 5

[1783] Output: Specific sound identification result

[1784] Specific operation: Performs voice recognition using the DeepSpeech library and identifies specific sounds such as warning sounds and horns.

[1785] Step 7:

[1786] The terminal notifies the user based on the identified specific sound using the built-in vibration motor and LCD display.

[1787] Input: The specific sound results identified in step 6

[1788] Output: Vibration and text notification

[1789] Specific operation: The device vibrates in response to a specific sound and displays a notification on the display to provide information to the user.

[1790] Step 8:

[1791] The device uses a camera to capture the user's facial expressions and analyzes their emotional state using the Microsoft Azure Emotion API.

[1792] Input: Video data of the user's face

[1793] Output: User's emotional state

[1794] Specific operation: The camera captures the user's face, passes the video data to the Emotion API to analyze emotions, and obtains the results.

[1795] Step 9:

[1796] The terminal appropriately changes the notification content and method based on the user's emotional state.

[1797] Input: The user's emotional state obtained in step 8

[1798] Output: Properly adjusted notification content and notification method

[1799] Specific operation: The tone and method of notification will change depending on the analyzed emotional state, notifying in a gentle tone if the user is anxious, and increasing the vibration if the user is impatient.

[1800] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1801] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1802] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1803] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1804] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1805] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1806] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1807] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, motorcycles, and other devices, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1808] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1809] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1810] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1811] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1812] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1813] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1814] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1815] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1816] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1817] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1818] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1819] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1820] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1821] The following is further disclosed regarding the above embodiment.

[1822] (Claim 1)

[1823] A smartphone application for supporting the daily lives of users with visual or hearing impairments,

[1824] means for detecting obstacles and signs from camera images;

[1825] means for converting detected obstacle and sign information into audio;

[1826] a means for recording environmental sounds and identifying specific sounds;

[1827] a means for notifying the identified specific sound by vibration and text display;

[1828] A system including:

[1829] (Claim 2)

[1830] 10. The system of claim 1, further comprising means for recognizing and executing voice commands.

[1831] (Claim 3)

[1832] 10. The system of claim 1, further comprising: means for generating a visual alert if the identified particular sound is an important sound.

[1833] "Example 1"

[1834] (Claim 1)

[1835] means for detecting obstacles and signs from camera images;

[1836] means for converting detected obstacle and sign information into audio;

[1837] a means for recording environmental sounds and identifying specific sounds;

[1838] a means for notifying the identified specific sound by vibration and text display;

[1839] means for performing grayscale conversion and optical character recognition through video analysis;

[1840] A means for converting text information into speech using speech synthesis technology;

[1841] a means for identifying important sounds through speech recognition technology;

[1842] A system including:

[1843] (Claim 2)

[1844] 10. The system of claim 1, further comprising means for recognizing and executing voice commands.

[1845] (Claim 3)

[1846] 10. The system of claim 1, further comprising: means for generating a visual alert if the identified particular sound is an important sound.

[1847] "Application Example 1"

[1848] (Claim 1)

[1849] An information processing device for supporting the daily lives of a user with a visual or hearing impairment,

[1850] means for detecting obstacles and signs from camera images;

[1851] means for converting detected obstacle and sign information into audio;

[1852] a means for recording environmental sounds and identifying specific sounds;

[1853] a means for notifying the identified specific sound by vibration and text display;

[1854] A means for detecting obstacles and guide signs in a physical store and notifying them by voice;

[1855] A means for recording announcements made in a physical store and notifying a user of identified announcements by vibration and text display;

[1856] A system including:

[1857] (Claim 2)

[1858] 10. The system of claim 1, further comprising means for recognizing and executing voice commands.

[1859] (Claim 3)

[1860] 10. The system of claim 1, further comprising: means for generating a visual alert if the identified particular sound is an important sound.

[1861] "Example 2: Combining Emotion Engines"

[1862] (Claim 1)

[1863] means for detecting obstacles and signs from camera images;

[1864] a means for converting detected obstacle and sign information into voice and notifying the information;

[1865] a means for recording environmental sounds and identifying specific sounds;

[1866] a means for notifying the identified specific sound by vibration and text display;

[1867] means for analyzing a user's facial expression and recognizing emotions;

[1868] means for modifying notification content and method based on the recognized emotion;

[1869] A system including:

[1870] (Claim 2)

[1871] 10. The system of claim 1, further comprising means for recognizing and executing voice commands.

[1872] (Claim 3)

[1873] 10. The system of claim 1, further comprising: means for generating a visual alert if the identified particular sound is an important sound.

[1874] "Application example 2 when combining emotion engines"

[1875] (Claim 1)

[1876] A smart device application for supporting the daily lives of visually or hearing impaired users, comprising:

[1877] means for detecting obstacles and signs from camera images;

[1878] means for converting detected obstacle and sign information into audio;

[1879] a means for recording environmental sounds and identifying specific sounds;

[1880] a means for notifying the identified specific sound by vibration and text display;

[1881] means for recognizing a user's emotions and appropriately modifying notification content based on the user's emotional state;

[1882] A means for simultaneously capturing camera images and microphone audio and detecting obstacles and specific sounds in real time;

[1883] A system including:

[1884] (Claim 2)

[1885] 10. The system of claim 1, further comprising means for recognizing and executing voice commands.

[1886] (Claim 3)

[1887] 10. The system of claim 1, further comprising: means for generating a visual alert if the identified particular sound is an important sound. [Explanation of symbols]

[1888] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A smartphone application for supporting the daily lives of users with visual or hearing impairments, means for detecting obstacles and signs from camera images; means for converting detected obstacle and sign information into audio; a means for recording environmental sounds and identifying specific sounds; a means for notifying the identified specific sound by vibration and text display; A system including:

2. The system of claim 1 further comprising means for recognizing and executing voice commands.

3. The system of claim 1 further comprising means for generating a visual alert if the identified particular sound is an important sound.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A