System

The system addresses the challenges faced by visually impaired individuals by providing comprehensive support through imaging, audio analysis, and real-time voice guidance, enabling safer and more independent daily activities.

JP2026017296APending Publication Date: 2026-02-04SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024118078
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2026-02-04

AI Technical Summary

Technical Problem

Visually impaired individuals face challenges in identifying obstacles, determining their location, distinguishing products, and recognizing people, which limits their ability to shop and go out safely and independently.

Method used

A system that includes imaging and audio recording means to capture surroundings, analyze images and sounds, and provide real-time voice notifications, location determination, and route guidance using wireless LAN signals, enabling comprehensive support for daily activities.

Benefits of technology

Enables visually impaired individuals to navigate obstacles, identify their location, recognize people, and obtain product information, thereby enhancing their safety and independence in daily life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026017296000001_ABST
    Figure 2026017296000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: The visually impaired person support system includes an imaging means for capturing a surrounding image, an image analysis means for analyzing the image captured by the imaging means and detecting an obstacle, and a notification means for notifying by voice on the basis of information detected by the image analysis means.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Visually impaired people face various problems in their daily lives, such as identifying obstacles, determining their current location, distinguishing products, and identifying people. This often limits their shopping and outings, reducing their quality of life. Conventional assistive technologies have struggled to comprehensively and efficiently resolve these problems. The purpose of this invention is to solve these problems and provide the visually impaired with the joy of shopping and outings. [Means for solving the problem]

[0005] The visually impaired support system of the present invention includes the following means.

[0006] 1. Provide an imaging means for capturing images of the surroundings, thereby obtaining visual information.

[0007] 2. An image analysis means is provided for analyzing the image captured by the imaging means and detecting obstacles, thereby identifying the presence of obstacles that cannot be visually confirmed.

[0008] 3. A notification means is provided that notifies the user by voice based on the information detected by the image analysis means, and provides the user with information about obstacles in real time.

[0009] 4. A location determination means is further provided that analyzes location information and determines the current location, thereby helping the user to understand their current location.

[0010] 5. Provide an audio recording means for capturing surrounding sounds, thereby obtaining audio information.

[0011] 6. A voice analysis means is provided that analyzes the voice captured by the voice recording means and identifies a specific person, allowing the user to identify people around them.

[0012] 7. A location determination means is provided that scans signals from wireless LAN access points and determines the current location within the building, allowing the user to know their current location within the building.

[0013] 8. A route guidance means is provided that calculates a route to the destination based on the location specifying means and provides voice guidance, thereby assisting the user in reaching the destination.

[0014] As a result, the present invention can provide comprehensive support for the daily lives of visually impaired people and allow them to enjoy shopping and going out.

[0015] Below, we provide definitions for each of the key terms contained in the claims.

[0016] "Imaging means" refers to a camera or sensor for capturing images of the user's surroundings.

[0017] "Image analysis means" refers to hardware or software that analyzes images captured by the imaging means and executes algorithms to detect obstacles and other objects.

[0018] "Notification means" refers to a speaker or speech synthesis engine for providing information to a user in audible or other forms.

[0019] "Location determination means" refers to a system that determines the user's current location using the signal strength of wireless LAN access points or other location determination technologies.

[0020] "Audio recording means" refers to a microphone or other audio input device for capturing ambient sounds.

[0021] "Audio analysis means" refers to hardware or software that analyzes audio captured by an audio recording means and executes an algorithm to identify a particular person.

[0022] The "route guidance means" refers to a system that calculates a route to a destination based on the location identification means and provides voice guidance to the user. [Brief explanation of the drawings]

[0023] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6]FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0024] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0025] First, the terms used in the following description will be explained.

[0026] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0027] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0028] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0029] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0030] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0031] [First embodiment]

[0032] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0033] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0034] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0035] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0036] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0037] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0038] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0039] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0040] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0041] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0042] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0043] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0044] The visually impaired support system of the present invention has a plurality of functions for solving various problems that users face in their daily lives. The embodiments of the present invention will be described in detail below.

[0045] Walking navigation function

[0046] 1. Image capture

[0047] The device captures images of the user's surroundings with its camera at regular intervals (e.g., every second).

[0048] The captured images are sent to a server in real time.

[0049] 2. Image Analysis

[0050] The server decodes the received image and runs an image recognition algorithm on it.

[0051] The server detects obstacles in the image, the distance to the obstacles, the signal status, and the presence or absence of walls.

[0052] 3. Information Generation and Notification

[0053] Based on the image analysis results, the server generates information to be notified to the user (e.g., "There is an obstacle 2 meters ahead").

[0054] The server sends the generated information to the terminal, which notifies the user using a voice synthesis engine.

[0055] Example: While the user is walking, the device announces, "There is an obstacle 2 meters ahead."

[0056] Mapping function for facilities accessible to the visually impaired

[0057] 1. Identifying your current location

[0058] The device periodically scans the signal strength of Wi-Fi access points in its current location.

[0059] The scan results are sent to the server.

[0060] 2. Location analysis and route guidance

[0061] The server compares the user's current location with a map database based on the signal from the wireless LAN access point.

[0062] The server determines the user's current location and calculates a route to the destination.

[0063] The server transmits route information to the terminal, and the terminal uses a voice synthesis engine to provide route guidance to the user.

[0064] Example: When a user enters the voice command "I want to go to the library," the device will provide voice guidance such as "Turn left and go 20 meters, then turn right."

[0065] Image Recognition Function

[0066] 1. Image capture

[0067] The user takes a photo of the product they are holding with the device's camera.

[0068] The captured image is sent to a server.

[0069] 2. Product Identification and Information Notification

[0070] The server uses an image recognition algorithm to identify the product in the image.

[0071] Based on the identification results, the server obtains detailed product information (name, price, description, etc.) from the Internet.

[0072] The server sends the acquired detailed product information to the terminal, which then notifies the user using a voice synthesis engine.

[0073] Example: When a user picks up an item and takes a picture of it with the camera, the device announces in a voice message, "This is the 347-page novel 'The Blue Bird' and costs 1,200 yen."

[0074] Registered user recognition function

[0075] 1. Audio capture

[0076] The device captures surrounding sounds using a microphone.

[0077] The captured audio data is sent to a server.

[0078] 2. Voice analysis and identity notification

[0079] The server uses a voice recognition algorithm to analyze the voice data and identify a specific person.

[0080] The server transmits the identification result to the terminal, and the terminal notifies the user using a voice synthesis engine.

[0081] Example: When a person in front of the user speaks, the device will announce in voice, "This is Tanaka-san."

[0082] The objective of the support system for visually impaired people of the present invention is to integrate these functions to efficiently and effectively solve the various problems that visually impaired people encounter in their daily lives, thereby enabling them to live safer and more independent lives.

[0083] The processing flow will be explained below.

[0084] Walking navigation function

[0085] Step 1:

[0086] The device captures images of the user's surroundings at regular intervals using a camera.

[0087] Step 2:

[0088] The device compresses the captured images and sends them to the server in real time.

[0089] Step 3:

[0090] The server decodes the received image and runs an image recognition algorithm on it.

[0091] Step 4:

[0092] The server detects obstacles in the image, the distance to the obstacles, the signal status, and the presence or absence of walls.

[0093] Step 5:

[0094] The server generates the information to be notified based on the results of image analysis (e.g., "There is an obstacle 2 meters ahead").

[0095] Step 6:

[0096] The server transmits the generated information to the terminal.

[0097] Step 7:

[0098] The information received by the terminal is passed to a speech synthesis engine to generate voice data.

[0099] Step 8:

[0100] The audio data generated by the terminal is notified to the user through the speaker.

[0101] Mapping function for facilities accessible to the visually impaired

[0102] Step 1:

[0103] The device scans the signal strength of wireless LAN access points.

[0104] Step 2:

[0105] The device sends the scan results to the server.

[0106] Step 3:

[0107] The server identifies the user's current location based on the signal from the wireless LAN access point by comparing it with a location information database.

[0108] Step 4:

[0109] The server determines the user's current location.

[0110] Step 5:

[0111] The server calculates the route to the destination based on the user's instructions (voice commands).

[0112] Step 6:

[0113] The server sends the route information to the terminal.

[0114] Step 7:

[0115] The route information received by the terminal is passed to a voice synthesis engine to generate voice data.

[0116] Step 8:

[0117] The voice guidance generated by the terminal is notified to the user through the speaker.

[0118] Image Recognition Function

[0119] Step 1:

[0120] The user takes a photo of the product they are holding with the device's camera.

[0121] Step 2:

[0122] The device sends the captured image to the server.

[0123] Step 3:

[0124] The server uses an image recognition algorithm to identify the product in the image.

[0125] Step 4:

[0126] Based on the product name identified by the server, detailed product information (such as name, price, and description) is obtained from the Internet.

[0127] Step 5:

[0128] The server organizes the detailed product information it has acquired and generates text data.

[0129] Step 6:

[0130] The server transmits the generated text data to the terminal.

[0131] Step 7:

[0132] The information received by the terminal is passed to a speech synthesis engine to generate voice data.

[0133] Step 8:

[0134] The audio data generated by the terminal is notified to the user through the speaker.

[0135] Registered user recognition function

[0136] Step 1:

[0137] The device captures surrounding sounds using a microphone.

[0138] Step 2:

[0139] The device sends the captured audio data to the server.

[0140] Step 3:

[0141] The server analyzes the voice data using a voice recognition algorithm.

[0142] Step 4:

[0143] The server compares the analysis results with a pre-registered voice database to identify the person.

[0144] Step 5:

[0145] The server generates the identified information as text data (e.g., "This is Mr. Tanaka").

[0146] Step 6:

[0147] The server transmits the generated text data to the terminal.

[0148] Step 7:

[0149] The information received by the terminal is passed to a speech synthesis engine to generate voice data.

[0150] Step 8:

[0151] The audio data generated by the terminal is notified to the user through the speaker.

[0152] Example 1

[0153] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0154] To solve the various challenges that visually impaired people face in their daily lives, it is necessary to properly obtain information about their surroundings and convey that information to them quickly and accurately. However, current systems lack comprehensive support that goes beyond simply detecting obstacles, covering things like determining their current location, identifying people, and even providing traffic light status and specific route guidance. This leaves visually impaired people at great risk.

[0155] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0156] In this invention, the server includes an imaging means for capturing an image of the surroundings, an image analysis means for analyzing the image captured by the imaging means and detecting obstacles, traffic light status, and the presence or absence of walls, an information generation means for generating notification information based on the information detected by the image analysis means, a notification means for notifying by voice the notification information generated by the information generation means, a location identification means for analyzing location information using a wireless LAN signal and identifying the current location, an audio recording means for capturing surrounding sounds, an audio analysis means for analyzing the sounds captured by the audio recording means and identifying a specific person, an information generation means for generating notification information based on the person information identified by the audio analysis means, and a notification means for notifying by voice the notification information generated by the information generation means.This enables visually impaired people to avoid obstacles while walking, identify their current location, and easily identify people, enabling them to live safe and independent lives.

[0157] An "imaging means" is a device or sensor for capturing an image of the surroundings.

[0158] "Image analysis means" refers to a set of algorithms and software for analyzing captured images and detecting obstacles, traffic light conditions, the presence or absence of walls, etc.

[0159] The "notification means" is a device or software that conveys the analyzed information to the user by voice.

[0160] The "information generation means" refers to logic or algorithms for generating information to be notified to the user based on the analysis results.

[0161] The "location determination means" is a device or software for determining the user's current location using a wireless LAN signal.

[0162] "Audio recording means" refers to a microphone or other audio recording device for capturing ambient audio.

[0163] "Audio analysis means" refers to algorithms or software that analyzes captured audio and identifies specific individuals.

[0164] The visually impaired support system of the present invention has a plurality of functions for solving various problems that users face in their daily lives. The embodiments of the present invention will be described in detail below.

[0165] Walking navigation function

[0166] Image capture

[0167] The device captures images of the user's surroundings at regular intervals (e.g., every second) using the built-in camera of a smartphone.

[0168] Sending images

[0169] The device immediately sends the captured image to the server using Wi-Fi or mobile data.

[0170] Image analysis

[0171] The server decodes the received images and uses image recognition algorithms such as TensorFlow and OpenCV to detect obstacles, traffic light status, and the presence or absence of walls.

[0172] Information Generation and Notification

[0173] Based on the analysis results, the server generates notification information such as "There is an obstacle 2 meters ahead." The information is sent from the server to the device, which then notifies the user using a voice synthesis engine such as Google Text-to-Speech.

[0174] Specific examples

[0175] While the user is walking, the device will announce to them by voice, "There is an obstacle 2 meters ahead."

[0176] Mapping function for facilities accessible to the visually impaired

[0177] Scan your current location

[0178] The device periodically scans the signal strength of surrounding Wi-Fi access points and sends the results of the scan to the server.

[0179] Location analysis and routing

[0180] The server uses signal data from wireless LAN access points to identify the user's current location using a map database such as Google Maps API and calculates the optimal route. The route information is sent to the device, which then uses a voice synthesis engine such as Google Text-to-Speech to provide route guidance to the user.

[0181] Specific examples

[0182] When a user inputs the voice command "I want to go to the library," the device will provide voice guidance such as "Turn left and go 20 meters, then turn right."

[0183] Image Recognition Function

[0184] Image capture

[0185] The user takes a photo of the product they are holding with the device's camera, and the captured image is sent to the server.

[0186] Product identification and information acquisition

[0187] The server identifies the product using an image recognition algorithm such as YOLOv3, and then uses the Amazon API or Rakuten API to obtain detailed product information (name, price, description, etc.).

[0188] Information Notification

[0189] The acquired product details are sent from the server to the terminal, which then notifies the user using a voice synthesis engine such as Google Text-to-Speech.

[0190] Specific examples

[0191] When the user picks up the product and takes a picture of it with the camera, the device announces in voice, "This is the 347-page novel 'The Blue Bird' and costs 1,200 yen."

[0192] Registered user recognition function

[0193] Audio capture

[0194] The device captures the surrounding sound with a microphone, and the captured sound data is sent to the server.

[0195] Audio analysis and notifications

[0196] The server analyzes the voice data using a speech recognition algorithm such as IBM Watson Speech to Text to identify a specific person, and the identification result is sent from the server to the device, which then notifies the user using a speech synthesis engine such as Google Text-to-Speech.

[0197] Specific examples

[0198] When a person in front of the user speaks, the device will announce aloud, "This is Tanaka-san."

[0199] Examples of prompt statements

[0200] 1. Walking navigation function

[0201] Prompt: "What is the process for the system that notifies you of nearby obstacles?"

[0202] 2. Mapping function for facilities accessible to the visually impaired

[0203] Prompt: "Please explain the process of the system that identifies your current location and provides directions to your destination."

[0204] 3. Image Recognition

[0205] Prompt: "Please tell me the process of the system that takes a picture of a product with a camera and provides detailed information about it."

[0206] 4. Registered User Recognition Function

[0207] Prompt: "Please explain the process for a system that analyzes surrounding sounds and identifies a specific person."

[0208] The above is an embodiment of the visually impaired support system of the present invention, and this system enables visually impaired people to live safer and more independent lives.

[0209] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0210] Walking navigation function

[0211] Step 1:

[0212] Image capture

[0213] The device captures images of the user's surroundings with a camera at regular intervals (every second). The input is image data obtained from the camera, and the output is the captured image file.

[0214] Specific operation: The device automatically launches the camera app and simulates pressing the shutter button every second.

[0215] Step 2:

[0216] Sending images

[0217] The image files captured by the device are sent to the server in real time. The input is the image file generated in step 1, and the output is the upload of image data to the server.

[0218] Specific operation: The device compresses the image data and uploads it to the server using the HTTPS protocol.

[0219] Step 3:

[0220] Image analysis

[0221] The server decodes the received image and analyzes it using an image recognition algorithm using TensorFlow, OpenCV, etc. The input is the image data sent to the server, and the output is the detection results such as obstacles, traffic light status, and the presence or absence of walls.

[0222] Specific operation: The server processes the received image, runs a recognition algorithm, and generates an analysis result.

[0223] Step 4:

[0224] information generation

[0225] The server generates information to be notified to the user based on the analysis results. The input is the analysis results obtained in step 3, and the output is a notification message.

[0226] What happens: The server uses a text generation algorithm to generate a message such as "There is an obstacle 2 meters ahead."

[0227] Step 5:

[0228] Information Notification

[0229] The server sends the generated notification message to the device, and the device uses a speech synthesis engine such as Google Text-to-Speech to notify the user audibly. The input is the generated notification message, and the output is the audio notification delivered to the user.

[0230] Specific operation: The device converts the received message into voice and notifies the user through the speaker, "There is an obstacle 2 meters ahead."

[0231] Mapping function for facilities accessible to the visually impaired

[0232] Step 1:

[0233] Scan your current location

[0234] The device periodically scans the signal strength of surrounding Wi-Fi access points. The input is the Wi-Fi signal captured by the Wi-Fi module, and the output is the signal strength data.

[0235] Specific operation: The device activates the Wi-Fi module and collects the surrounding network signal strength.

[0236] Step 2:

[0237] Sending scan results

[0238] The terminal sends the collected signal strength data to the server. The input is the signal strength data collected in step 1, and the output is the upload to the server.

[0239] Specific operation: The terminal compresses the signal data and uploads it to the server via HTTPS protocol.

[0240] Step 3:

[0241] Location analysis

[0242] The server analyzes the location information based on the signal data of the wireless LAN access point and determines the current location. The input is the transmitted signal strength data, and the output is the determined current location information.

[0243] What it does: The server analyzes the signal data and locates it by comparing it with a map database (e.g., Google Maps API).

[0244] Step 4:

[0245] Route calculation

[0246] The server calculates the optimal route from the specified current location to the destination. The input is the current location information and the destination information, and the output is the route information.

[0247] Specific operation: The server executes a route calculation algorithm (e.g., A search algorithm) to calculate the route to the destination.

[0248] Step 5:

[0249] Route guidance

[0250] The server sends route information to the device, which then uses a speech synthesis engine such as Google Text-to-Speech to provide route guidance to the user. The input is the calculated route information, and the output is voice guidance provided to the user.

[0251] Specific operation: The device converts the route information it receives into voice and notifies the user, "Turn left, go 20 meters, then turn right."

[0252] Image Recognition Function

[0253] Step 1:

[0254] Image capture

[0255] The user takes a photo of the product they are holding with the device's camera. The input is image data acquired from the camera, and the output is the captured image file.

[0256] Specific operation: The device automatically launches the camera app, and the user presses the shutter button.

[0257] Step 2:

[0258] Sending images

[0259] The device sends the captured image file to the server. The input is the captured image file, and the output is the upload of the image data to the server.

[0260] Specific operation: The device compresses the image data and uploads it to the server using the HTTPS protocol.

[0261] Step 3:

[0262] Product Identification

[0263] The server uses an image recognition algorithm (e.g., YOLOv3) to identify the product in the image. The input is the submitted image data, and the output is the identified product information.

[0264] Specific operation: The server analyzes the received image and identifies it as "This is a novel called 'The Blue Bird'."

[0265] Step 4:

[0266] Information acquisition

[0267] The server retrieves detailed product information (such as name, price, description, etc.) from the Internet. The input is the identified product information, and the output is the retrieved product detail data.

[0268] Specific operation: The server calls the API to obtain product data.

[0269] Step 5:

[0270] Information Notification

[0271] The server sends the acquired product details to the terminal, which then notifies the user using a speech synthesis engine such as Google Text-to-Speech. The input is the acquired product details data, and the output is a voice notification delivered to the user.

[0272] Specific operation: The device converts the received product information into voice and notifies the user, "This is the 347-page novel 'The Blue Bird', priced at 1,200 yen."

[0273] Registered user recognition function

[0274] Step 1:

[0275] Audio capture

[0276] The device captures the surrounding sound with a microphone. The input is the audio data collected by the microphone, and the output is the captured audio file.

[0277] Specific operation: The device activates the microphone and collects surrounding sounds.

[0278] Step 2:

[0279] Sending Audio

[0280] The device sends the captured audio file to the server. The input is the captured audio file, and the output is the audio data uploaded to the server.

[0281] Specific operation: The device compresses the audio data and uploads it to the server using the HTTPS protocol.

[0282] Step 3:

[0283] Audio analysis

[0284] The server analyzes the voice data using a speech recognition algorithm (e.g., IBM Watson Speech to Text) to identify a specific person. The input is the transmitted voice data, and the output is the identified person's information.

[0285] Specific operation: The server processes the voice data and generates an identification result such as "This is Mr. Tanaka."

[0286] Step 4:

[0287] Generating the identification results

[0288] The server generates notification information based on the identification result. The input is the identified person's information, and the output is the generated notification message.

[0289] Specific operation: The server generates a text message and sends it to the device.

[0290] Step 5:

[0291] Notification of identity

[0292] The device uses a speech synthesis engine such as Google Text-to-Speech to convert the notification message into speech and notify the user. The input is the generated notification message, and the output is the audio notification delivered to the user.

[0293] Specific operation: The device uses the speaker to announce, "This is Tanaka-san."

[0294] (Application example 1)

[0295] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0296] Due to their unique needs, it is difficult for visually impaired people to shop safely and comfortably in brick-and-mortar stores. To solve this problem, a system is needed that allows visually impaired people to independently select and navigate products in brick-and-mortar stores. In addition, since it is difficult for them to accurately grasp product information in real time, appropriate support is required.

[0297] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0298] In this invention, the server includes an imaging means for capturing an image of the surroundings, an image analysis means for analyzing the image captured by the imaging means and detecting obstacles, a notification means for providing a voice notification based on information detected by the image analysis means, an image recognition means for identifying products, an information acquisition means for acquiring detailed information about the identified products, and a notification means for providing a voice notification based on information acquired by the information acquisition means. This enables visually impaired people to receive product information by voice while safely moving around in a physical store.

[0299] An "imaging means" is a device that captures an image of the surroundings.

[0300] The "image analysis means" is a device that analyzes the image captured by the imaging means and detects objects such as obstacles.

[0301] The "notification means" is a device that notifies the user of information detected by the image analysis means and other information by voice.

[0302] An "image recognition means" is a device for identifying an item in a captured image.

[0303] The "information acquisition means" is a device that acquires detailed information (such as the name and price) of the identified product from the Internet or the like.

[0304] The "location specifying means" is a device that specifies the current location by analyzing the location information performed by the image analysis means.

[0305] An "audio recording means" is a device that captures ambient sounds.

[0306] The "voice analysis means" is a device that analyzes the voice captured by the voice recording means and identifies a specific person.

[0307] The present invention relates to a support system for visually impaired people to shop safely and comfortably in a brick-and-mortar store. Hereinafter, an embodiment of the present invention will be specifically described.

[0308] System Configuration

[0309] The system includes the following main components:

[0310] 1. An imaging means to capture images of the surroundings

[0311] 2. Image analysis means for analyzing images captured by the imaging means and detecting obstacles

[0312] 3. Notification means that notifies by voice based on information detected by image analysis means

[0313] 4. Image recognition methods for identifying products

[0314] 5. Information acquisition means for acquiring detailed information on identified products

[0315] 6. Notification method for notifying information by voice

[0316] Hardware and Software

[0317] Hardware:

[0318] Smartphone, smart glasses, or head-mounted display

[0319] Web camera (imaging means)

[0320] Microphone (audio recording means)

[0321] Speaker (notification means)

[0322] software:

[0323] OpenCV (library for image capture and analysis)

[0324] Google Text-to-Speech (gTTS) (voice notifications)

[0325] Image recognition API (product identification)

[0326] Internet connection (information acquisition)

[0327] Processing steps

[0328] When a user activates the system, the imaging means captures images of the surrounding area at regular intervals. The captured images are sent to a server in real time, and the image analysis means detects obstacles in the images and analyzes their positions. The notification means then notifies the user of the obstacle information by voice.

[0329] At the same time, if the user wants to identify a product, they take a photo of the product using a smartphone or smart glasses. The captured image is sent to a server, and the image recognition means identifies the product. Detailed information about the identified product is obtained from the Internet by an information acquisition means, and the information is notified to the user by voice via a notification means.

[0330] Specific examples

[0331] For example, when a user is walking in a shopping mall, the system will notify them that there is an obstacle two meters ahead, helping them to walk safely. Also, when a user picks up a product and takes a picture of it with the camera, the system will announce the product information in a voice message, saying, "This is a 347-page book, priced at 1,200 yen."

[0332] Prompt Sentence Examples

[0333] Example prompt for image recognition:

[0334] "Please recognize the item in this image and provide the item name and distance."

[0335] Example of a prompt for audio notification:

[0336] "2 meters ahead, there is an obstacle. It is a table."

[0337] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0338] Step 1:

[0339] The imaging means captures images of the surroundings at regular intervals (e.g., every second). The input is video data from the camera, which is captured and saved as image data. The output is the saved image data.

[0340] Step 2:

[0341] The device sends the captured image data to the server. The input is the image data obtained in step 1, which is uploaded to the server via the Internet. The output is the image data received by the server.

[0342] Step 3:

[0343] The server analyzes the image data it receives and identifies the presence and location of obstacles. The input is the image data received by the server, which is analyzed using an image analysis algorithm (e.g., OpenCV). The output is the location information of obstacles.

[0344] Step 4:

[0345] The server generates the content to be notified to the user based on the analysis results. The input is the obstacle location information obtained in step 3, and an appropriate notification message is created using the generative AI model. The output is the notification message (text data).

[0346] Step 5:

[0347] The server sends a notification message to the device, and the device notifies the user by voice. The input is the notification message created in step 4, which is input to a speech synthesis engine (e.g., gTTS) to generate voice data and play it through the speaker. The output is a voice notification.

[0348] Step 6:

[0349] When a user wants to identify a product, they take a picture of the product with the camera on their device. The input is the video data from the camera, which is captured and saved as image data. The output is the saved image data.

[0350] Step 7:

[0351] The device sends the captured product image to the server. The input is the image data obtained in step 6, which is uploaded to the server via the Internet. The output is the image data received by the server.

[0352] Step 8:

[0353] The server analyzes the product image and identifies the product. The input is the image data obtained in step 7, and the product is identified using an image recognition algorithm (e.g., image recognition API). The output is the identified product data (e.g., product name, price).

[0354] Step 9:

[0355] The server acquires detailed information based on the identified product data. The input is the product data acquired in step 8, and detailed information is collected from the Internet, etc. using an information acquisition means. The output is detailed product information.

[0356] Step 10:

[0357] The server sends the product information to the terminal, which then notifies the user by voice. The input is the detailed product information obtained in step 9, which is input into a speech synthesis engine (e.g., gTTS) to generate voice data and play it back through the speaker. The output is a voice notification of the product information.

[0358] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0359] The support system for visually impaired people of the present invention has multiple functions for solving various problems that users face in their daily lives. This system incorporates a walking navigation function, a facility mapping function for visually impaired people, an image recognition function, a registered user recognition function, and an emotion engine that recognizes the user's emotions. An embodiment of the present invention will be described in detail below.

[0360] Walking navigation function

[0361] 1. Image capture

[0362] The device captures images of the user's surroundings with its camera at regular intervals (e.g., every second).

[0363] The captured images are sent to a server in real time.

[0364] 2. Image Analysis

[0365] The server decodes the received image and runs an image recognition algorithm on it.

[0366] The server detects obstacles in the image, the distance to the obstacles, the signal status, and the presence or absence of walls.

[0367] 3. Information Generation and Notification

[0368] Based on the image analysis results, the server generates information to be notified to the user (e.g., "There is an obstacle 2 meters ahead").

[0369] The server sends the generated information to the terminal, which notifies the user using a voice synthesis engine.

[0370] Example: While the user is walking, the device announces, "There is an obstacle 2 meters ahead."

[0371] Mapping function for facilities accessible to the visually impaired

[0372] 1. Identifying your current location

[0373] The device periodically scans the signal strength of Wi-Fi access points in its current location.

[0374] The scan results are sent to the server.

[0375] 2. Location analysis and route guidance

[0376] The server compares the user's current location with a map database based on the signal from the wireless LAN access point.

[0377] The server determines the user's current location and calculates a route to the destination.

[0378] The server transmits route information to the terminal, and the terminal uses a voice synthesis engine to provide route guidance to the user.

[0379] Example: When a user enters the voice command "I want to go to the library," the device will provide voice guidance such as "Turn left and go 20 meters, then turn right."

[0380] Image Recognition Function

[0381] 1. Image capture

[0382] The user takes a photo of the product they are holding with the device's camera.

[0383] The captured image is sent to a server.

[0384] 2. Product Identification and Information Notification

[0385] The server uses an image recognition algorithm to identify the product in the image.

[0386] Based on the identification results, the server obtains detailed product information (name, price, description, etc.) from the Internet.

[0387] The server sends the acquired detailed product information to the terminal, which then notifies the user using a voice synthesis engine.

[0388] Example: When a user picks up an item and takes a picture of it with the camera, the device announces in a voice message, "This is the 347-page novel 'The Blue Bird' and costs 1,200 yen."

[0389] Registered user recognition function

[0390] 1. Audio capture

[0391] The device captures surrounding sounds using a microphone.

[0392] The captured audio data is sent to a server.

[0393] 2. Voice analysis and identity notification

[0394] The server uses a voice recognition algorithm to analyze the voice data and identify a specific person.

[0395] The server transmits the identification result to the terminal, and the terminal notifies the user using a voice synthesis engine.

[0396] Example: When a person in front of the user speaks, the device will announce in voice, "This is Tanaka-san."

[0397] Emotion Engine

[0398] 1. Voice and facial expression capture

[0399] The device captures the user's voice and facial expressions using sensors and microphones.

[0400] The captured data is sent to a server.

[0401] 2. Emotion analysis

[0402] The server uses an emotion engine to parse the user's emotional state from the captured data.

[0403] The analysis results are classified as emotional states such as "joy," "sadness," and "stress."

[0404] 3. Emotion-based feedback

[0405] The server generates notification information based on the analyzed emotional state (e.g., "Walk a little more slowly").

[0406] The server sends the generated information to the terminal, which notifies the user using a voice synthesis engine.

[0407] Example: If the device determines that the user is feeling stressed, it will provide voice advice such as "Relax and take a deep breath."

[0408] By integrating the above functions, the support system for visually impaired people of the present invention can provide support for efficiently and effectively resolving the various problems that visually impaired people encounter in their daily lives. In addition, by incorporating an emotion engine, it can provide support that is adapted to the user's emotional state, allowing visually impaired people to live their lives with greater peace of mind.

[0409] The processing flow will be explained below.

[0410] Emotion Engine

[0411] Step 1:

[0412] The device captures the user's voice and facial expressions at regular intervals using sensors and microphones. The sampling period for audio and video data is set appropriately (e.g., every second).

[0413] Step 2:

[0414] The captured voice and facial expression data is transmitted to a server in real time.

[0415] Step 3:

[0416] The server decodes the received voice and facial expression data and runs emotion recognition algorithms.

[0417] Step 4:

[0418] Using an emotion recognition algorithm, the server analyzes the user's emotional state (happiness, sadness, anger, stress, etc.) and assigns a score to each emotion.

[0419] Step 5:

[0420] The server generates a feedback message that corresponds to the emotional state (e.g., "Take a break" or "Relax and take a deep breath").

[0421] Step 6:

[0422] The server sends the generated feedback message to the terminal.

[0423] Step 7:

[0424] The terminal passes the received feedback message to a speech synthesis engine to generate speech data.

[0425] Step 8:

[0426] The audio data generated by the terminal is notified to the user through the speaker.

[0427] Walking navigation function

[0428] Step 1:

[0429] The device captures images of the user's surroundings at regular intervals using a camera.

[0430] Step 2:

[0431] The captured images are compressed and sent to a server in real time.

[0432] Step 3:

[0433] The server decodes the received image and runs an image recognition algorithm on it.

[0434] Step 4:

[0435] The server detects obstacles in the image, the distance to the obstacles, the signal status, and the presence or absence of walls.

[0436] Step 5:

[0437] The server generates the information to be notified based on the results of image analysis (e.g., "There is an obstacle 2 meters ahead").

[0438] Step 6:

[0439] The server transmits the generated information to the terminal.

[0440] Step 7:

[0441] The information received by the terminal is passed to a speech synthesis engine to generate voice data.

[0442] Step 8:

[0443] The audio data generated by the terminal is notified to the user through the speaker.

[0444] Mapping function for facilities accessible to the visually impaired

[0445] Step 1:

[0446] The device scans the signal strength of wireless LAN access points.

[0447] Step 2:

[0448] The scan results are sent to the server.

[0449] Step 3:

[0450] The server identifies the user's current location based on the signal from the wireless LAN access point and compares it with a map database.

[0451] Step 4:

[0452] The server determines the user's current location.

[0453] Step 5:

[0454] The server calculates the route to the destination based on the user's instructions (voice commands).

[0455] Step 6:

[0456] The server sends the route information to the terminal.

[0457] Step 7:

[0458] The route information received by the terminal is passed to a voice synthesis engine to generate voice data.

[0459] Step 8:

[0460] The voice guidance generated by the terminal is notified to the user through the speaker.

[0461] Image Recognition Function

[0462] Step 1:

[0463] The user takes a photo of the product they are holding with the device's camera.

[0464] Step 2:

[0465] The captured image is sent to a server.

[0466] Step 3:

[0467] The server uses an image recognition algorithm to identify the product in the image.

[0468] Step 4:

[0469] Based on the product name identified by the server, detailed product information (such as name, price, and description) is obtained from the Internet.

[0470] Step 5:

[0471] The server organizes the detailed product information it has acquired and generates text data.

[0472] Step 6:

[0473] The server transmits the generated text data to the terminal.

[0474] Step 7:

[0475] The information received by the terminal is passed to a speech synthesis engine to generate voice data.

[0476] Step 8:

[0477] The audio data generated by the terminal is notified to the user through the speaker.

[0478] Registered user recognition function

[0479] Step 1:

[0480] The device captures surrounding sounds using a microphone.

[0481] Step 2:

[0482] The captured audio data is sent to a server.

[0483] Step 3:

[0484] The server analyzes the voice data using a voice recognition algorithm.

[0485] Step 4:

[0486] The analysis results are compared with a pre-registered voice database, and the server identifies the person.

[0487] Step 5:

[0488] The server generates the identified information as text data (e.g., "This is Mr. Tanaka").

[0489] Step 6:

[0490] The server transmits the generated text data to the terminal.

[0491] Step 7:

[0492] The information received by the terminal is passed to a speech synthesis engine to generate voice data.

[0493] Step 8:

[0494] The audio data generated by the terminal is notified to the user through the speaker.

[0495] Example 2

[0496] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0497] There is a need for a system that comprehensively resolves the multiple obstacles and difficulties that visually impaired people face in their daily lives. Conventional systems only individually address a wide range of requirements, such as obstacle detection, location identification, person identification, product information acquisition, and even user emotion analysis, and do not provide a unified solution. Furthermore, it is difficult to provide appropriate feedback based on the user's real-time situation and emotions, which limits the effectiveness of these systems in improving the safety and comfort of visually impaired people.

[0498] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0499] In this invention, the server includes an imaging means for capturing images of the surroundings, an image analysis means for analyzing the images and detecting obstacles, a notification means for providing audio notification based on the detected information, a location identification means for periodically scanning wireless signals to identify the current location, a route guidance means for analyzing voice commands and providing route guidance, a product identification means for capturing images of products and obtaining identification information, a voice analysis means for capturing surrounding voices and identifying specific people, and an emotion analysis means for analyzing the user's emotional state and providing feedback. This makes it possible to provide a system integrating multiple functions to centrally solve various problems in the daily lives of visually impaired people and significantly improve safety and comfort.

[0500] An "imaging means" is a device that captures an image of the surroundings.

[0501] "Image analysis means" is a technology for analyzing captured images and detecting obstacles and their positions.

[0502] The "notification means" is a technology for providing a voice notification based on the analyzed information.

[0503] "Location determination means" is a technology that periodically scans for radio signals to determine the current location.

[0504] "Route guidance means" is a technology for analyzing voice commands and providing guidance on the route to a destination.

[0505] The "product identification means" is a technology for analyzing the captured image of a product and obtaining its identification information.

[0506] "Audio analysis means" is a technology for capturing surrounding sounds and identifying specific people.

[0507] "Emotion analysis means" is a technology for analyzing the user's emotional state and providing feedback based on that state.

[0508] The visually impaired person support system of the present invention is mainly composed of a server, a terminal, and a user. This system is designed to provide multiple integrated functions to enable visually impaired people to spend their daily lives comfortably and safely. Specific embodiments of the present invention will be described in detail below.

[0509] Device Features

[0510] The device has the function of capturing images of the user's surroundings at regular intervals (e.g., every second) using a camera while the user is walking. The hardware used is a smartphone camera, and the software is a camera app. The captured image data is sent to a server in real time. The device also has the function of periodically scanning the signal strength of wireless LAN access points and sending the scan results to the server. It also has the function of capturing images of products held by the user and sending them to the server. As an audio recording function, it also captures surrounding sounds using a microphone.

[0511] Server Features

[0512] The server decodes the received image data and runs an image recognition algorithm. The software used is TensorFlow. The server detects obstacles in the image, the distance to the obstacles, the signal status, and the presence or absence of walls. It also compares the user's current location with a map database based on the signal from the Wi-Fi access point and calculates a route to the destination. The software used is Google Maps API. The server also uses an image recognition algorithm to identify products and retrieve detailed product information (such as name, price, and description) from the Internet. It also has the ability to analyze voice data using a voice recognition algorithm to identify specific people. Finally, an emotion engine is used to analyze the user's emotional state from the captured data. The software used is Affectiva.

[0513] User Notification

[0514] The server generates the notification information (e.g., "There is an obstacle 2 meters ahead") based on the analysis results and sends it to the device. The device then notifies the user using a speech synthesis engine. The software used is Google Text-to-Speech.

[0515] Specific examples

[0516] For example, if a user's smartphone camera captures images in real time while walking and the server detects an obstacle, the device will announce, "There is an obstacle two meters ahead." If the user enters the voice command, "I want to go to the library," the device will provide voice guidance, saying, "Turn left, walk 20 meters, then turn right." If the user picks up an item and takes a picture of it with the camera, the device will announce, "This is a 347-page novel, priced at 1,200 yen." If a person in front of the user speaks, the device will announce, "This is Mr. Tanaka." Finally, if the device analyzes that the user is feeling stressed, it will provide voice advice, saying, "Relax and take a deep breath."

[0517] Example prompts to input to the generative AI model

[0518] 1. Prompt to explain specific examples of walking navigation features:

[0519] "Regarding navigation features for safe walking for the visually impaired, please explain how the system notifies users of obstacles in real time as they walk."

[0520] 2. Prompt to explain the visually impaired accessibility mapping feature:

[0521] "Please explain with specific examples the voice guidance function that allows users to find directions to their destination."

[0522] 3. Prompt to explain specific examples of image recognition features:

[0523] "Please explain how you can provide audio information about an item when the user takes a photo of it."

[0524] 4. Prompt to explain specific examples of subscriber recognition features:

[0525] "Please provide a concrete example of a feature that uses voice to identify and notify the user of specific people in their vicinity."

[0526] 5. Prompt to explain a specific example of an emotion engine:

[0527] "Please give a concrete example of a system that analyzes a user's emotional state and provides appropriate feedback."

[0528] The support system for the visually impaired of the present invention integrates multiple functions in this way, aiming to provide a unified solution to the various problems that visually impaired people encounter in their daily lives. In addition, by utilizing an emotion engine, it is possible to provide appropriate support according to the user's psychological state, improving the quality of life.

[0529] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0530] Processing steps of walking navigation function

[0531] Step 1: Capture an image

[0532] Input: Surrounding image (real-time)

[0533] How it works: The device uses the smartphone camera to capture an image of the user's surroundings every second.

[0534] Data processing: Capture as image data

[0535] Output: Captured image data

[0536] Step 2: Sending images

[0537] Input: Captured image data

[0538] How it works: The device sends captured image data to the server in real time.

[0539] Data processing: data encoding and transmission processing

[0540] Output: Image data sent to the server

[0541] Step 3: Image analysis

[0542] Input: Received image data

[0543] How it works: The server decodes the received image data and uses TensorFlow to run image recognition algorithms to detect obstacles in the image, distance, signal status, and the presence or absence of walls.

[0544] Data calculation: Analysis using image recognition algorithms

[0545] Output: Obstacle detection results and analysis data

[0546] Step 4: Generate information

[0547] Input: Image analysis results

[0548] Operation: Based on the analysis results, the server generates information to notify the user (e.g., "There is an obstacle 2 meters ahead").

[0549] Data calculation: Transformation of analysis results and generation of notification messages

[0550] Output: Information message

[0551] Step 5: Notification

[0552] Input: Notification message

[0553] What it does: The server generates a notification message and sends it to the device, which then uses Google Text-to-Speech to notify the user audibly.

[0554] Data operations: message conversion and notification processing

[0555] Output: Synthesized notification

[0556] Blind Accessible Facilities Mapping Feature Processing Steps

[0557] Step 1: Locate your location

[0558] Input: Wi-Fi access point signal strength

[0559] How it works: Your device periodically scans the signal strength of Wi-Fi access points.

[0560] Data Processing: Signal Strength Data Collection

[0561] Output: Scan result data

[0562] Step 2: Submit scan results

[0563] Input: Scan result data

[0564] Operation: The device sends the scan result data to the server.

[0565] Data processing: data encoding and transmission processing

[0566] Output: Scan result data sent to the server

[0567] Step 3: Location analysis

[0568] Input: Scan result data

[0569] How it works: The server compares the scanned data with a map database to determine the user's current location. It uses the Google Maps API.

[0570] Data calculation: Matching signal strength data with map data, and determining location

[0571] Output: Current location information

[0572] Step 4: Route calculation

[0573] Input: Current location

[0574] How it works: The server calculates the route to the destination based on the current location information.

[0575] Data calculation: Execution of route calculation algorithms

[0576] Output: Route information

[0577] Step 5: Directions

[0578] Input: Route information

[0579] How it works: The server sends route information to the device, which then uses Google Text-to-Speech to provide voice directions.

[0580] Data calculation: Route information voice conversion and notification processing

[0581] Output: Speech-synthesized route directions

[0582] Image Recognition Function Processing Steps

[0583] Step 1: Capture an image

[0584] Input: Item held in hand (real-time)

[0585] How it works: A user takes a picture of a product with their smartphone camera.

[0586] Data processing: Capture as image data

[0587] Output: Captured image data

[0588] Step 2: Sending images

[0589] Input: Captured image data

[0590] Action: The device sends the captured image data to the server.

[0591] Data processing: data encoding and transmission processing

[0592] Output: Image data sent to the server

[0593] Step 3: Product Identification

[0594] Input: Received image data

[0595] How it works: The server uses image recognition algorithms to identify the product in the image.

[0596] Data calculation: Running image recognition algorithms

[0597] Output: Product identification results

[0598] Step 4: Information Acquisition

[0599] Input: Product identification result

[0600] Operation: Based on the identification results, the server retrieves detailed product information (name, price, description, etc.) from the Internet.

[0601] Data Calculation: Search and Extraction of Details

[0602] Output: Product details

[0603] Step 5: Information Notification

[0604] Input: Product details

[0605] How it works: The server sends product details to the device, which then uses Google Text-to-Speech to notify the user aloud.

[0606] Data Computing: Information Speech Conversion and Notification Processing

[0607] Output: Speech-synthesized product information

[0608] Process steps for subscriber recognition

[0609] Step 1: Capture audio

[0610] Input: Ambient audio (real-time)

[0611] How it works: Your device captures ambient sound with its microphone.

[0612] Data processing: Capture as audio data

[0613] Output: Captured audio data

[0614] Step 2: Sending audio

[0615] Input: Captured audio data

[0616] What it does: The device sends the captured audio data to the server.

[0617] Data processing: data encoding and transmission processing

[0618] Output: Audio data sent to the server

[0619] Step 3: Voice analysis and person identification

[0620] Input: Received audio data

[0621] How it works: The server analyzes the audio data using a speech recognition algorithm to identify a specific person. The software used is Google Speech-to-Text.

[0622] Data Computing: Speech Recognition Algorithms and Person Identification

[0623] Output: Person identification result

[0624] Step 4: Notification of identification results

[0625] Input: Person identification result

[0626] How it works: The server sends the identification results to the device, which then notifies the user using Google Text-to-Speech.

[0627] Data processing: speech conversion and notification of the recognition results

[0628] Output: Synthesized voice notification of person identification

[0629] Emotion Engine Processing Steps

[0630] Step 1: Capture your voice and facial expressions

[0631] Input: User's voice and facial expressions (real-time)

[0632] How it works: The device captures the user's voice and facial expressions using sensors and microphones.

[0633] Data processing: Capture as voice and facial expression data

[0634] Output: Captured audio and facial expression data

[0635] Step 2: Sending data

[0636] Input: Captured voice and facial expression data

[0637] Action: The device sends the captured data to the server.

[0638] Data processing: data encoding and transmission processing

[0639] Output: Voice and facial expression data sent to the server

[0640] Step 3: Sentiment Analysis

[0641] Input: Received voice and facial expression data

[0642] How it works: The server uses an emotion analysis algorithm to analyze the user's emotional state from the captured data. The software used is Affectiva.

[0643] Data Computing: Running sentiment analysis algorithms and classifying emotional states

[0644] Output: Emotion analysis results

[0645] Step 4: Emotion-based feedback

[0646] Input: Sentiment analysis results

[0647] How it works: The server generates notification information (e.g., "Walk a little slower") based on the analyzed emotional state.

[0648] Data calculation: Transformation of analysis results and generation of notification messages

[0649] Output: Information message

[0650] Step 5: Notification

[0651] Input: Notification message

[0652] What it does: The server generates a notification message and sends it to the device, which then uses Google Text-to-Speech to notify the user audibly.

[0653] Data operations: message conversion and notification processing

[0654] Output: Synthesized notification

[0655] Through the above processing steps, the visually impaired person support system of the present invention provides the user with a safe and comfortable life.

[0656] (Application example 2)

[0657] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0658] Visually impaired people face many challenges in their daily lives. In particular, shopping and traveling in brick-and-mortar stores are often extremely inconvenient and dangerous for them. They also have difficulty recognizing obstacles, identifying specific people, and obtaining product information, which limits their independent lifestyles. Effective support systems that can solve these problems and enable visually impaired people to live their daily lives with peace of mind are needed.

[0659] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an imaging means for capturing images of the surroundings, an image analysis means for analyzing the images captured by the imaging means and detecting obstacles, a notification means for providing a voice notification based on information detected by the image analysis means, a registered person recognition means for identifying registered people and providing a voice notification based on the identification result, and an emotion analysis and notification means for analyzing the user's emotional state and providing feedback based on the emotional state. This allows visually impaired people to move safely while avoiding obstacles, to identify specific people to communicate smoothly, and to receive support according to their emotional state.

[0660] 1. "Imaging means" means a device for capturing images of the surroundings, including a camera or sensor.

[0661] 2. "Image Analysis Means" means the algorithms or processes that analyze captured images to detect obstacles and extract other information.

[0662] 3. "Notification means" refers to an interface for notifying the user of necessary information by voice or other means.

[0663] 4. "Walking navigation means" is a function that provides real-time voice navigation to obstacles.

[0664] 5. "Means for recognizing registered persons" refers to a function that identifies registered persons and sends notifications based on the results of that identification.

[0665] 6. "Emotion analysis and notification means" means a means for analyzing the user's emotional state and providing feedback according to that state in the form of voice or other means.

[0666] 7. "Location identification means" is a function in which the image analysis means analyzes location information and identifies the user's current location.

[0667] 8. "Store guidance means" is a function that provides audio guidance to the user about each area within a physical store.

[0668] This invention provides a support system that enables visually impaired people to navigate and shop safely in brick-and-mortar stores. The system has the following main functions:

[0669] System configuration

[0670] 1. Imaging Method

[0671] It uses smart glasses or a smartphone camera to capture images of the surroundings.

[0672] Example: As a user walks through a physical store, a camera captures images of the user's surroundings at regular intervals.

[0673] 2. Image analysis methods

[0674] The captured image is sent to the server, where it is analyzed in real time.

[0675] The software used includes image analysis algorithms and TensorFlow. The server recognizes obstacles and products and sends that information to the notification means.

[0676] Example: If an obstacle is detected, the analysis result will be generated as "There is an obstacle 2 meters ahead."

[0677] 3. Means of notification

[0678] Based on the analysis results, the information is notified to the user by voice. Voice synthesis uses a voice engine such as pyttsx3.

[0679] Example: A speech synthesis engine informs the user, "There is an obstacle 2 meters ahead."

[0680] 4. Walking navigation methods

[0681] Real-time voice navigation is provided while the user is moving, which is achieved by combining imaging means, image analysis means, and notification means.

[0682] Example: When a user arrives at a specific area (e.g., fresh produce), they are notified, "You have arrived at the fish section."

[0683] 5. Registered User Identification Method

[0684] It captures surrounding sounds and identifies specific people, using the microphone in smart glasses or smartphones to record the sounds.

[0685] The server performs voice analysis and notifies the user of the results of recognizing a specific person.

[0686] Example: When a person in front of the user speaks, the user is notified that "This is a store clerk."

[0687] 6. Sentiment Analysis and Notification Methods

[0688] It captures the user's voice and facial expressions and analyzes their emotional state using an emotion analysis model powered by TensorFlow.

[0689] Appropriate feedback is provided via voice based on the analysis results.

[0690] For example, if the user is analyzed as feeling stressed, they will be notified to "relax and take a deep breath."

[0691] Specific examples

[0692] Consider a scenario in which the system would work effectively when a visually impaired person visits a supermarket. The user walks through various sections of the store, and the smart glasses provide real-time information about their surroundings. For example, when the user scans an item they have picked up with the camera, they are notified by voice, "This is an apple. The price is 150 yen." When they arrive at a specific area, they are informed, "You have arrived at the fish section."

[0693] Prompt Sentence Examples

[0694] An example of a specific prompt sentence for the generative AI model is shown below.

[0695] A smart glasses application for visually impaired people in shopping malls provides walking navigation, product recognition, staff recognition, and emotional feedback. As an example, when a user scans an item they have picked up, the application asks, "How can the application recognize the product from the image captured by the camera and announce its name, price, and description by voice?"

[0696] This will enable visually impaired people to live their daily lives more safely and independently.

[0697] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0698] Step 1:

[0699] The device uses an imaging means (camera) to capture images of the surroundings. The imaging means captures images from the camera at regular intervals (e.g., every second) according to the user's viewpoint, and sends the image data to a server. The input is the image captured from the camera, and the output is the image data sent to the server.

[0700] Step 2:

[0701] The server analyzes the image data it receives. The image analysis means uses TensorFlow's image recognition algorithm to identify obstacles, products, and people in the image. The input is the image data sent to the server, and the output is the analysis results (obstacle position, product information, and person information). Specifically, the server extracts features from the image and performs analysis using a pre-trained model.

[0702] Step 3:

[0703] The server generates the information to be notified based on the analysis results. The notification means converts the analysis results (e.g., "There is an obstacle 2 meters ahead") into voice using a speech synthesis engine (pyttsx3) and notifies the user. The input is the analysis results from the image analysis means, and the output is a voice notification. This is achieved by the server converting the analysis results into text and sending it to the speech synthesis engine.

[0704] Step 4:

[0705] When the user takes a specific action (e.g., picking up a product), the image recognition function is activated. The product is photographed with the device's camera and the image is sent to the server. The input is the image of the product, and the output is the data sent to the server. The server identifies the product, obtains product information (name, price, etc.), and notifies the user by voice.

[0706] Step 5:

[0707] When the user moves, the walking navigation means operates. The terminal captures images in real time using the imaging means and sends them to the server. The server analyzes obstacle information and notifies the user of navigation information. The input is image data captured in real time, and the output is navigation information. The server processes the data in real time and generates navigation information.

[0708] Step 6:

[0709] The device's microphone captures surrounding sounds and activates the registered user recognition means. The voice data is sent to the server, which then uses a voice recognition algorithm to identify the person. The input is the captured voice data, and the output is specific person information. The server notifies the user of the identification result through a voice synthesis engine.

[0710] Step 7:

[0711] The terminal captures the user's voice and facial expressions, and the emotion analysis and notification means are activated. The server uses an emotion analysis algorithm to analyze the emotional state and provide appropriate feedback. The input is the captured voice and facial expression data, and the output is feedback information based on the emotional state. The server analyzes the emotional state and generates information to be notified.

[0712] Step 8:

[0713] When a user goes shopping in a physical store, the store guidance means is activated. The server identifies the user's current location, calculates the route to the destination, and provides voice guidance. The input is data obtained by scanning the signal strength of wireless LAN access points, and the output is route guidance information. The server sends the route guidance information to a speech synthesis engine, which notifies the user by voice.

[0714] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0715] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0716] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0717] [Second embodiment]

[0718] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0719] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0720] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0721] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0722] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0723] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0724] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0725] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0726] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0727] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0728] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0729] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0730] The visually impaired support system of the present invention has a plurality of functions for solving various problems that users face in their daily lives. The embodiments of the present invention will be described in detail below.

[0731] Walking navigation function

[0732] 1. Image capture

[0733] The device captures images of the user's surroundings with its camera at regular intervals (e.g., every second).

[0734] The captured images are sent to a server in real time.

[0735] 2. Image Analysis

[0736] The server decodes the received image and runs an image recognition algorithm on it.

[0737] The server detects obstacles in the image, the distance to the obstacles, the signal status, and the presence or absence of walls.

[0738] 3. Information Generation and Notification

[0739] Based on the image analysis results, the server generates information to be notified to the user (e.g., "There is an obstacle 2 meters ahead").

[0740] The server sends the generated information to the terminal, which notifies the user using a voice synthesis engine.

[0741] Example: While the user is walking, the device announces, "There is an obstacle 2 meters ahead."

[0742] Mapping function for facilities accessible to the visually impaired

[0743] 1. Identifying your current location

[0744] The device periodically scans the signal strength of Wi-Fi access points in its current location.

[0745] The scan results are sent to the server.

[0746] 2. Location analysis and route guidance

[0747] The server compares the user's current location with a map database based on the signal from the wireless LAN access point.

[0748] The server determines the user's current location and calculates a route to the destination.

[0749] The server transmits route information to the terminal, and the terminal uses a voice synthesis engine to provide route guidance to the user.

[0750] Example: When a user enters the voice command "I want to go to the library," the device will provide voice guidance such as "Turn left and go 20 meters, then turn right."

[0751] Image Recognition Function

[0752] 1. Image capture

[0753] The user takes a photo of the product they are holding with the device's camera.

[0754] The captured image is sent to a server.

[0755] 2. Product Identification and Information Notification

[0756] The server uses an image recognition algorithm to identify the product in the image.

[0757] Based on the identification results, the server obtains detailed product information (name, price, description, etc.) from the Internet.

[0758] The server sends the acquired detailed product information to the terminal, which then notifies the user using a voice synthesis engine.

[0759] Example: When a user picks up an item and takes a picture of it with the camera, the device announces in a voice message, "This is the 347-page novel 'The Blue Bird' and costs 1,200 yen."

[0760] Registered user recognition function

[0761] 1. Audio capture

[0762] The device captures surrounding sounds using a microphone.

[0763] The captured audio data is sent to a server.

[0764] 2. Voice analysis and identity notification

[0765] The server uses a voice recognition algorithm to analyze the voice data and identify a specific person.

[0766] The server transmits the identification result to the terminal, and the terminal notifies the user using a voice synthesis engine.

[0767] Example: When a person in front of the user speaks, the device will announce in voice, "This is Tanaka-san."

[0768] The objective of the support system for visually impaired people of the present invention is to integrate these functions to efficiently and effectively solve the various problems that visually impaired people encounter in their daily lives, thereby enabling them to live safer and more independent lives.

[0769] The processing flow will be explained below.

[0770] Walking navigation function

[0771] Step 1:

[0772] The device captures images of the user's surroundings at regular intervals using a camera.

[0773] Step 2:

[0774] The device compresses the captured images and sends them to the server in real time.

[0775] Step 3:

[0776] The server decodes the received image and runs an image recognition algorithm on it.

[0777] Step 4:

[0778] The server detects obstacles in the image, the distance to the obstacles, the signal status, and the presence or absence of walls.

[0779] Step 5:

[0780] The server generates the information to be notified based on the results of image analysis (e.g., "There is an obstacle 2 meters ahead").

[0781] Step 6:

[0782] The server transmits the generated information to the terminal.

[0783] Step 7:

[0784] The information received by the terminal is passed to a speech synthesis engine to generate voice data.

[0785] Step 8:

[0786] The audio data generated by the terminal is notified to the user through the speaker.

[0787] Mapping function for facilities accessible to the visually impaired

[0788] Step 1:

[0789] The device scans the signal strength of wireless LAN access points.

[0790] Step 2:

[0791] The device sends the scan results to the server.

[0792] Step 3:

[0793] The server identifies the user's current location based on the signal from the wireless LAN access point by comparing it with a location information database.

[0794] Step 4:

[0795] The server determines the user's current location.

[0796] Step 5:

[0797] The server calculates the route to the destination based on the user's instructions (voice commands).

[0798] Step 6:

[0799] The server sends the route information to the terminal.

[0800] Step 7:

[0801] The route information received by the terminal is passed to a voice synthesis engine to generate voice data.

[0802] Step 8:

[0803] The voice guidance generated by the terminal is notified to the user through the speaker.

[0804] Image Recognition Function

[0805] Step 1:

[0806] The user takes a photo of the product they are holding with the device's camera.

[0807] Step 2:

[0808] The device sends the captured image to the server.

[0809] Step 3:

[0810] The server uses an image recognition algorithm to identify the product in the image.

[0811] Step 4:

[0812] Based on the product name identified by the server, detailed product information (such as name, price, and description) is obtained from the Internet.

[0813] Step 5:

[0814] The server organizes the detailed product information it has acquired and generates text data.

[0815] Step 6:

[0816] The server transmits the generated text data to the terminal.

[0817] Step 7:

[0818] The information received by the terminal is passed to a speech synthesis engine to generate voice data.

[0819] Step 8:

[0820] The audio data generated by the terminal is notified to the user through the speaker.

[0821] Registered user recognition function

[0822] Step 1:

[0823] The device captures surrounding sounds using a microphone.

[0824] Step 2:

[0825] The device sends the captured audio data to the server.

[0826] Step 3:

[0827] The server analyzes the voice data using a voice recognition algorithm.

[0828] Step 4:

[0829] The server compares the analysis results with a pre-registered voice database to identify the person.

[0830] Step 5:

[0831] The server generates the identified information as text data (e.g., "This is Mr. Tanaka").

[0832] Step 6:

[0833] The server transmits the generated text data to the terminal.

[0834] Step 7:

[0835] The information received by the terminal is passed to a speech synthesis engine to generate voice data.

[0836] Step 8:

[0837] The audio data generated by the terminal is notified to the user through the speaker.

[0838] Example 1

[0839] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0840] To solve the various challenges that visually impaired people face in their daily lives, it is necessary to properly obtain information about their surroundings and convey that information to them quickly and accurately. However, current systems lack comprehensive support that goes beyond simply detecting obstacles, covering things like determining their current location, identifying people, and even providing traffic light status and specific route guidance. This leaves visually impaired people at great risk.

[0841] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0842] In this invention, the server includes an imaging means for capturing an image of the surroundings, an image analysis means for analyzing the image captured by the imaging means and detecting obstacles, traffic light status, and the presence or absence of walls, an information generation means for generating notification information based on the information detected by the image analysis means, a notification means for notifying by voice the notification information generated by the information generation means, a location identification means for analyzing location information using a wireless LAN signal and identifying the current location, an audio recording means for capturing surrounding sounds, an audio analysis means for analyzing the sounds captured by the audio recording means and identifying a specific person, an information generation means for generating notification information based on the person information identified by the audio analysis means, and a notification means for notifying by voice the notification information generated by the information generation means.This enables visually impaired people to avoid obstacles while walking, identify their current location, and easily identify people, enabling them to live safe and independent lives.

[0843] An "imaging means" is a device or sensor for capturing an image of the surroundings.

[0844] "Image analysis means" refers to a set of algorithms and software for analyzing captured images and detecting obstacles, traffic light conditions, the presence or absence of walls, etc.

[0845] The "notification means" is a device or software that conveys the analyzed information to the user by voice.

[0846] The "information generation means" refers to logic or algorithms for generating information to be notified to the user based on the analysis results.

[0847] The "location determination means" is a device or software for determining the user's current location using a wireless LAN signal.

[0848] "Audio recording means" refers to a microphone or other audio recording device for capturing ambient audio.

[0849] "Audio analysis means" refers to algorithms or software that analyzes captured audio and identifies specific individuals.

[0850] The visually impaired support system of the present invention has a plurality of functions for solving various problems that users face in their daily lives. The embodiments of the present invention will be described in detail below.

[0851] Walking navigation function

[0852] Image capture

[0853] The device captures images of the user's surroundings at regular intervals (e.g., every second) using the built-in camera of a smartphone.

[0854] Sending images

[0855] The device immediately sends the captured image to the server using Wi-Fi or mobile data.

[0856] Image analysis

[0857] The server decodes the received images and uses image recognition algorithms such as TensorFlow and OpenCV to detect obstacles, traffic light status, and the presence or absence of walls.

[0858] Information Generation and Notification

[0859] Based on the analysis results, the server generates notification information such as "There is an obstacle 2 meters ahead." The information is sent from the server to the device, which then notifies the user using a voice synthesis engine such as Google Text-to-Speech.

[0860] Specific examples

[0861] While the user is walking, the device will announce to them by voice, "There is an obstacle 2 meters ahead."

[0862] Mapping function for facilities accessible to the visually impaired

[0863] Scan your current location

[0864] The device periodically scans the signal strength of surrounding Wi-Fi access points and sends the results of the scan to the server.

[0865] Location analysis and routing

[0866] The server uses signal data from wireless LAN access points to identify the user's current location using a map database such as Google Maps API and calculates the optimal route. The route information is sent to the device, which then uses a voice synthesis engine such as Google Text-to-Speech to provide route guidance to the user.

[0867] Specific examples

[0868] When a user inputs the voice command "I want to go to the library," the device will provide voice guidance such as "Turn left and go 20 meters, then turn right."

[0869] Image Recognition Function

[0870] Image capture

[0871] The user takes a photo of the product they are holding with the device's camera, and the captured image is sent to the server.

[0872] Product identification and information acquisition

[0873] The server identifies the product using an image recognition algorithm such as YOLOv3, and then uses the Amazon API or Rakuten API to obtain detailed product information (name, price, description, etc.).

[0874] Information Notification

[0875] The acquired product details are sent from the server to the terminal, which then notifies the user using a voice synthesis engine such as Google Text-to-Speech.

[0876] Specific examples

[0877] When the user picks up the product and takes a picture of it with the camera, the device announces in voice, "This is the 347-page novel 'The Blue Bird' and costs 1,200 yen."

[0878] Registered user recognition function

[0879] Audio capture

[0880] The device captures the surrounding sound with a microphone, and the captured sound data is sent to the server.

[0881] Audio analysis and notifications

[0882] The server analyzes the voice data using a speech recognition algorithm such as IBM Watson Speech to Text to identify a specific person, and the identification result is sent from the server to the device, which then notifies the user using a speech synthesis engine such as Google Text-to-Speech.

[0883] Specific examples

[0884] When a person in front of the user speaks, the device will announce aloud, "This is Tanaka-san."

[0885] Examples of prompt statements

[0886] 1. Walking navigation function

[0887] Prompt: "What is the process for the system that notifies you of nearby obstacles?"

[0888] 2. Mapping function for facilities accessible to the visually impaired

[0889] Prompt: "Please explain the process of the system that identifies your current location and provides directions to your destination."

[0890] 3. Image Recognition

[0891] Prompt: "Please tell me the process of the system that takes a picture of a product with a camera and provides detailed information about it."

[0892] 4. Registered User Recognition Function

[0893] Prompt: "Please explain the process for a system that analyzes surrounding sounds and identifies a specific person."

[0894] The above is an embodiment of the visually impaired support system of the present invention, and this system enables visually impaired people to live safer and more independent lives.

[0895] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0896] Walking navigation function

[0897] Step 1:

[0898] Image capture

[0899] The device captures images of the user's surroundings with a camera at regular intervals (every second). The input is image data obtained from the camera, and the output is the captured image file.

[0900] Specific operation: The device automatically launches the camera app and simulates pressing the shutter button every second.

[0901] Step 2:

[0902] Sending images

[0903] The image files captured by the device are sent to the server in real time. The input is the image file generated in step 1, and the output is the upload of image data to the server.

[0904] Specific operation: The device compresses the image data and uploads it to the server using the HTTPS protocol.

[0905] Step 3:

[0906] Image analysis

[0907] The server decodes the received image and analyzes it using an image recognition algorithm using TensorFlow, OpenCV, etc. The input is the image data sent to the server, and the output is the detection results such as obstacles, traffic light status, and the presence or absence of walls.

[0908] Specific operation: The server processes the received image, runs a recognition algorithm, and generates an analysis result.

[0909] Step 4:

[0910] information generation

[0911] The server generates information to be notified to the user based on the analysis results. The input is the analysis results obtained in step 3, and the output is a notification message.

[0912] What happens: The server uses a text generation algorithm to generate a message such as "There is an obstacle 2 meters ahead."

[0913] Step 5:

[0914] Information Notification

[0915] The server sends the generated notification message to the device, and the device uses a speech synthesis engine such as Google Text-to-Speech to notify the user audibly. The input is the generated notification message, and the output is the audio notification delivered to the user.

[0916] Specific operation: The device converts the received message into voice and notifies the user through the speaker, "There is an obstacle 2 meters ahead."

[0917] Mapping function for facilities accessible to the visually impaired

[0918] Step 1:

[0919] Scan your current location

[0920] The device periodically scans the signal strength of surrounding Wi-Fi access points. The input is the Wi-Fi signal captured by the Wi-Fi module, and the output is the signal strength data.

[0921] Specific operation: The device activates the Wi-Fi module and collects the surrounding network signal strength.

[0922] Step 2:

[0923] Sending scan results

[0924] The terminal sends the collected signal strength data to the server. The input is the signal strength data collected in step 1, and the output is the upload to the server.

[0925] Specific operation: The terminal compresses the signal data and uploads it to the server via HTTPS protocol.

[0926] Step 3:

[0927] location analysis

[0928] The server analyzes the location information based on the signal data of the wireless LAN access point and determines the current location. The input is the transmitted signal strength data, and the output is the determined current location information.

[0929] What it does: The server analyzes the signal data and locates it by comparing it with a map database (e.g., Google Maps API).

[0930] Step 4:

[0931] Route calculation

[0932] The server calculates the optimal route from the specified current location to the destination. The input is the current location information and the destination information, and the output is the route information.

[0933] Specific operation: The server executes a route calculation algorithm (e.g., A search algorithm) to calculate the route to the destination.

[0934] Step 5:

[0935] Route guidance

[0936] The server sends route information to the device, which then uses a speech synthesis engine such as Google Text-to-Speech to provide route guidance to the user. The input is the calculated route information, and the output is voice guidance provided to the user.

[0937] Specific operation: The device converts the route information it receives into voice and notifies the user, "Turn left, go 20 meters, then turn right."

[0938] Image Recognition Function

[0939] Step 1:

[0940] Image capture

[0941] The user takes a photo of the product they are holding with the device's camera. The input is image data acquired from the camera, and the output is the captured image file.

[0942] Specific operation: The device automatically launches the camera app, and the user presses the shutter button.

[0943] Step 2:

[0944] Sending images

[0945] The device sends the captured image file to the server. The input is the captured image file, and the output is the upload of the image data to the server.

[0946] Specific operation: The device compresses the image data and uploads it to the server using the HTTPS protocol.

[0947] Step 3:

[0948] Product Identification

[0949] The server uses an image recognition algorithm (e.g., YOLOv3) to identify the product in the image. The input is the submitted image data, and the output is the identified product information.

[0950] Specific operation: The server analyzes the received image and identifies it as "This is a novel called 'The Blue Bird'."

[0951] Step 4:

[0952] Information acquisition

[0953] The server retrieves detailed product information (such as name, price, description, etc.) from the Internet. The input is the identified product information, and the output is the retrieved product detail data.

[0954] Specific operation: The server calls the API to obtain product data.

[0955] Step 5:

[0956] Information Notification

[0957] The server sends the acquired product details to the terminal, which then notifies the user using a speech synthesis engine such as Google Text-to-Speech. The input is the acquired product details data, and the output is a voice notification delivered to the user.

[0958] Specific operation: The device converts the received product information into voice and notifies the user, "This is the 347-page novel 'The Blue Bird', priced at 1,200 yen."

[0959] Registered user recognition function

[0960] Step 1:

[0961] Audio capture

[0962] The device captures the surrounding sound with a microphone. The input is the audio data collected by the microphone, and the output is the captured audio file.

[0963] Specific operation: The device activates the microphone and collects surrounding sounds.

[0964] Step 2:

[0965] Sending Audio

[0966] The device sends the captured audio file to the server. The input is the captured audio file, and the output is the audio data uploaded to the server.

[0967] Specific operation: The device compresses the audio data and uploads it to the server using the HTTPS protocol.

[0968] Step 3:

[0969] Audio analysis

[0970] The server analyzes the voice data using a speech recognition algorithm (e.g., IBM Watson Speech to Text) to identify a specific person. The input is the transmitted voice data, and the output is the identified person's information.

[0971] Specific operation: The server processes the voice data and generates an identification result such as "This is Mr. Tanaka."

[0972] Step 4:

[0973] Generating the identification results

[0974] The server generates notification information based on the identification result. The input is the identified person's information, and the output is the generated notification message.

[0975] Specific operation: The server generates a text message and sends it to the device.

[0976] Step 5:

[0977] Notification of identity

[0978] The device uses a speech synthesis engine such as Google Text-to-Speech to convert the notification message into speech and notify the user. The input is the generated notification message, and the output is the audio notification delivered to the user.

[0979] Specific operation: The device uses the speaker to announce, "This is Tanaka-san."

[0980] (Application example 1)

[0981] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0982] Due to their unique needs, it is difficult for visually impaired people to shop safely and comfortably in brick-and-mortar stores. To solve this problem, a system is needed that allows visually impaired people to independently select and navigate products in brick-and-mortar stores. In addition, since it is difficult for them to accurately grasp product information in real time, appropriate support is required.

[0983] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0984] In this invention, the server includes an imaging means for capturing an image of the surroundings, an image analysis means for analyzing the image captured by the imaging means and detecting obstacles, a notification means for providing a voice notification based on information detected by the image analysis means, an image recognition means for identifying products, an information acquisition means for acquiring detailed information about the identified products, and a notification means for providing a voice notification based on information acquired by the information acquisition means. This enables visually impaired people to receive product information by voice while safely moving around in a physical store.

[0985] An "imaging means" is a device that captures an image of the surroundings.

[0986] The "image analysis means" is a device that analyzes the image captured by the imaging means and detects objects such as obstacles.

[0987] The "notification means" is a device that notifies the user of information detected by the image analysis means and other information by voice.

[0988] An "image recognition means" is a device for identifying an item in a captured image.

[0989] The "information acquisition means" is a device that acquires detailed information (such as the name and price) of the identified product from the Internet or the like.

[0990] The "location specifying means" is a device that specifies the current location by analyzing the location information performed by the image analysis means.

[0991] An "audio recording means" is a device that captures ambient sounds.

[0992] The "voice analysis means" is a device that analyzes the voice captured by the voice recording means and identifies a specific person.

[0993] The present invention relates to a support system for visually impaired people to shop safely and comfortably in a brick-and-mortar store. Hereinafter, an embodiment of the present invention will be specifically described.

[0994] System Configuration

[0995] The system includes the following main components:

[0996] 1. An imaging means to capture images of the surroundings

[0997] 2. Image analysis means for analyzing images captured by the imaging means and detecting obstacles

[0998] 3. Notification means that notifies by voice based on information detected by image analysis means

[0999] 4. Image recognition methods for identifying products

[1000] 5. Information acquisition means for acquiring detailed information on identified products

[1001] 6. Notification method for notifying information by voice

[1002] Hardware and Software

[1003] Hardware:

[1004] Smartphone, smart glasses, or head-mounted display

[1005] Web camera (imaging means)

[1006] Microphone (audio recording means)

[1007] Speaker (notification means)

[1008] software:

[1009] OpenCV (library for image capture and analysis)

[1010] Google Text-to-Speech (gTTS) (voice notifications)

[1011] Image recognition API (product identification)

[1012] Internet connection (information acquisition)

[1013] Processing steps

[1014] When a user activates the system, the imaging means captures images of the surrounding area at regular intervals. The captured images are sent to a server in real time, and the image analysis means detects obstacles in the images and analyzes their positions. The notification means then notifies the user of the obstacle information by voice.

[1015] At the same time, if the user wants to identify a product, they take a photo of the product using a smartphone or smart glasses. The captured image is sent to a server, and the image recognition means identifies the product. Detailed information about the identified product is obtained from the Internet by an information acquisition means, and the information is notified to the user by voice via a notification means.

[1016] Specific examples

[1017] For example, when a user is walking in a shopping mall, the system will notify them that there is an obstacle two meters ahead, helping them to walk safely. Also, when a user picks up a product and takes a picture of it with the camera, the system will announce the product information in a voice message, saying, "This is a 347-page book, priced at 1,200 yen."

[1018] Prompt Sentence Examples

[1019] Example prompt for image recognition:

[1020] "Please recognize the item in this image and provide the item name and distance."

[1021] Example of a prompt for audio notification:

[1022] "2 meters ahead, there is an obstacle. It is a table."

[1023] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1024] Step 1:

[1025] The imaging means captures images of the surroundings at regular intervals (e.g., every second). The input is video data from the camera, which is captured and saved as image data. The output is the saved image data.

[1026] Step 2:

[1027] The device sends the captured image data to the server. The input is the image data obtained in step 1, which is uploaded to the server via the Internet. The output is the image data received by the server.

[1028] Step 3:

[1029] The server analyzes the image data it receives and identifies the presence and location of obstacles. The input is the image data received by the server, which is analyzed using an image analysis algorithm (e.g., OpenCV). The output is the location information of obstacles.

[1030] Step 4:

[1031] The server generates the content to be notified to the user based on the analysis results. The input is the obstacle location information obtained in step 3, and an appropriate notification message is created using the generative AI model. The output is the notification message (text data).

[1032] Step 5:

[1033] The server sends a notification message to the device, and the device notifies the user by voice. The input is the notification message created in step 4, which is input to a speech synthesis engine (e.g., gTTS) to generate voice data and play it through the speaker. The output is a voice notification.

[1034] Step 6:

[1035] When a user wants to identify a product, they take a picture of the product with the camera on their device. The input is the video data from the camera, which is captured and saved as image data. The output is the saved image data.

[1036] Step 7:

[1037] The device sends the captured product image to the server. The input is the image data obtained in step 6, which is uploaded to the server via the Internet. The output is the image data received by the server.

[1038] Step 8:

[1039] The server analyzes the product image and identifies the product. The input is the image data obtained in step 7, and the product is identified using an image recognition algorithm (e.g., image recognition API). The output is the identified product data (e.g., product name, price).

[1040] Step 9:

[1041] The server acquires detailed information based on the identified product data. The input is the product data acquired in step 8, and detailed information is collected from the Internet, etc. using an information acquisition means. The output is detailed product information.

[1042] Step 10:

[1043] The server sends the product information to the terminal, which then notifies the user by voice. The input is the detailed product information obtained in step 9, which is input into a speech synthesis engine (e.g., gTTS) to generate voice data and play it back through the speaker. The output is a voice notification of the product information.

[1044] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1045] The support system for visually impaired people of the present invention has multiple functions for solving various problems that users face in their daily lives. This system incorporates a walking navigation function, a facility mapping function for visually impaired people, an image recognition function, a registered user recognition function, and an emotion engine that recognizes the user's emotions. An embodiment of the present invention will be described in detail below.

[1046] Walking navigation function

[1047] 1. Image capture

[1048] The device captures images of the user's surroundings with its camera at regular intervals (e.g., every second).

[1049] The captured images are sent to a server in real time.

[1050] 2. Image Analysis

[1051] The server decodes the received image and runs an image recognition algorithm on it.

[1052] The server detects obstacles in the image, the distance to the obstacles, the signal status, and the presence or absence of walls.

[1053] 3. Information Generation and Notification

[1054] Based on the image analysis results, the server generates information to be notified to the user (e.g., "There is an obstacle 2 meters ahead").

[1055] The server sends the generated information to the terminal, which notifies the user using a voice synthesis engine.

[1056] Example: While the user is walking, the device announces, "There is an obstacle 2 meters ahead."

[1057] Mapping function for facilities accessible to the visually impaired

[1058] 1. Identifying your current location

[1059] The device periodically scans the signal strength of Wi-Fi access points in its current location.

[1060] The scan results are sent to the server.

[1061] 2. Location analysis and route guidance

[1062] The server compares the user's current location with a map database based on the signal from the wireless LAN access point.

[1063] The server determines the user's current location and calculates a route to the destination.

[1064] The server transmits route information to the terminal, and the terminal uses a voice synthesis engine to provide route guidance to the user.

[1065] Example: When a user enters the voice command "I want to go to the library," the device will provide voice guidance such as "Turn left and go 20 meters, then turn right."

[1066] Image Recognition Function

[1067] 1. Image capture

[1068] The user takes a photo of the product they are holding with the device's camera.

[1069] The captured image is sent to a server.

[1070] 2. Product Identification and Information Notification

[1071] The server uses an image recognition algorithm to identify the product in the image.

[1072] Based on the identification results, the server obtains detailed product information (name, price, description, etc.) from the Internet.

[1073] The server sends the acquired detailed product information to the terminal, which then notifies the user using a voice synthesis engine.

[1074] Example: When a user picks up an item and takes a picture of it with the camera, the device announces in a voice message, "This is the 347-page novel 'The Blue Bird' and costs 1,200 yen."

[1075] Registered user recognition function

[1076] 1. Audio capture

[1077] The device captures surrounding sounds using a microphone.

[1078] The captured audio data is sent to a server.

[1079] 2. Voice analysis and identity notification

[1080] The server uses a voice recognition algorithm to analyze the voice data and identify a specific person.

[1081] The server transmits the identification result to the terminal, and the terminal notifies the user using a voice synthesis engine.

[1082] Example: When a person in front of the user speaks, the device will announce in voice, "This is Tanaka-san."

[1083] Emotion Engine

[1084] 1. Voice and facial expression capture

[1085] The device captures the user's voice and facial expressions using sensors and microphones.

[1086] The captured data is sent to a server.

[1087] 2. Emotion analysis

[1088] The server uses an emotion engine to parse the user's emotional state from the captured data.

[1089] The analysis results are classified as emotional states such as "joy," "sadness," and "stress."

[1090] 3. Emotion-based feedback

[1091] The server generates notification information based on the analyzed emotional state (e.g., "Walk a little more slowly").

[1092] The server sends the generated information to the terminal, which notifies the user using a voice synthesis engine.

[1093] Example: If the device determines that the user is feeling stressed, it will provide voice advice such as "Relax and take a deep breath."

[1094] By integrating the above functions, the support system for visually impaired people of the present invention can provide support for efficiently and effectively resolving the various problems that visually impaired people encounter in their daily lives. In addition, by incorporating an emotion engine, it can provide support that is adapted to the user's emotional state, allowing visually impaired people to live their lives with greater peace of mind.

[1095] The processing flow will be explained below.

[1096] Emotion Engine

[1097] Step 1:

[1098] The device captures the user's voice and facial expressions at regular intervals using sensors and microphones. The sampling period for audio and video data is set appropriately (e.g., every second).

[1099] Step 2:

[1100] The captured voice and facial expression data is transmitted to a server in real time.

[1101] Step 3:

[1102] The server decodes the received voice and facial expression data and runs emotion recognition algorithms.

[1103] Step 4:

[1104] Using an emotion recognition algorithm, the server analyzes the user's emotional state (happiness, sadness, anger, stress, etc.) and assigns a score to each emotion.

[1105] Step 5:

[1106] The server generates a feedback message that corresponds to the emotional state (e.g., "Take a break" or "Relax and take a deep breath").

[1107] Step 6:

[1108] The server sends the generated feedback message to the terminal.

[1109] Step 7:

[1110] The terminal passes the received feedback message to a speech synthesis engine to generate speech data.

[1111] Step 8:

[1112] The audio data generated by the terminal is notified to the user through the speaker.

[1113] Walking navigation function

[1114] Step 1:

[1115] The device captures images of the user's surroundings at regular intervals using a camera.

[1116] Step 2:

[1117] The captured images are compressed and sent to a server in real time.

[1118] Step 3:

[1119] The server decodes the received image and runs an image recognition algorithm on it.

[1120] Step 4:

[1121] The server detects obstacles in the image, the distance to the obstacles, the signal status, and the presence or absence of walls.

[1122] Step 5:

[1123] The server generates the information to be notified based on the results of image analysis (e.g., "There is an obstacle 2 meters ahead").

[1124] Step 6:

[1125] The server transmits the generated information to the terminal.

[1126] Step 7:

[1127] The information received by the terminal is passed to a speech synthesis engine to generate voice data.

[1128] Step 8:

[1129] The audio data generated by the terminal is notified to the user through the speaker.

[1130] Mapping function for facilities accessible to the visually impaired

[1131] Step 1:

[1132] The device scans the signal strength of wireless LAN access points.

[1133] Step 2:

[1134] The scan results are sent to the server.

[1135] Step 3:

[1136] The server identifies the user's current location based on the signal from the wireless LAN access point and compares it with a map database.

[1137] Step 4:

[1138] The server determines the user's current location.

[1139] Step 5:

[1140] The server calculates the route to the destination based on the user's instructions (voice commands).

[1141] Step 6:

[1142] The server sends the route information to the terminal.

[1143] Step 7:

[1144] The route information received by the terminal is passed to a voice synthesis engine to generate voice data.

[1145] Step 8:

[1146] The voice guidance generated by the terminal is notified to the user through the speaker.

[1147] Image Recognition Function

[1148] Step 1:

[1149] The user takes a photo of the product they are holding with the device's camera.

[1150] Step 2:

[1151] The captured image is sent to a server.

[1152] Step 3:

[1153] The server uses an image recognition algorithm to identify the product in the image.

[1154] Step 4:

[1155] Based on the product name identified by the server, detailed product information (such as name, price, and description) is obtained from the Internet.

[1156] Step 5:

[1157] The server organizes the detailed product information it has acquired and generates text data.

[1158] Step 6:

[1159] The server transmits the generated text data to the terminal.

[1160] Step 7:

[1161] The information received by the terminal is passed to a speech synthesis engine to generate voice data.

[1162] Step 8:

[1163] The audio data generated by the terminal is notified to the user through the speaker.

[1164] Registered user recognition function

[1165] Step 1:

[1166] The device captures surrounding sounds using a microphone.

[1167] Step 2:

[1168] The captured audio data is sent to a server.

[1169] Step 3:

[1170] The server analyzes the voice data using a voice recognition algorithm.

[1171] Step 4:

[1172] The analysis results are compared with a pre-registered voice database, and the server identifies the person.

[1173] Step 5:

[1174] The server generates the identified information as text data (e.g., "This is Mr. Tanaka").

[1175] Step 6:

[1176] The server transmits the generated text data to the terminal.

[1177] Step 7:

[1178] The information received by the terminal is passed to a speech synthesis engine to generate voice data.

[1179] Step 8:

[1180] The audio data generated by the terminal is notified to the user through the speaker.

[1181] Example 2

[1182] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1183] There is a need for a system that comprehensively resolves the multiple obstacles and difficulties that visually impaired people face in their daily lives. Conventional systems only individually address a wide range of requirements, such as obstacle detection, location identification, person identification, product information acquisition, and even user emotion analysis, and do not provide a unified solution. Furthermore, it is difficult to provide appropriate feedback based on the user's real-time situation and emotions, which limits the effectiveness of these systems in improving the safety and comfort of visually impaired people.

[1184] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1185] In this invention, the server includes an imaging means for capturing images of the surroundings, an image analysis means for analyzing the images and detecting obstacles, a notification means for providing audio notification based on the detected information, a location identification means for periodically scanning wireless signals to identify the current location, a route guidance means for analyzing voice commands and providing route guidance, a product identification means for capturing images of products and obtaining identification information, a voice analysis means for capturing surrounding voices and identifying specific people, and an emotion analysis means for analyzing the user's emotional state and providing feedback. This makes it possible to provide a system integrating multiple functions to centrally solve various problems in the daily lives of visually impaired people and significantly improve safety and comfort.

[1186] An "imaging means" is a device that captures an image of the surroundings.

[1187] "Image analysis means" is a technology for analyzing captured images and detecting obstacles and their positions.

[1188] The "notification means" is a technology for providing a voice notification based on the analyzed information.

[1189] "Location determination means" is a technology that periodically scans for radio signals to determine the current location.

[1190] "Route guidance means" is a technology for analyzing voice commands and providing guidance on the route to a destination.

[1191] The "product identification means" is a technology for analyzing the captured image of a product and obtaining its identification information.

[1192] "Audio analysis means" is a technology for capturing surrounding sounds and identifying specific people.

[1193] "Emotion analysis means" is a technology for analyzing the user's emotional state and providing feedback based on that state.

[1194] The visually impaired person support system of the present invention is mainly composed of a server, a terminal, and a user. This system is designed to provide multiple integrated functions to enable visually impaired people to spend their daily lives comfortably and safely. Specific embodiments of the present invention will be described in detail below.

[1195] Device Features

[1196] The device has the function of capturing images of the user's surroundings at regular intervals (e.g., every second) using a camera while the user is walking. The hardware used is a smartphone camera, and the software is a camera app. The captured image data is sent to a server in real time. The device also has the function of periodically scanning the signal strength of wireless LAN access points and sending the scan results to the server. It also has the function of capturing images of products held by the user and sending them to the server. As an audio recording function, it also captures surrounding sounds using a microphone.

[1197] Server Features

[1198] The server decodes the received image data and runs an image recognition algorithm. The software used is TensorFlow. The server detects obstacles in the image, the distance to the obstacles, the signal status, and the presence or absence of walls. It also compares the user's current location with a map database based on the signal from the Wi-Fi access point and calculates a route to the destination. The software used is Google Maps API. The server also uses an image recognition algorithm to identify products and retrieve detailed product information (such as name, price, and description) from the Internet. It also has the ability to analyze voice data using a voice recognition algorithm to identify specific people. Finally, an emotion engine is used to analyze the user's emotional state from the captured data. The software used is Affectiva.

[1199] User Notification

[1200] The server generates the notification information (e.g., "There is an obstacle 2 meters ahead") based on the analysis results and sends it to the device. The device then notifies the user using a speech synthesis engine. The software used is Google Text-to-Speech.

[1201] Specific examples

[1202] For example, if a user's smartphone camera captures images in real time while walking and the server detects an obstacle, the device will announce, "There is an obstacle two meters ahead." If the user enters the voice command, "I want to go to the library," the device will provide voice guidance, saying, "Turn left, walk 20 meters, then turn right." If the user picks up an item and takes a picture of it with the camera, the device will announce, "This is a 347-page novel, priced at 1,200 yen." If a person in front of the user speaks, the device will announce, "This is Mr. Tanaka." Finally, if the device analyzes that the user is feeling stressed, it will provide voice advice, saying, "Relax and take a deep breath."

[1203] Example prompts to input to the generative AI model

[1204] 1. Prompt to explain specific examples of walking navigation features:

[1205] "Regarding navigation features for safe walking for the visually impaired, please explain how the system notifies users of obstacles in real time as they walk."

[1206] 2. Prompt to explain the visually impaired accessibility mapping feature:

[1207] "Please explain with specific examples the voice guidance function that allows users to find directions to their destination."

[1208] 3. Prompt to explain specific examples of image recognition features:

[1209] "Please explain how you can provide audio information about an item when the user takes a photo of it."

[1210] 4. Prompt to explain specific examples of subscriber recognition features:

[1211] "Please provide a concrete example of a feature that uses voice to identify and notify the user of specific people in their vicinity."

[1212] 5. Prompt to explain a specific example of an emotion engine:

[1213] "Please give a concrete example of a system that analyzes a user's emotional state and provides appropriate feedback."

[1214] The support system for the visually impaired of the present invention integrates multiple functions in this way, aiming to provide a unified solution to the various problems that visually impaired people encounter in their daily lives. In addition, by utilizing an emotion engine, it is possible to provide appropriate support according to the user's psychological state, improving the quality of life.

[1215] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1216] Processing steps of walking navigation function

[1217] Step 1: Capture an image

[1218] Input: Surrounding image (real-time)

[1219] How it works: The device uses the smartphone camera to capture an image of the user's surroundings every second.

[1220] Data processing: Capture as image data

[1221] Output: Captured image data

[1222] Step 2: Sending images

[1223] Input: Captured image data

[1224] How it works: The device sends captured image data to the server in real time.

[1225] Data processing: data encoding and transmission processing

[1226] Output: Image data sent to the server

[1227] Step 3: Image analysis

[1228] Input: Received image data

[1229] How it works: The server decodes the received image data and uses TensorFlow to run image recognition algorithms to detect obstacles in the image, distance, signal status, and the presence or absence of walls.

[1230] Data calculation: Analysis using image recognition algorithms

[1231] Output: Obstacle detection results and analysis data

[1232] Step 4: Generate information

[1233] Input: Image analysis results

[1234] Operation: Based on the analysis results, the server generates information to notify the user (e.g., "There is an obstacle 2 meters ahead").

[1235] Data calculation: Transformation of analysis results and generation of notification messages

[1236] Output: Information message

[1237] Step 5: Notification

[1238] Input: Notification message

[1239] What it does: The server generates a notification message and sends it to the device, which then uses Google Text-to-Speech to notify the user audibly.

[1240] Data operations: message conversion and notification processing

[1241] Output: Synthesized notification

[1242] Blind Accessible Facilities Mapping Feature Processing Steps

[1243] Step 1: Locate your location

[1244] Input: Wi-Fi access point signal strength

[1245] How it works: Your device periodically scans the signal strength of Wi-Fi access points.

[1246] Data Processing: Signal Strength Data Collection

[1247] Output: Scan result data

[1248] Step 2: Submit scan results

[1249] Input: Scan result data

[1250] Operation: The device sends the scan result data to the server.

[1251] Data processing: data encoding and transmission processing

[1252] Output: Scan result data sent to the server

[1253] Step 3: Location analysis

[1254] Input: Scan result data

[1255] How it works: The server compares the scanned data with a map database to determine the user's current location. It uses the Google Maps API.

[1256] Data calculation: Matching signal strength data with map data, and determining location

[1257] Output: Current location information

[1258] Step 4: Route calculation

[1259] Input: Current location

[1260] How it works: The server calculates the route to the destination based on the current location information.

[1261] Data calculation: Execution of route calculation algorithms

[1262] Output: Route information

[1263] Step 5: Directions

[1264] Input: Route information

[1265] How it works: The server sends route information to the device, which then uses Google Text-to-Speech to provide voice directions.

[1266] Data calculation: Route information voice conversion and notification processing

[1267] Output: Speech-synthesized route directions

[1268] Image Recognition Function Processing Steps

[1269] Step 1: Capture an image

[1270] Input: Item held in hand (real-time)

[1271] How it works: A user takes a picture of a product with their smartphone camera.

[1272] Data processing: Capture as image data

[1273] Output: Captured image data

[1274] Step 2: Sending images

[1275] Input: Captured image data

[1276] Action: The device sends the captured image data to the server.

[1277] Data processing: data encoding and transmission processing

[1278] Output: Image data sent to the server

[1279] Step 3: Product Identification

[1280] Input: Received image data

[1281] How it works: The server uses image recognition algorithms to identify the product in the image.

[1282] Data calculation: Running image recognition algorithms

[1283] Output: Product identification results

[1284] Step 4: Information Acquisition

[1285] Input: Product identification result

[1286] Operation: Based on the identification results, the server retrieves detailed product information (name, price, description, etc.) from the Internet.

[1287] Data Calculation: Search and Extraction of Details

[1288] Output: Product details

[1289] Step 5: Information Notification

[1290] Input: Product details

[1291] How it works: The server sends product details to the device, which then uses Google Text-to-Speech to notify the user aloud.

[1292] Data Computing: Information Speech Conversion and Notification Processing

[1293] Output: Speech-synthesized product information

[1294] Process steps for subscriber recognition

[1295] Step 1: Capture audio

[1296] Input: Ambient audio (real-time)

[1297] How it works: Your device captures ambient sound with its microphone.

[1298] Data processing: Capture as audio data

[1299] Output: Captured audio data

[1300] Step 2: Sending audio

[1301] Input: Captured audio data

[1302] What it does: The device sends the captured audio data to the server.

[1303] Data processing: data encoding and transmission processing

[1304] Output: Audio data sent to the server

[1305] Step 3: Voice analysis and person identification

[1306] Input: Received audio data

[1307] How it works: The server analyzes the audio data using a speech recognition algorithm to identify a specific person. The software used is Google Speech-to-Text.

[1308] Data Computing: Speech Recognition Algorithms and Person Identification

[1309] Output: Person identification result

[1310] Step 4: Notification of identification results

[1311] Input: Person identification result

[1312] How it works: The server sends the identification results to the device, which then notifies the user using Google Text-to-Speech.

[1313] Data processing: speech conversion and notification of the recognition results

[1314] Output: Synthesized voice notification of person identification

[1315] Emotion Engine Processing Steps

[1316] Step 1: Capture your voice and facial expressions

[1317] Input: User's voice and facial expressions (real-time)

[1318] How it works: The device captures the user's voice and facial expressions using sensors and microphones.

[1319] Data processing: Capture as voice and facial expression data

[1320] Output: Captured audio and facial expression data

[1321] Step 2: Sending data

[1322] Input: Captured voice and facial expression data

[1323] Action: The device sends the captured data to the server.

[1324] Data processing: data encoding and transmission processing

[1325] Output: Voice and facial expression data sent to the server

[1326] Step 3: Sentiment Analysis

[1327] Input: Received voice and facial expression data

[1328] How it works: The server uses an emotion analysis algorithm to analyze the user's emotional state from the captured data. The software used is Affectiva.

[1329] Data Computing: Running sentiment analysis algorithms and classifying emotional states

[1330] Output: Emotion analysis results

[1331] Step 4: Emotion-based feedback

[1332] Input: Sentiment analysis results

[1333] How it works: The server generates notification information (e.g., "Walk a little slower") based on the analyzed emotional state.

[1334] Data calculation: Transformation of analysis results and generation of notification messages

[1335] Output: Information message

[1336] Step 5: Notification

[1337] Input: Notification message

[1338] What it does: The server generates a notification message and sends it to the device, which then uses Google Text-to-Speech to notify the user audibly.

[1339] Data operations: message conversion and notification processing

[1340] Output: Synthesized notification

[1341] Through the above processing steps, the visually impaired person support system of the present invention provides the user with a safe and comfortable life.

[1342] (Application example 2)

[1343] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1344] Visually impaired people face many challenges in their daily lives. In particular, shopping and traveling in brick-and-mortar stores are often extremely inconvenient and dangerous for them. They also have difficulty recognizing obstacles, identifying specific people, and obtaining product information, which limits their independent lifestyles. Effective support systems that can solve these problems and enable visually impaired people to live their daily lives with peace of mind are needed.

[1345] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an imaging means for capturing images of the surroundings, an image analysis means for analyzing the images captured by the imaging means and detecting obstacles, a notification means for providing a voice notification based on information detected by the image analysis means, a registered person recognition means for identifying registered people and providing a voice notification based on the identification result, and an emotion analysis and notification means for analyzing the user's emotional state and providing feedback based on the emotional state. This allows visually impaired people to move safely while avoiding obstacles, to identify specific people to communicate smoothly, and to receive support according to their emotional state.

[1346] 1. "Imaging means" means a device for capturing images of the surroundings, including a camera or sensor.

[1347] 2. "Image Analysis Means" means the algorithms or processes that analyze captured images to detect obstacles and extract other information.

[1348] 3. "Notification means" refers to an interface for notifying the user of necessary information by voice or other means.

[1349] 4. "Walking navigation means" is a function that provides real-time voice navigation to obstacles.

[1350] 5. "Means for recognizing registered persons" refers to a function that identifies registered persons and sends notifications based on the results of that identification.

[1351] 6. "Emotion analysis and notification means" means a means for analyzing the user's emotional state and providing feedback according to that state in the form of voice or other means.

[1352] 7. "Location identification means" is a function in which the image analysis means analyzes location information and identifies the user's current location.

[1353] 8. "Store guidance means" is a function that provides audio guidance to the user about each area within a physical store.

[1354] This invention provides a support system that enables visually impaired people to navigate and shop safely in brick-and-mortar stores. The system has the following main functions:

[1355] System configuration

[1356] 1. Imaging Method

[1357] It uses smart glasses or a smartphone camera to capture images of the surroundings.

[1358] Example: As a user walks through a physical store, a camera captures images of the user's surroundings at regular intervals.

[1359] 2. Image analysis methods

[1360] The captured image is sent to the server, where it is analyzed in real time.

[1361] The software used includes image analysis algorithms and TensorFlow. The server recognizes obstacles and products and sends that information to the notification means.

[1362] Example: If an obstacle is detected, the analysis result will be generated as "There is an obstacle 2 meters ahead."

[1363] 3. Means of notification

[1364] Based on the analysis results, the information is notified to the user by voice. Voice synthesis uses a voice engine such as pyttsx3.

[1365] Example: A speech synthesis engine informs the user, "There is an obstacle 2 meters ahead."

[1366] 4. Walking navigation methods

[1367] Real-time voice navigation is provided while the user is moving, which is achieved by combining imaging means, image analysis means, and notification means.

[1368] Example: When a user arrives at a specific area (e.g., fresh produce), they are notified, "You have arrived at the fish section."

[1369] 5. Registered User Identification Method

[1370] It captures surrounding sounds and identifies specific people, using the microphone in smart glasses or smartphones to record the sounds.

[1371] The server performs voice analysis and notifies the user of the results of recognizing a specific person.

[1372] Example: When a person in front of the user speaks, the user is notified that "This is a store clerk."

[1373] 6. Sentiment Analysis and Notification Methods

[1374] It captures the user's voice and facial expressions and analyzes their emotional state using an emotion analysis model powered by TensorFlow.

[1375] Appropriate feedback is provided via voice based on the analysis results.

[1376] For example, if the user is analyzed as feeling stressed, they will be notified to "relax and take a deep breath."

[1377] Specific examples

[1378] Consider a scenario in which the system would work effectively when a visually impaired person visits a supermarket. The user walks through various sections of the store, and the smart glasses provide real-time information about their surroundings. For example, when the user scans an item they have picked up with the camera, they are notified by voice, "This is an apple. The price is 150 yen." When they arrive at a specific area, they are informed, "You have arrived at the fish section."

[1379] Prompt Sentence Examples

[1380] An example of a specific prompt sentence for the generative AI model is shown below.

[1381] A smart glasses application for visually impaired people in shopping malls provides walking navigation, product recognition, staff recognition, and emotional feedback. As an example, when a user scans an item they have picked up, the application asks, "How can the application recognize the product from the image captured by the camera and announce its name, price, and description by voice?"

[1382] This will enable visually impaired people to live their daily lives more safely and independently.

[1383] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1384] Step 1:

[1385] The device uses an imaging means (camera) to capture images of the surroundings. The imaging means captures images from the camera at regular intervals (e.g., every second) according to the user's viewpoint, and sends the image data to a server. The input is the image captured from the camera, and the output is the image data sent to the server.

[1386] Step 2:

[1387] The server analyzes the image data it receives. The image analysis means uses TensorFlow's image recognition algorithm to identify obstacles, products, and people in the image. The input is the image data sent to the server, and the output is the analysis results (obstacle position, product information, and person information). Specifically, the server extracts features from the image and performs analysis using a pre-trained model.

[1388] Step 3:

[1389] The server generates the information to be notified based on the analysis results. The notification means converts the analysis results (e.g., "There is an obstacle 2 meters ahead") into voice using a speech synthesis engine (pyttsx3) and notifies the user. The input is the analysis results from the image analysis means, and the output is a voice notification. This is achieved by the server converting the analysis results into text and sending it to the speech synthesis engine.

[1390] Step 4:

[1391] When the user takes a specific action (e.g., picking up a product), the image recognition function is activated. The product is photographed with the device's camera and the image is sent to the server. The input is the image of the product, and the output is the data sent to the server. The server identifies the product, obtains product information (name, price, etc.), and notifies the user by voice.

[1392] Step 5:

[1393] When the user moves, the walking navigation means operates. The terminal captures images in real time using the imaging means and sends them to the server. The server analyzes obstacle information and notifies the user of navigation information. The input is image data captured in real time, and the output is navigation information. The server processes the data in real time and generates navigation information.

[1394] Step 6:

[1395] The device's microphone captures surrounding sounds and activates the registered user recognition means. The voice data is sent to the server, which then uses a voice recognition algorithm to identify the person. The input is the captured voice data, and the output is specific person information. The server notifies the user of the identification result through a voice synthesis engine.

[1396] Step 7:

[1397] The terminal captures the user's voice and facial expressions, and the emotion analysis and notification means are activated. The server uses an emotion analysis algorithm to analyze the emotional state and provide appropriate feedback. The input is the captured voice and facial expression data, and the output is feedback information based on the emotional state. The server analyzes the emotional state and generates information to be notified.

[1398] Step 8:

[1399] When a user goes shopping in a physical store, the store guidance means is activated. The server identifies the user's current location, calculates the route to the destination, and provides voice guidance. The input is data obtained by scanning the signal strength of wireless LAN access points, and the output is route guidance information. The server sends the route guidance information to a speech synthesis engine, which notifies the user by voice.

[1400] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1401] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1402] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1403] [Third embodiment]

[1404] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1405] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1406] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1407] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1408] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1409] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1410] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1411] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1412] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1413] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1414] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1415] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1416] The visually impaired support system of the present invention has a plurality of functions for solving various problems that users face in their daily lives. The embodiments of the present invention will be described in detail below.

[1417] Walking navigation function

[1418] 1. Image capture

[1419] The device captures images of the user's surroundings with its camera at regular intervals (e.g., every second).

[1420] The captured images are sent to a server in real time.

[1421] 2. Image Analysis

[1422] The server decodes the received image and runs an image recognition algorithm on it.

[1423] The server detects obstacles in the image, the distance to the obstacles, the signal status, and the presence or absence of walls.

[1424] 3. Information Generation and Notification

[1425] Based on the image analysis results, the server generates information to be notified to the user (e.g., "There is an obstacle 2 meters ahead").

[1426] The server sends the generated information to the terminal, which notifies the user using a voice synthesis engine.

[1427] Example: While the user is walking, the device announces, "There is an obstacle 2 meters ahead."

[1428] Mapping function for facilities accessible to the visually impaired

[1429] 1. Identifying your current location

[1430] The device periodically scans the signal strength of Wi-Fi access points in its current location.

[1431] The scan results are sent to the server.

[1432] 2. Location analysis and route guidance

[1433] The server compares the user's current location with a map database based on the signal from the wireless LAN access point.

[1434] The server determines the user's current location and calculates a route to the destination.

[1435] The server transmits route information to the terminal, and the terminal uses a voice synthesis engine to provide route guidance to the user.

[1436] Example: When a user enters the voice command "I want to go to the library," the device will provide voice guidance such as "Turn left and go 20 meters, then turn right."

[1437] Image Recognition Function

[1438] 1. Image capture

[1439] The user takes a photo of the product they are holding with the device's camera.

[1440] The captured image is sent to a server.

[1441] 2. Product Identification and Information Notification

[1442] The server uses an image recognition algorithm to identify the product in the image.

[1443] Based on the identification results, the server obtains detailed product information (name, price, description, etc.) from the Internet.

[1444] The server sends the acquired detailed product information to the terminal, which then notifies the user using a voice synthesis engine.

[1445] Example: When a user picks up an item and takes a picture of it with the camera, the device announces in a voice message, "This is the 347-page novel 'The Blue Bird' and costs 1,200 yen."

[1446] Registered user recognition function

[1447] 1. Audio capture

[1448] The device captures surrounding sounds using a microphone.

[1449] The captured audio data is sent to a server.

[1450] 2. Voice analysis and identity notification

[1451] The server uses a voice recognition algorithm to analyze the voice data and identify a specific person.

[1452] The server transmits the identification result to the terminal, and the terminal notifies the user using a voice synthesis engine.

[1453] Example: When a person in front of the user speaks, the device will announce in voice, "This is Tanaka-san."

[1454] The objective of the support system for visually impaired people of the present invention is to integrate these functions to efficiently and effectively solve the various problems that visually impaired people encounter in their daily lives, thereby enabling them to live safer and more independent lives.

[1455] The processing flow will be explained below.

[1456] Walking navigation function

[1457] Step 1:

[1458] The device captures images of the user's surroundings at regular intervals using a camera.

[1459] Step 2:

[1460] The device compresses the captured images and sends them to the server in real time.

[1461] Step 3:

[1462] The server decodes the received image and runs an image recognition algorithm on it.

[1463] Step 4:

[1464] The server detects obstacles in the image, the distance to the obstacles, the signal status, and the presence or absence of walls.

[1465] Step 5:

[1466] The server generates the information to be notified based on the results of image analysis (e.g., "There is an obstacle 2 meters ahead").

[1467] Step 6:

[1468] The server transmits the generated information to the terminal.

[1469] Step 7:

[1470] The information received by the terminal is passed to a speech synthesis engine to generate voice data.

[1471] Step 8:

[1472] The audio data generated by the terminal is notified to the user through the speaker.

[1473] Mapping function for facilities accessible to the visually impaired

[1474] Step 1:

[1475] The device scans the signal strength of wireless LAN access points.

[1476] Step 2:

[1477] The device sends the scan results to the server.

[1478] Step 3:

[1479] The server identifies the user's current location based on the signal from the wireless LAN access point by comparing it with a location information database.

[1480] Step 4:

[1481] The server determines the user's current location.

[1482] Step 5:

[1483] The server calculates the route to the destination based on the user's instructions (voice commands).

[1484] Step 6:

[1485] The server sends the route information to the terminal.

[1486] Step 7:

[1487] The route information received by the terminal is passed to a voice synthesis engine to generate voice data.

[1488] Step 8:

[1489] The voice guidance generated by the terminal is notified to the user through the speaker.

[1490] Image Recognition Function

[1491] Step 1:

[1492] The user takes a photo of the product they are holding with the device's camera.

[1493] Step 2:

[1494] The device sends the captured image to the server.

[1495] Step 3:

[1496] The server uses an image recognition algorithm to identify the product in the image.

[1497] Step 4:

[1498] Based on the product name identified by the server, detailed product information (such as name, price, and description) is obtained from the Internet.

[1499] Step 5:

[1500] The server organizes the detailed product information it has acquired and generates text data.

[1501] Step 6:

[1502] The server transmits the generated text data to the terminal.

[1503] Step 7:

[1504] The information received by the terminal is passed to a speech synthesis engine to generate voice data.

[1505] Step 8:

[1506] The audio data generated by the terminal is notified to the user through the speaker.

[1507] Registered user recognition function

[1508] Step 1:

[1509] The device captures surrounding sounds using a microphone.

[1510] Step 2:

[1511] The device sends the captured audio data to the server.

[1512] Step 3:

[1513] The server analyzes the voice data using a voice recognition algorithm.

[1514] Step 4:

[1515] The server compares the analysis results with a pre-registered voice database to identify the person.

[1516] Step 5:

[1517] The server generates the identified information as text data (e.g., "This is Mr. Tanaka").

[1518] Step 6:

[1519] The server transmits the generated text data to the terminal.

[1520] Step 7:

[1521] The information received by the terminal is passed to a speech synthesis engine to generate voice data.

[1522] Step 8:

[1523] The audio data generated by the terminal is notified to the user through the speaker.

[1524] Example 1

[1525] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1526] To solve the various challenges that visually impaired people face in their daily lives, it is necessary to properly obtain information about their surroundings and convey that information to them quickly and accurately. However, current systems lack comprehensive support that goes beyond simply detecting obstacles, covering things like determining their current location, identifying people, and even providing traffic light status and specific route guidance. This leaves visually impaired people at great risk.

[1527] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1528] In this invention, the server includes an imaging means for capturing an image of the surroundings, an image analysis means for analyzing the image captured by the imaging means and detecting obstacles, traffic light status, and the presence or absence of walls, an information generation means for generating notification information based on the information detected by the image analysis means, a notification means for notifying by voice the notification information generated by the information generation means, a location identification means for analyzing location information using a wireless LAN signal and identifying the current location, an audio recording means for capturing surrounding sounds, an audio analysis means for analyzing the sounds captured by the audio recording means and identifying a specific person, an information generation means for generating notification information based on the person information identified by the audio analysis means, and a notification means for notifying by voice the notification information generated by the information generation means.This enables visually impaired people to avoid obstacles while walking, identify their current location, and easily identify people, enabling them to live safe and independent lives.

[1529] An "imaging means" is a device or sensor for capturing an image of the surroundings.

[1530] "Image analysis means" refers to a set of algorithms and software for analyzing captured images and detecting obstacles, traffic light conditions, the presence or absence of walls, etc.

[1531] The "notification means" is a device or software that conveys the analyzed information to the user by voice.

[1532] The "information generation means" refers to logic or algorithms for generating information to be notified to the user based on the analysis results.

[1533] The "location determination means" is a device or software for determining the user's current location using a wireless LAN signal.

[1534] "Audio recording means" refers to a microphone or other audio recording device for capturing ambient audio.

[1535] "Audio analysis means" refers to algorithms or software that analyzes captured audio and identifies specific individuals.

[1536] The visually impaired support system of the present invention has a plurality of functions for solving various problems that users face in their daily lives. The embodiments of the present invention will be described in detail below.

[1537] Walking navigation function

[1538] Image capture

[1539] The device captures images of the user's surroundings at regular intervals (e.g., every second) using the built-in camera of a smartphone.

[1540] Sending images

[1541] The device immediately sends the captured image to the server using Wi-Fi or mobile data.

[1542] Image analysis

[1543] The server decodes the received images and uses image recognition algorithms such as TensorFlow and OpenCV to detect obstacles, traffic light status, and the presence or absence of walls.

[1544] Information Generation and Notification

[1545] Based on the analysis results, the server generates notification information such as "There is an obstacle 2 meters ahead." The information is sent from the server to the device, which then notifies the user using a voice synthesis engine such as Google Text-to-Speech.

[1546] Specific examples

[1547] While the user is walking, the device will announce to them by voice, "There is an obstacle 2 meters ahead."

[1548] Mapping function for facilities accessible to the visually impaired

[1549] Scan your current location

[1550] The device periodically scans the signal strength of surrounding Wi-Fi access points and sends the results of the scan to the server.

[1551] Location analysis and routing

[1552] The server uses signal data from wireless LAN access points to identify the user's current location using a map database such as Google Maps API and calculates the optimal route. The route information is sent to the device, which then uses a voice synthesis engine such as Google Text-to-Speech to provide route guidance to the user.

[1553] Specific examples

[1554] When a user inputs the voice command "I want to go to the library," the device will provide voice guidance such as "Turn left and go 20 meters, then turn right."

[1555] Image Recognition Function

[1556] Image capture

[1557] The user takes a photo of the product they are holding with the device's camera, and the captured image is sent to the server.

[1558] Product identification and information acquisition

[1559] The server identifies the product using an image recognition algorithm such as YOLOv3, and then uses the Amazon API or Rakuten API to obtain detailed product information (name, price, description, etc.).

[1560] Information Notification

[1561] The acquired product details are sent from the server to the terminal, which then notifies the user using a voice synthesis engine such as Google Text-to-Speech.

[1562] Specific examples

[1563] When the user picks up the product and takes a picture of it with the camera, the device announces in voice, "This is the 347-page novel 'The Blue Bird' and costs 1,200 yen."

[1564] Registered user recognition function

[1565] Audio capture

[1566] The device captures the surrounding sound with a microphone, and the captured sound data is sent to the server.

[1567] Audio analysis and notifications

[1568] The server analyzes the voice data using a speech recognition algorithm such as IBM Watson Speech to Text to identify a specific person, and the identification result is sent from the server to the device, which then notifies the user using a speech synthesis engine such as Google Text-to-Speech.

[1569] Specific examples

[1570] When a person in front of the user speaks, the device will announce aloud, "This is Tanaka-san."

[1571] Examples of prompt statements

[1572] 1. Walking navigation function

[1573] Prompt: "What is the process for the system that notifies you of nearby obstacles?"

[1574] 2. Mapping function for facilities accessible to the visually impaired

[1575] Prompt: "Please explain the process of the system that identifies your current location and provides directions to your destination."

[1576] 3. Image Recognition

[1577] Prompt: "Please tell me the process of the system that takes a picture of a product with a camera and provides detailed information about it."

[1578] 4. Registered User Recognition Function

[1579] Prompt: "Please explain the process for a system that analyzes surrounding sounds and identifies a specific person."

[1580] The above is an embodiment of the visually impaired support system of the present invention, and this system enables visually impaired people to live safer and more independent lives.

[1581] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1582] Walking navigation function

[1583] Step 1:

[1584] Image capture

[1585] The device captures images of the user's surroundings with a camera at regular intervals (every second). The input is image data obtained from the camera, and the output is the captured image file.

[1586] Specific operation: The device automatically launches the camera app and simulates pressing the shutter button every second.

[1587] Step 2:

[1588] Sending images

[1589] The image files captured by the device are sent to the server in real time. The input is the image file generated in step 1, and the output is the upload of image data to the server.

[1590] Specific operation: The device compresses the image data and uploads it to the server using the HTTPS protocol.

[1591] Step 3:

[1592] Image analysis

[1593] The server decodes the received image and analyzes it using an image recognition algorithm using TensorFlow, OpenCV, etc. The input is the image data sent to the server, and the output is the detection results such as obstacles, traffic light status, and the presence or absence of walls.

[1594] Specific operation: The server processes the received image, runs a recognition algorithm, and generates an analysis result.

[1595] Step 4:

[1596] information generation

[1597] The server generates information to be notified to the user based on the analysis results. The input is the analysis results obtained in step 3, and the output is a notification message.

[1598] What happens: The server uses a text generation algorithm to generate a message such as "There is an obstacle 2 meters ahead."

[1599] Step 5:

[1600] Information Notification

[1601] The server sends the generated notification message to the device, and the device uses a speech synthesis engine such as Google Text-to-Speech to notify the user audibly. The input is the generated notification message, and the output is the audio notification delivered to the user.

[1602] Specific operation: The device converts the received message into voice and notifies the user through the speaker, "There is an obstacle 2 meters ahead."

[1603] Mapping function for facilities accessible to the visually impaired

[1604] Step 1:

[1605] Scan your current location

[1606] The device periodically scans the signal strength of surrounding Wi-Fi access points. The input is the Wi-Fi signal captured by the Wi-Fi module, and the output is the signal strength data.

[1607] Specific operation: The device activates the Wi-Fi module and collects the surrounding network signal strength.

[1608] Step 2:

[1609] Sending scan results

[1610] The terminal sends the collected signal strength data to the server. The input is the signal strength data collected in step 1, and the output is the upload to the server.

[1611] Specific operation: The terminal compresses the signal data and uploads it to the server via HTTPS protocol.

[1612] Step 3:

[1613] Location analysis

[1614] The server analyzes the location information based on the signal data of the wireless LAN access point and determines the current location. The input is the transmitted signal strength data, and the output is the determined current location information.

[1615] What it does: The server analyzes the signal data and locates it by comparing it with a map database (e.g., Google Maps API).

[1616] Step 4:

[1617] Route calculation

[1618] The server calculates the optimal route from the specified current location to the destination. The input is the current location information and the destination information, and the output is the route information.

[1619] Specific operation: The server executes a route calculation algorithm (e.g., A search algorithm) to calculate the route to the destination.

[1620] Step 5:

[1621] Route guidance

[1622] The server sends route information to the device, which then uses a speech synthesis engine such as Google Text-to-Speech to provide route guidance to the user. The input is the calculated route information, and the output is voice guidance provided to the user.

[1623] Specific operation: The device converts the route information it receives into voice and notifies the user, "Turn left, go 20 meters, then turn right."

[1624] Image Recognition Function

[1625] Step 1:

[1626] Image capture

[1627] The user takes a photo of the product they are holding with the device's camera. The input is image data acquired from the camera, and the output is the captured image file.

[1628] Specific operation: The device automatically launches the camera app, and the user presses the shutter button.

[1629] Step 2:

[1630] Sending images

[1631] The device sends the captured image file to the server. The input is the captured image file, and the output is the upload of the image data to the server.

[1632] Specific operation: The device compresses the image data and uploads it to the server using the HTTPS protocol.

[1633] Step 3:

[1634] Product Identification

[1635] The server uses an image recognition algorithm (e.g., YOLOv3) to identify the product in the image. The input is the submitted image data, and the output is the identified product information.

[1636] Specific operation: The server analyzes the received image and identifies it as "This is a novel called 'The Blue Bird'."

[1637] Step 4:

[1638] Information acquisition

[1639] The server retrieves detailed product information (such as name, price, description, etc.) from the Internet. The input is the identified product information, and the output is the retrieved product detail data.

[1640] Specific operation: The server calls the API to obtain product data.

[1641] Step 5:

[1642] Information Notification

[1643] The server sends the acquired product details to the terminal, which then notifies the user using a speech synthesis engine such as Google Text-to-Speech. The input is the acquired product details data, and the output is a voice notification delivered to the user.

[1644] Specific operation: The device converts the received product information into voice and notifies the user, "This is the 347-page novel 'The Blue Bird', priced at 1,200 yen."

[1645] Registered user recognition function

[1646] Step 1:

[1647] Audio capture

[1648] The device captures the surrounding sound with a microphone. The input is the audio data collected by the microphone, and the output is the captured audio file.

[1649] Specific operation: The device activates the microphone and collects surrounding sounds.

[1650] Step 2:

[1651] Sending Audio

[1652] The device sends the captured audio file to the server. The input is the captured audio file, and the output is the audio data uploaded to the server.

[1653] Specific operation: The device compresses the audio data and uploads it to the server using the HTTPS protocol.

[1654] Step 3:

[1655] Audio analysis

[1656] The server analyzes the voice data using a speech recognition algorithm (e.g., IBM Watson Speech to Text) to identify a specific person. The input is the transmitted voice data, and the output is the identified person's information.

[1657] Specific operation: The server processes the voice data and generates an identification result such as "This is Mr. Tanaka."

[1658] Step 4:

[1659] Generating the identification results

[1660] The server generates notification information based on the identification result. The input is the identified person's information, and the output is the generated notification message.

[1661] Specific operation: The server generates a text message and sends it to the device.

[1662] Step 5:

[1663] Notification of identity

[1664] The device uses a speech synthesis engine such as Google Text-to-Speech to convert the notification message into speech and notify the user. The input is the generated notification message, and the output is the audio notification delivered to the user.

[1665] Specific operation: The device uses the speaker to announce, "This is Tanaka-san."

[1666] (Application example 1)

[1667] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1668] Due to their unique needs, it is difficult for visually impaired people to shop safely and comfortably in brick-and-mortar stores. To solve this problem, a system is needed that allows visually impaired people to independently select and navigate products in brick-and-mortar stores. In addition, since it is difficult for them to accurately grasp product information in real time, appropriate support is required.

[1669] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1670] In this invention, the server includes an imaging means for capturing an image of the surroundings, an image analysis means for analyzing the image captured by the imaging means and detecting obstacles, a notification means for providing a voice notification based on information detected by the image analysis means, an image recognition means for identifying products, an information acquisition means for acquiring detailed information about the identified products, and a notification means for providing a voice notification based on information acquired by the information acquisition means. This enables visually impaired people to receive product information by voice while safely moving around in a physical store.

[1671] An "imaging means" is a device that captures an image of the surroundings.

[1672] The "image analysis means" is a device that analyzes the image captured by the imaging means and detects objects such as obstacles.

[1673] The "notification means" is a device that notifies the user of information detected by the image analysis means and other information by voice.

[1674] An "image recognition means" is a device for identifying an item in a captured image.

[1675] The "information acquisition means" is a device that acquires detailed information (such as the name and price) of the identified product from the Internet or the like.

[1676] The "location specifying means" is a device that specifies the current location by analyzing the location information performed by the image analysis means.

[1677] An "audio recording means" is a device that captures ambient sounds.

[1678] The "voice analysis means" is a device that analyzes the voice captured by the voice recording means and identifies a specific person.

[1679] The present invention relates to a support system for visually impaired people to shop safely and comfortably in a brick-and-mortar store. Hereinafter, an embodiment of the present invention will be specifically described.

[1680] System Configuration

[1681] The system includes the following main components:

[1682] 1. An imaging means to capture images of the surroundings

[1683] 2. Image analysis means for analyzing images captured by the imaging means and detecting obstacles

[1684] 3. Notification means that notifies by voice based on information detected by image analysis means

[1685] 4. Image recognition methods for identifying products

[1686] 5. Information acquisition means for acquiring detailed information on identified products

[1687] 6. Notification method for notifying information by voice

[1688] Hardware and Software

[1689] Hardware:

[1690] Smartphone, smart glasses, or head-mounted display

[1691] Web camera (imaging means)

[1692] Microphone (audio recording means)

[1693] Speaker (notification means)

[1694] software:

[1695] OpenCV (library for image capture and analysis)

[1696] Google Text-to-Speech (gTTS) (voice notifications)

[1697] Image recognition API (product identification)

[1698] Internet connection (information acquisition)

[1699] Processing steps

[1700] When a user activates the system, the imaging means captures images of the surrounding area at regular intervals. The captured images are sent to a server in real time, and the image analysis means detects obstacles in the images and analyzes their positions. The notification means then notifies the user of the obstacle information by voice.

[1701] At the same time, if the user wants to identify a product, they take a photo of the product using a smartphone or smart glasses. The captured image is sent to a server, and the image recognition means identifies the product. Detailed information about the identified product is obtained from the Internet by an information acquisition means, and the information is notified to the user by voice via a notification means.

[1702] Specific examples

[1703] For example, when a user is walking in a shopping mall, the system will notify them that there is an obstacle two meters ahead, helping them to walk safely. Also, when a user picks up a product and takes a picture of it with the camera, the system will announce the product information in a voice message, saying, "This is a 347-page book, priced at 1,200 yen."

[1704] Prompt Sentence Examples

[1705] Example prompt for image recognition:

[1706] "Please recognize the item in this image and provide the item name and distance."

[1707] Example of a prompt for audio notification:

[1708] "2 meters ahead, there is an obstacle. It is a table."

[1709] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1710] Step 1:

[1711] The imaging means captures images of the surroundings at regular intervals (e.g., every second). The input is video data from the camera, which is captured and saved as image data. The output is the saved image data.

[1712] Step 2:

[1713] The device sends the captured image data to the server. The input is the image data obtained in step 1, which is uploaded to the server via the Internet. The output is the image data received by the server.

[1714] Step 3:

[1715] The server analyzes the image data it receives and identifies the presence and location of obstacles. The input is the image data received by the server, which is analyzed using an image analysis algorithm (e.g., OpenCV). The output is the location information of obstacles.

[1716] Step 4:

[1717] The server generates the content to be notified to the user based on the analysis results. The input is the obstacle location information obtained in step 3, and an appropriate notification message is created using the generative AI model. The output is the notification message (text data).

[1718] Step 5:

[1719] The server sends a notification message to the device, and the device notifies the user by voice. The input is the notification message created in step 4, which is input to a speech synthesis engine (e.g., gTTS) to generate voice data and play it through the speaker. The output is a voice notification.

[1720] Step 6:

[1721] When a user wants to identify a product, they take a picture of the product with the camera on their device. The input is the video data from the camera, which is captured and saved as image data. The output is the saved image data.

[1722] Step 7:

[1723] The device sends the captured product image to the server. The input is the image data obtained in step 6, which is uploaded to the server via the Internet. The output is the image data received by the server.

[1724] Step 8:

[1725] The server analyzes the product image and identifies the product. The input is the image data obtained in step 7, and the product is identified using an image recognition algorithm (e.g., image recognition API). The output is the identified product data (e.g., product name, price).

[1726] Step 9:

[1727] The server acquires detailed information based on the identified product data. The input is the product data acquired in step 8, and detailed information is collected from the Internet, etc. using an information acquisition means. The output is detailed product information.

[1728] Step 10:

[1729] The server sends the product information to the terminal, which then notifies the user by voice. The input is the detailed product information obtained in step 9, which is input into a speech synthesis engine (e.g., gTTS) to generate voice data and play it back through the speaker. The output is a voice notification of the product information.

[1730] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1731] The support system for visually impaired people of the present invention has multiple functions for solving various problems that users face in their daily lives. This system incorporates a walking navigation function, a facility mapping function for visually impaired people, an image recognition function, a registered user recognition function, and an emotion engine that recognizes the user's emotions. An embodiment of the present invention will be described in detail below.

[1732] Walking navigation function

[1733] 1. Image capture

[1734] The device captures images of the user's surroundings with its camera at regular intervals (e.g., every second).

[1735] The captured images are sent to a server in real time.

[1736] 2. Image Analysis

[1737] The server decodes the received image and runs an image recognition algorithm on it.

[1738] The server detects obstacles in the image, the distance to the obstacles, the signal status, and the presence or absence of walls.

[1739] 3. Information Generation and Notification

[1740] Based on the image analysis results, the server generates information to be notified to the user (e.g., "There is an obstacle 2 meters ahead").

[1741] The server sends the generated information to the terminal, which notifies the user using a voice synthesis engine.

[1742] Example: While the user is walking, the device announces, "There is an obstacle 2 meters ahead."

[1743] Mapping function for facilities accessible to the visually impaired

[1744] 1. Identifying your current location

[1745] The device periodically scans the signal strength of Wi-Fi access points in its current location.

[1746] The scan results are sent to the server.

[1747] 2. Location analysis and route guidance

[1748] The server compares the user's current location with a map database based on the signal from the wireless LAN access point.

[1749] The server determines the user's current location and calculates a route to the destination.

[1750] The server transmits route information to the terminal, and the terminal uses a voice synthesis engine to provide route guidance to the user.

[1751] Example: When a user enters the voice command "I want to go to the library," the device will provide voice guidance such as "Turn left and go 20 meters, then turn right."

[1752] Image Recognition Function

[1753] 1. Image capture

[1754] The user takes a photo of the product they are holding with the device's camera.

[1755] The captured image is sent to a server.

[1756] 2. Product Identification and Information Notification

[1757] The server uses an image recognition algorithm to identify the product in the image.

[1758] Based on the identification results, the server obtains detailed product information (name, price, description, etc.) from the Internet.

[1759] The server sends the acquired detailed product information to the terminal, which then notifies the user using a voice synthesis engine.

[1760] Example: When a user picks up an item and takes a picture of it with the camera, the device announces in a voice message, "This is the 347-page novel 'The Blue Bird' and costs 1,200 yen."

[1761] Registered user recognition function

[1762] 1. Audio capture

[1763] The device captures surrounding sounds using a microphone.

[1764] The captured audio data is sent to a server.

[1765] 2. Voice analysis and identity notification

[1766] The server uses a voice recognition algorithm to analyze the voice data and identify a specific person.

[1767] The server transmits the identification result to the terminal, and the terminal notifies the user using a voice synthesis engine.

[1768] Example: When a person in front of the user speaks, the device will announce in voice, "This is Tanaka-san."

[1769] Emotion Engine

[1770] 1. Voice and facial expression capture

[1771] The device captures the user's voice and facial expressions using sensors and microphones.

[1772] The captured data is sent to a server.

[1773] 2. Emotion analysis

[1774] The server uses an emotion engine to parse the user's emotional state from the captured data.

[1775] The analysis results are classified as emotional states such as "joy," "sadness," and "stress."

[1776] 3. Emotion-based feedback

[1777] The server generates notification information based on the analyzed emotional state (e.g., "Walk a little more slowly").

[1778] The server sends the generated information to the terminal, which notifies the user using a voice synthesis engine.

[1779] Example: If the device determines that the user is feeling stressed, it will provide voice advice such as "Relax and take a deep breath."

[1780] By integrating the above functions, the support system for visually impaired people of the present invention can provide support for efficiently and effectively resolving the various problems that visually impaired people encounter in their daily lives. In addition, by incorporating an emotion engine, it can provide support that is adapted to the user's emotional state, allowing visually impaired people to live their lives with greater peace of mind.

[1781] The processing flow will be explained below.

[1782] Emotion Engine

[1783] Step 1:

[1784] The device captures the user's voice and facial expressions at regular intervals using sensors and microphones. The sampling period for audio and video data is set appropriately (e.g., every second).

[1785] Step 2:

[1786] The captured voice and facial expression data is transmitted to a server in real time.

[1787] Step 3:

[1788] The server decodes the received voice and facial expression data and runs emotion recognition algorithms.

[1789] Step 4:

[1790] Using an emotion recognition algorithm, the server analyzes the user's emotional state (happiness, sadness, anger, stress, etc.) and assigns a score to each emotion.

[1791] Step 5:

[1792] The server generates a feedback message that corresponds to the emotional state (e.g., "Take a break" or "Relax and take a deep breath").

[1793] Step 6:

[1794] The server sends the generated feedback message to the terminal.

[1795] Step 7:

[1796] The terminal passes the received feedback message to a speech synthesis engine to generate speech data.

[1797] Step 8:

[1798] The audio data generated by the terminal is notified to the user through the speaker.

[1799] Walking navigation function

[1800] Step 1:

[1801] The device captures images of the user's surroundings at regular intervals using a camera.

[1802] Step 2:

[1803] The captured images are compressed and sent to a server in real time.

[1804] Step 3:

[1805] The server decodes the received image and runs an image recognition algorithm on it.

[1806] Step 4:

[1807] The server detects obstacles in the image, the distance to the obstacles, the signal status, and the presence or absence of walls.

[1808] Step 5:

[1809] The server generates the information to be notified based on the results of image analysis (e.g., "There is an obstacle 2 meters ahead").

[1810] Step 6:

[1811] The server transmits the generated information to the terminal.

[1812] Step 7:

[1813] The information received by the terminal is passed to a speech synthesis engine to generate voice data.

[1814] Step 8:

[1815] The audio data generated by the terminal is notified to the user through the speaker.

[1816] Mapping function for facilities accessible to the visually impaired

[1817] Step 1:

[1818] The device scans the signal strength of wireless LAN access points.

[1819] Step 2:

[1820] The scan results are sent to the server.

[1821] Step 3:

[1822] The server identifies the user's current location based on the signal from the wireless LAN access point and compares it with a map database.

[1823] Step 4:

[1824] The server determines the user's current location.

[1825] Step 5:

[1826] The server calculates the route to the destination based on the user's instructions (voice commands).

[1827] Step 6:

[1828] The server sends the route information to the terminal.

[1829] Step 7:

[1830] The route information received by the terminal is passed to a voice synthesis engine to generate voice data.

[1831] Step 8:

[1832] The voice guidance generated by the terminal is notified to the user through the speaker.

[1833] Image Recognition Function

[1834] Step 1:

[1835] The user takes a photo of the product they are holding with the device's camera.

[1836] Step 2:

[1837] The captured image is sent to a server.

[1838] Step 3:

[1839] The server uses an image recognition algorithm to identify the product in the image.

[1840] Step 4:

[1841] Based on the product name identified by the server, detailed product information (such as name, price, and description) is obtained from the Internet.

[1842] Step 5:

[1843] The server organizes the detailed product information it has acquired and generates text data.

[1844] Step 6:

[1845] The server transmits the generated text data to the terminal.

[1846] Step 7:

[1847] The information received by the terminal is passed to a speech synthesis engine to generate voice data.

[1848] Step 8:

[1849] The audio data generated by the terminal is notified to the user through the speaker.

[1850] Registered user recognition function

[1851] Step 1:

[1852] The device captures surrounding sounds using a microphone.

[1853] Step 2:

[1854] The captured audio data is sent to a server.

[1855] Step 3:

[1856] The server analyzes the voice data using a voice recognition algorithm.

[1857] Step 4:

[1858] The analysis results are compared with a pre-registered voice database, and the server identifies the person.

[1859] Step 5:

[1860] The server generates the identified information as text data (e.g., "This is Mr. Tanaka").

[1861] Step 6:

[1862] The server transmits the generated text data to the terminal.

[1863] Step 7:

[1864] The information received by the terminal is passed to a speech synthesis engine to generate voice data.

[1865] Step 8:

[1866] The audio data generated by the terminal is notified to the user through the speaker.

[1867] Example 2

[1868] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1869] There is a need for a system that comprehensively resolves the multiple obstacles and difficulties that visually impaired people face in their daily lives. Conventional systems only individually address a wide range of requirements, such as obstacle detection, location identification, person identification, product information acquisition, and even user emotion analysis, and do not provide a unified solution. Furthermore, it is difficult to provide appropriate feedback based on the user's real-time situation and emotions, which limits the effectiveness of these systems in improving the safety and comfort of visually impaired people.

[1870] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1871] In this invention, the server includes an imaging means for capturing images of the surroundings, an image analysis means for analyzing the images and detecting obstacles, a notification means for providing audio notification based on the detected information, a location identification means for periodically scanning wireless signals to identify the current location, a route guidance means for analyzing voice commands and providing route guidance, a product identification means for capturing images of products and obtaining identification information, a voice analysis means for capturing surrounding voices and identifying specific people, and an emotion analysis means for analyzing the user's emotional state and providing feedback. This makes it possible to provide a system integrating multiple functions to centrally solve various problems in the daily lives of visually impaired people and significantly improve safety and comfort.

[1872] An "imaging means" is a device that captures an image of the surroundings.

[1873] "Image analysis means" is a technology for analyzing captured images and detecting obstacles and their positions.

[1874] The "notification means" is a technology for providing a voice notification based on the analyzed information.

[1875] "Location determination means" is a technology that periodically scans for radio signals to determine the current location.

[1876] "Route guidance means" is a technology for analyzing voice commands and providing guidance on the route to a destination.

[1877] The "product identification means" is a technology for analyzing the captured image of a product and obtaining its identification information.

[1878] "Audio analysis means" is a technology for capturing surrounding sounds and identifying specific people.

[1879] "Emotion analysis means" is a technology for analyzing the user's emotional state and providing feedback based on that state.

[1880] The visually impaired person support system of the present invention is mainly composed of a server, a terminal, and a user. This system is designed to provide multiple integrated functions to enable visually impaired people to spend their daily lives comfortably and safely. Specific embodiments of the present invention will be described in detail below.

[1881] Device Features

[1882] The device has the function of capturing images of the user's surroundings at regular intervals (e.g., every second) using a camera while the user is walking. The hardware used is a smartphone camera, and the software is a camera app. The captured image data is sent to a server in real time. The device also has the function of periodically scanning the signal strength of wireless LAN access points and sending the scan results to the server. It also has the function of capturing images of products held by the user and sending them to the server. As an audio recording function, it also captures surrounding sounds using a microphone.

[1883] Server Features

[1884] The server decodes the received image data and runs an image recognition algorithm. The software used is TensorFlow. The server detects obstacles in the image, the distance to the obstacles, the signal status, and the presence or absence of walls. It also compares the user's current location with a map database based on the signal from the Wi-Fi access point and calculates a route to the destination. The software used is Google Maps API. The server also uses an image recognition algorithm to identify products and retrieve detailed product information (such as name, price, and description) from the Internet. It also has the ability to analyze voice data using a voice recognition algorithm to identify specific people. Finally, an emotion engine is used to analyze the user's emotional state from the captured data. The software used is Affectiva.

[1885] User Notification

[1886] The server generates the notification information (e.g., "There is an obstacle 2 meters ahead") based on the analysis results and sends it to the device. The device then notifies the user using a speech synthesis engine. The software used is Google Text-to-Speech.

[1887] Specific examples

[1888] For example, if a user's smartphone camera captures images in real time while walking and the server detects an obstacle, the device will announce, "There is an obstacle two meters ahead." If the user enters the voice command, "I want to go to the library," the device will provide voice guidance, saying, "Turn left, walk 20 meters, then turn right." If the user picks up an item and takes a picture of it with the camera, the device will announce, "This is a 347-page novel, priced at 1,200 yen." If a person in front of the user speaks, the device will announce, "This is Mr. Tanaka." Finally, if the device analyzes that the user is feeling stressed, it will provide voice advice, saying, "Relax and take a deep breath."

[1889] Example prompts to input to the generative AI model

[1890] 1. Prompt to explain specific examples of walking navigation features:

[1891] "Regarding navigation features for safe walking for the visually impaired, please explain how the system notifies users of obstacles in real time as they walk."

[1892] 2. Prompt to explain the visually impaired accessibility mapping feature:

[1893] "Please explain with specific examples the voice guidance function that allows users to find directions to their destination."

[1894] 3. Prompt to explain specific examples of image recognition features:

[1895] "Please explain how you can provide audio information about an item when the user takes a photo of it."

[1896] 4. Prompt to explain specific examples of subscriber recognition features:

[1897] "Please provide a concrete example of a feature that uses voice to identify and notify the user of specific people in their vicinity."

[1898] 5. Prompt to explain a specific example of an emotion engine:

[1899] "Please give a concrete example of a system that analyzes a user's emotional state and provides appropriate feedback."

[1900] The support system for the visually impaired of the present invention integrates multiple functions in this way, aiming to provide a unified solution to the various problems that visually impaired people encounter in their daily lives. In addition, by utilizing an emotion engine, it is possible to provide appropriate support according to the user's psychological state, improving the quality of life.

[1901] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1902] Processing steps of walking navigation function

[1903] Step 1: Capture an image

[1904] Input: Surrounding image (real-time)

[1905] How it works: The device uses the smartphone camera to capture an image of the user's surroundings every second.

[1906] Data processing: Capture as image data

[1907] Output: Captured image data

[1908] Step 2: Sending images

[1909] Input: Captured image data

[1910] How it works: The device sends captured image data to the server in real time.

[1911] Data processing: data encoding and transmission processing

[1912] Output: Image data sent to the server

[1913] Step 3: Image analysis

[1914] Input: Received image data

[1915] How it works: The server decodes the received image data and uses TensorFlow to run image recognition algorithms to detect obstacles in the image, distance, signal status, and the presence or absence of walls.

[1916] Data calculation: Analysis using image recognition algorithms

[1917] Output: Obstacle detection results and analysis data

[1918] Step 4: Generate information

[1919] Input: Image analysis results

[1920] Operation: Based on the analysis results, the server generates information to notify the user (e.g., "There is an obstacle 2 meters ahead").

[1921] Data calculation: Transformation of analysis results and generation of notification messages

[1922] Output: Information message

[1923] Step 5: Notification

[1924] Input: Notification message

[1925] What it does: The server generates a notification message and sends it to the device, which then uses Google Text-to-Speech to notify the user audibly.

[1926] Data operations: message conversion and notification processing

[1927] Output: Synthesized notification

[1928] Blind Accessible Facilities Mapping Feature Processing Steps

[1929] Step 1: Locate your location

[1930] Input: Wi-Fi access point signal strength

[1931] How it works: Your device periodically scans the signal strength of Wi-Fi access points.

[1932] Data Processing: Signal Strength Data Collection

[1933] Output: Scan result data

[1934] Step 2: Submit scan results

[1935] Input: Scan result data

[1936] Operation: The device sends the scan result data to the server.

[1937] Data processing: data encoding and transmission processing

[1938] Output: Scan result data sent to the server

[1939] Step 3: Location analysis

[1940] Input: Scan result data

[1941] How it works: The server compares the scanned data with a map database to determine the user's current location. It uses the Google Maps API.

[1942] Data calculation: Matching signal strength data with map data, and determining location

[1943] Output: Current location information

[1944] Step 4: Route calculation

[1945] Input: Current location

[1946] How it works: The server calculates the route to the destination based on the current location information.

[1947] Data calculation: Execution of route calculation algorithms

[1948] Output: Route information

[1949] Step 5: Directions

[1950] Input: Route information

[1951] How it works: The server sends route information to the device, which then uses Google Text-to-Speech to provide voice directions.

[1952] Data calculation: Route information voice conversion and notification processing

[1953] Output: Speech-synthesized route directions

[1954] Image Recognition Function Processing Steps

[1955] Step 1: Capture an image

[1956] Input: Item held in hand (real-time)

[1957] How it works: A user takes a picture of a product with their smartphone camera.

[1958] Data processing: Capture as image data

[1959] Output: Captured image data

[1960] Step 2: Sending images

[1961] Input: Captured image data

[1962] Action: The device sends the captured image data to the server.

[1963] Data processing: data encoding and transmission processing

[1964] Output: Image data sent to the server

[1965] Step 3: Product Identification

[1966] Input: Received image data

[1967] How it works: The server uses image recognition algorithms to identify the product in the image.

[1968] Data calculation: Running image recognition algorithms

[1969] Output: Product identification results

[1970] Step 4: Information Acquisition

[1971] Input: Product identification result

[1972] Operation: Based on the identification results, the server retrieves detailed product information (name, price, description, etc.) from the Internet.

[1973] Data Calculation: Search and Extraction of Details

[1974] Output: Product details

[1975] Step 5: Information Notification

[1976] Input: Product details

[1977] How it works: The server sends product details to the device, which then uses Google Text-to-Speech to notify the user aloud.

[1978] Data Computing: Information Speech Conversion and Notification Processing

[1979] Output: Speech-synthesized product information

[1980] Process steps for subscriber recognition

[1981] Step 1: Capture audio

[1982] Input: Ambient audio (real-time)

[1983] How it works: Your device captures ambient sound with its microphone.

[1984] Data processing: Capture as audio data

[1985] Output: Captured audio data

[1986] Step 2: Sending audio

[1987] Input: Captured audio data

[1988] What it does: The device sends the captured audio data to the server.

[1989] Data processing: data encoding and transmission processing

[1990] Output: Audio data sent to the server

[1991] Step 3: Voice analysis and person identification

[1992] Input: Received audio data

[1993] How it works: The server analyzes the audio data using a speech recognition algorithm to identify a specific person. The software used is Google Speech-to-Text.

[1994] Data Computing: Speech Recognition Algorithms and Person Identification

[1995] Output: Person identification result

[1996] Step 4: Notification of identification results

[1997] Input: Person identification result

[1998] How it works: The server sends the identification results to the device, which then notifies the user using Google Text-to-Speech.

[1999] Data processing: speech conversion and notification of the recognition results

[2000] Output: Synthesized voice notification of person identification

[2001] Emotion Engine Processing Steps

[2002] Step 1: Capture your voice and facial expressions

[2003] Input: User's voice and facial expressions (real-time)

[2004] How it works: The device captures the user's voice and facial expressions using sensors and microphones.

[2005] Data processing: Capture as voice and facial expression data

[2006] Output: Captured audio and facial expression data

[2007] Step 2: Sending data

[2008] Input: Captured voice and facial expression data

[2009] Action: The device sends the captured data to the server.

[2010] Data processing: data encoding and transmission processing

[2011] Output: Voice and facial expression data sent to the server

[2012] Step 3: Sentiment Analysis

[2013] Input: Received voice and facial expression data

[2014] How it works: The server uses an emotion analysis algorithm to analyze the user's emotional state from the captured data. The software used is Affectiva.

[2015] Data Computing: Running sentiment analysis algorithms and classifying emotional states

[2016] Output: Emotion analysis results

[2017] Step 4: Emotion-based feedback

[2018] Input: Sentiment analysis results

[2019] How it works: The server generates notification information (e.g., "Walk a little slower") based on the analyzed emotional state.

[2020] Data calculation: Transformation of analysis results and generation of notification messages

[2021] Output: Information message

[2022] Step 5: Notification

[2023] Input: Notification message

[2024] What it does: The server generates a notification message and sends it to the device, which then uses Google Text-to-Speech to notify the user audibly.

[2025] Data operations: message conversion and notification processing

[2026] Output: Synthesized notification

[2027] Through the above processing steps, the visually impaired person support system of the present invention provides the user with a safe and comfortable life.

[2028] (Application example 2)

[2029] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[2030] Visually impaired people face many challenges in their daily lives. In particular, shopping and traveling in brick-and-mortar stores are often extremely inconvenient and dangerous for them. They also have difficulty recognizing obstacles, identifying specific people, and obtaining product information, which limits their independent lifestyles. Effective support systems that can solve these problems and enable visually impaired people to live their daily lives with peace of mind are needed.

[2031] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an imaging means for capturing images of the surroundings, an image analysis means for analyzing the images captured by the imaging means and detecting obstacles, a notification means for providing a voice notification based on information detected by the image analysis means, a registered person recognition means for identifying registered people and providing a voice notification based on the identification result, and an emotion analysis and notification means for analyzing the user's emotional state and providing feedback based on the emotional state. This allows visually impaired people to move safely while avoiding obstacles, to identify specific people to communicate smoothly, and to receive support according to their emotional state.

[2032] 1. "Imaging means" means a device for capturing images of the surroundings, including a camera or sensor.

[2033] 2. "Image Analysis Means" means the algorithms or processes that analyze captured images to detect obstacles and extract other information.

[2034] 3. "Notification means" refers to an interface for notifying the user of necessary information by voice or other means.

[2035] 4. "Walking navigation means" is a function that provides real-time voice navigation to obstacles.

[2036] 5. "Means for recognizing registered persons" refers to a function that identifies registered persons and sends notifications based on the results of that identification.

[2037] 6. "Emotion analysis and notification means" means a means for analyzing the user's emotional state and providing feedback according to that state in the form of voice or other means.

[2038] 7. "Location identification means" is a function in which the image analysis means analyzes location information and identifies the user's current location.

[2039] 8. "Store guidance means" is a function that provides audio guidance to the user about each area within a physical store.

[2040] This invention provides a support system that enables visually impaired people to navigate and shop safely in brick-and-mortar stores. The system has the following main functions:

[2041] System configuration

[2042] 1. Imaging Method

[2043] It uses smart glasses or a smartphone camera to capture images of the surroundings.

[2044] Example: As a user walks through a physical store, a camera captures images of the user's surroundings at regular intervals.

[2045] 2. Image analysis methods

[2046] The captured image is sent to the server, where it is analyzed in real time.

[2047] The software used includes image analysis algorithms and TensorFlow. The server recognizes obstacles and products and sends that information to the notification means.

[2048] Example: If an obstacle is detected, the analysis result will be generated as "There is an obstacle 2 meters ahead."

[2049] 3. Means of notification

[2050] Based on the analysis results, the information is notified to the user by voice. Voice synthesis uses a voice engine such as pyttsx3.

[2051] Example: A speech synthesis engine informs the user, "There is an obstacle 2 meters ahead."

[2052] 4. Walking navigation methods

[2053] Real-time voice navigation is provided while the user is moving, which is achieved by combining imaging means, image analysis means, and notification means.

[2054] Example: When a user arrives at a specific area (e.g., fresh produce), they are notified, "You have arrived at the fish section."

[2055] 5. Registered User Identification Method

[2056] It captures surrounding sounds and identifies specific people, using the microphone in smart glasses or smartphones to record the sounds.

[2057] The server performs voice analysis and notifies the user of the results of recognizing a specific person.

[2058] Example: When a person in front of the user speaks, the user is notified that "This is a store clerk."

[2059] 6. Sentiment Analysis and Notification Methods

[2060] It captures the user's voice and facial expressions and analyzes their emotional state using an emotion analysis model powered by TensorFlow.

[2061] Appropriate feedback is provided via voice based on the analysis results.

[2062] For example, if the user is analyzed as feeling stressed, they will be notified to "relax and take a deep breath."

[2063] Specific examples

[2064] Consider a scenario in which the system would work effectively when a visually impaired person visits a supermarket. The user walks through various sections of the store, and the smart glasses provide real-time information about their surroundings. For example, when the user scans an item they have picked up with the camera, they are notified by voice, "This is an apple. The price is 150 yen." When they arrive at a specific area, they are informed, "You have arrived at the fish section."

[2065] Prompt Sentence Examples

[2066] An example of a specific prompt sentence for the generative AI model is shown below.

[2067] A smart glasses application for visually impaired people in shopping malls provides walking navigation, product recognition, staff recognition, and emotional feedback. As an example, when a user scans an item they have picked up, the application asks, "How can the application recognize the product from the image captured by the camera and announce its name, price, and description by voice?"

[2068] This will enable visually impaired people to live their daily lives more safely and independently.

[2069] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2070] Step 1:

[2071] The device uses an imaging means (camera) to capture images of the surroundings. The imaging means captures images from the camera at regular intervals (e.g., every second) according to the user's viewpoint, and sends the image data to a server. The input is the image captured from the camera, and the output is the image data sent to the server.

[2072] Step 2:

[2073] The server analyzes the image data it receives. The image analysis means uses TensorFlow's image recognition algorithm to identify obstacles, products, and people in the image. The input is the image data sent to the server, and the output is the analysis results (obstacle position, product information, and person information). Specifically, the server extracts features from the image and performs analysis using a pre-trained model.

[2074] Step 3:

[2075] The server generates the information to be notified based on the analysis results. The notification means converts the analysis results (e.g., "There is an obstacle 2 meters ahead") into voice using a speech synthesis engine (pyttsx3) and notifies the user. The input is the analysis results from the image analysis means, and the output is a voice notification. This is achieved by the server converting the analysis results into text and sending it to the speech synthesis engine.

[2076] Step 4:

[2077] When the user takes a specific action (e.g., picking up a product), the image recognition function is activated. The product is photographed with the device's camera and the image is sent to the server. The input is the image of the product, and the output is the data sent to the server. The server identifies the product, obtains product information (name, price, etc.), and notifies the user by voice.

[2078] Step 5:

[2079] When the user moves, the walking navigation means operates. The terminal captures images in real time using the imaging means and sends them to the server. The server analyzes obstacle information and notifies the user of navigation information. The input is image data captured in real time, and the output is navigation information. The server processes the data in real time and generates navigation information.

[2080] Step 6:

[2081] The device's microphone captures surrounding sounds and activates the registered user recognition means. The voice data is sent to the server, which then uses a voice recognition algorithm to identify the person. The input is the captured voice data, and the output is specific person information. The server notifies the user of the identification result through a voice synthesis engine.

[2082] Step 7:

[2083] The terminal captures the user's voice and facial expressions, and the emotion analysis and notification means are activated. The server uses an emotion analysis algorithm to analyze the emotional state and provide appropriate feedback. The input is the captured voice and facial expression data, and the output is feedback information based on the emotional state. The server analyzes the emotional state and generates information to be notified.

[2084] Step 8:

[2085] When a user goes shopping in a physical store, the store guidance means is activated. The server identifies the user's current location, calculates the route to the destination, and provides voice guidance. The input is data obtained by scanning the signal strength of wireless LAN access points, and the output is route guidance information. The server sends the route guidance information to a speech synthesis engine, which notifies the user by voice.

[2086] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[2087] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2088] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[2089] [Fourth embodiment]

[2090] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[2091] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[2092] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[2093] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[2094] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[2095] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[2096] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[2097] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[2098] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[2099] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[2100] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[2101] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[2102] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2103] The visually impaired support system of the present invention has a plurality of functions for solving various problems that users face in their daily lives. The embodiments of the present invention will be described in detail below.

[2104] Walking navigation function

[2105] 1. Image capture

[2106] The device captures images of the user's surroundings with its camera at regular intervals (e.g., every second).

[2107] The captured images are sent to a server in real time.

[2108] 2. Image Analysis

[2109] The server decodes the received image and runs an image recognition algorithm on it.

[2110] The server detects obstacles in the image, the distance to the obstacles, the signal status, and the presence or absence of walls.

[2111] 3. Information Generation and Notification

[2112] Based on the image analysis results, the server generates information to be notified to the user (e.g., "There is an obstacle 2 meters ahead").

[2113] The server sends the generated information to the terminal, which notifies the user using a voice synthesis engine.

[2114] Example: While the user is walking, the device announces, "There is an obstacle 2 meters ahead."

[2115] Mapping function for facilities accessible to the visually impaired

[2116] 1. Identifying your current location

[2117] The device periodically scans the signal strength of Wi-Fi access points in its current location.

[2118] The scan results are sent to the server.

[2119] 2. Location analysis and route guidance

[2120] The server compares the user's current location with a map database based on the signal from the wireless LAN access point.

[2121] The server determines the user's current location and calculates a route to the destination.

[2122] The server transmits route information to the terminal, and the terminal uses a voice synthesis engine to provide route guidance to the user.

[2123] Example: When a user enters the voice command "I want to go to the library," the device will provide voice guidance such as "Turn left and go 20 meters, then turn right."

[2124] Image Recognition Function

[2125] 1. Image capture

[2126] The user takes a photo of the product they are holding with the device's camera.

[2127] The captured image is sent to a server.

[2128] 2. Product Identification and Information Notification

[2129] The server uses an image recognition algorithm to identify the product in the image.

[2130] Based on the identification results, the server obtains detailed product information (name, price, description, etc.) from the Internet.

[2131] The server sends the acquired detailed product information to the terminal, which then notifies the user using a voice synthesis engine.

[2132] Example: When a user picks up an item and takes a picture of it with the camera, the device announces in a voice message, "This is the 347-page novel 'The Blue Bird' and costs 1,200 yen."

[2133] Registered user recognition function

[2134] 1. Audio capture

[2135] The device captures surrounding sounds using a microphone.

[2136] The captured audio data is sent to a server.

[2137] 2. Voice analysis and identity notification

[2138] The server uses a voice recognition algorithm to analyze the voice data and identify a specific person.

[2139] The server transmits the identification result to the terminal, and the terminal notifies the user using a voice synthesis engine.

[2140] Example: When a person in front of the user speaks, the device will announce in voice, "This is Tanaka-san."

[2141] The objective of the support system for visually impaired people of the present invention is to integrate these functions to efficiently and effectively solve the various problems that visually impaired people encounter in their daily lives, thereby enabling them to live safer and more independent lives.

[2142] The processing flow will be explained below.

[2143] Walking navigation function

[2144] Step 1:

[2145] The device captures images of the user's surroundings at regular intervals using a camera.

[2146] Step 2:

[2147] The device compresses the captured images and sends them to the server in real time.

[2148] Step 3:

[2149] The server decodes the received image and runs an image recognition algorithm on it.

[2150] Step 4:

[2151] The server detects obstacles in the image, the distance to the obstacles, the signal status, and the presence or absence of walls.

[2152] Step 5:

[2153] The server generates the information to be notified based on the results of image analysis (e.g., "There is an obstacle 2 meters ahead").

[2154] Step 6:

[2155] The server transmits the generated information to the terminal.

[2156] Step 7:

[2157] The information received by the terminal is passed to a speech synthesis engine to generate voice data.

[2158] Step 8:

[2159] The audio data generated by the terminal is notified to the user through the speaker.

[2160] Mapping function for facilities accessible to the visually impaired

[2161] Step 1:

[2162] The device scans the signal strength of wireless LAN access points.

[2163] Step 2:

[2164] The device sends the scan results to the server.

[2165] Step 3:

[2166] The server identifies the user's current location based on the signal from the wireless LAN access point by comparing it with a location information database.

[2167] Step 4:

[2168] The server determines the user's current location.

[2169] Step 5:

[2170] The server calculates the route to the destination based on the user's instructions (voice commands).

[2171] Step 6:

[2172] The server sends the route information to the terminal.

[2173] Step 7:

[2174] The route information received by the terminal is passed to a voice synthesis engine to generate voice data.

[2175] Step 8:

[2176] The voice guidance generated by the terminal is notified to the user through the speaker.

[2177] Image Recognition Function

[2178] Step 1:

[2179] The user takes a photo of the product they are holding with the device's camera.

[2180] Step 2:

[2181] The device sends the captured image to the server.

[2182] Step 3:

[2183] The server uses an image recognition algorithm to identify the product in the image.

[2184] Step 4:

[2185] Based on the product name identified by the server, detailed product information (such as name, price, and description) is obtained from the Internet.

[2186] Step 5:

[2187] The server organizes the detailed product information it has acquired and generates text data.

[2188] Step 6:

[2189] The server transmits the generated text data to the terminal.

[2190] Step 7:

[2191] The information received by the terminal is passed to a speech synthesis engine to generate voice data.

[2192] Step 8:

[2193] The audio data generated by the terminal is notified to the user through the speaker.

[2194] Registered user recognition function

[2195] Step 1:

[2196] The device captures surrounding sounds using a microphone.

[2197] Step 2:

[2198] The device sends the captured audio data to the server.

[2199] Step 3:

[2200] The server analyzes the voice data using a voice recognition algorithm.

[2201] Step 4:

[2202] The server compares the analysis results with a pre-registered voice database to identify the person.

[2203] Step 5:

[2204] The server generates the identified information as text data (e.g., "This is Mr. Tanaka").

[2205] Step 6:

[2206] The server transmits the generated text data to the terminal.

[2207] Step 7:

[2208] The information received by the terminal is passed to a speech synthesis engine to generate voice data.

[2209] Step 8:

[2210] The audio data generated by the terminal is notified to the user through the speaker.

[2211] Example 1

[2212] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2213] To solve the various challenges that visually impaired people face in their daily lives, it is necessary to properly obtain information about their surroundings and convey that information to them quickly and accurately. However, current systems lack comprehensive support that goes beyond simply detecting obstacles, covering things like determining their current location, identifying people, and even providing traffic light status and specific route guidance. This leaves visually impaired people at great risk.

[2214] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[2215] In this invention, the server includes an imaging means for capturing an image of the surroundings, an image analysis means for analyzing the image captured by the imaging means and detecting obstacles, traffic light status, and the presence or absence of walls, an information generation means for generating notification information based on the information detected by the image analysis means, a notification means for notifying by voice the notification information generated by the information generation means, a location identification means for analyzing location information using a wireless LAN signal and identifying the current location, an audio recording means for capturing surrounding sounds, an audio analysis means for analyzing the sounds captured by the audio recording means and identifying a specific person, an information generation means for generating notification information based on the person information identified by the audio analysis means, and a notification means for notifying by voice the notification information generated by the information generation means.This enables visually impaired people to avoid obstacles while walking, identify their current location, and easily identify people, enabling them to live safe and independent lives.

[2216] An "imaging means" is a device or sensor for capturing an image of the surroundings.

[2217] "Image analysis means" refers to a set of algorithms and software for analyzing captured images and detecting obstacles, traffic light conditions, the presence or absence of walls, etc.

[2218] The "notification means" is a device or software that conveys the analyzed information to the user by voice.

[2219] The "information generation means" refers to logic or algorithms for generating information to be notified to the user based on the analysis results.

[2220] The "location determination means" is a device or software for determining the user's current location using a wireless LAN signal.

[2221] "Audio recording means" refers to a microphone or other audio recording device for capturing ambient audio.

[2222] "Audio analysis means" refers to algorithms or software that analyzes captured audio and identifies specific individuals.

[2223] The visually impaired support system of the present invention has a plurality of functions for solving various problems that users face in their daily lives. The embodiments of the present invention will be described in detail below.

[2224] Walking navigation function

[2225] Image capture

[2226] The device captures images of the user's surroundings at regular intervals (e.g., every second) using the built-in camera of a smartphone.

[2227] Sending images

[2228] The device immediately sends the captured image to the server using Wi-Fi or mobile data.

[2229] Image analysis

[2230] The server decodes the received images and uses image recognition algorithms such as TensorFlow and OpenCV to detect obstacles, traffic light status, and the presence or absence of walls.

[2231] Information Generation and Notification

[2232] Based on the analysis results, the server generates notification information such as "There is an obstacle 2 meters ahead." The information is sent from the server to the device, which then notifies the user using a voice synthesis engine such as Google Text-to-Speech.

[2233] Specific examples

[2234] While the user is walking, the device will announce to them by voice, "There is an obstacle 2 meters ahead."

[2235] Mapping function for facilities accessible to the visually impaired

[2236] Scan your current location

[2237] The device periodically scans the signal strength of surrounding Wi-Fi access points and sends the results of the scan to the server.

[2238] Location analysis and routing

[2239] The server uses signal data from wireless LAN access points to identify the user's current location using a map database such as Google Maps API and calculates the optimal route. The route information is sent to the device, which then uses a voice synthesis engine such as Google Text-to-Speech to provide route guidance to the user.

[2240] Specific examples

[2241] When a user inputs the voice command "I want to go to the library," the device will provide voice guidance such as "Turn left and go 20 meters, then turn right."

[2242] Image Recognition Function

[2243] Image capture

[2244] The user takes a photo of the product they are holding with the device's camera, and the captured image is sent to the server.

[2245] Product identification and information acquisition

[2246] The server identifies the product using an image recognition algorithm such as YOLOv3, and then uses the Amazon API or Rakuten API to obtain detailed product information (name, price, description, etc.).

[2247] Information Notification

[2248] The acquired product details are sent from the server to the terminal, which then notifies the user using a voice synthesis engine such as Google Text-to-Speech.

[2249] Specific examples

[2250] When the user picks up the product and takes a picture of it with the camera, the device announces in voice, "This is the 347-page novel 'The Blue Bird' and costs 1,200 yen."

[2251] Registered user recognition function

[2252] Audio capture

[2253] The device captures the surrounding sound with a microphone, and the captured sound data is sent to the server.

[2254] Audio analysis and notifications

[2255] The server analyzes the voice data using a speech recognition algorithm such as IBM Watson Speech to Text to identify a specific person, and the identification result is sent from the server to the device, which then notifies the user using a speech synthesis engine such as Google Text-to-Speech.

[2256] Specific examples

[2257] When a person in front of the user speaks, the device will announce aloud, "This is Tanaka-san."

[2258] Examples of prompt statements

[2259] 1. Walking navigation function

[2260] Prompt: "What is the process for the system that notifies you of nearby obstacles?"

[2261] 2. Mapping function for facilities accessible to the visually impaired

[2262] Prompt: "Please explain the process of the system that identifies your current location and provides directions to your destination."

[2263] 3. Image Recognition

[2264] Prompt: "Please tell me the process of the system that takes a picture of a product with a camera and provides detailed information about it."

[2265] 4. Registered User Recognition Function

[2266] Prompt: "Please explain the process for a system that analyzes surrounding sounds and identifies a specific person."

[2267] The above is an embodiment of the visually impaired support system of the present invention, and this system enables visually impaired people to live safer and more independent lives.

[2268] The flow of the identification process in the first embodiment will be described with reference to FIG.

[2269] Walking navigation function

[2270] Step 1:

[2271] Image capture

[2272] The device captures images of the user's surroundings with a camera at regular intervals (every second). The input is image data obtained from the camera, and the output is the captured image file.

[2273] Specific operation: The device automatically launches the camera app and simulates pressing the shutter button every second.

[2274] Step 2:

[2275] Sending images

[2276] The image files captured by the device are sent to the server in real time. The input is the image file generated in step 1, and the output is the upload of image data to the server.

[2277] Specific operation: The device compresses the image data and uploads it to the server using the HTTPS protocol.

[2278] Step 3:

[2279] Image analysis

[2280] The server decodes the received image and analyzes it using an image recognition algorithm using TensorFlow, OpenCV, etc. The input is the image data sent to the server, and the output is the detection results such as obstacles, traffic light status, and the presence or absence of walls.

[2281] Specific operation: The server processes the received image, runs a recognition algorithm, and generates an analysis result.

[2282] Step 4:

[2283] information generation

[2284] The server generates information to be notified to the user based on the analysis results. The input is the analysis results obtained in step 3, and the output is a notification message.

[2285] What happens: The server uses a text generation algorithm to generate a message such as "There is an obstacle 2 meters ahead."

[2286] Step 5:

[2287] Information Notification

[2288] The server sends the generated notification message to the device, and the device uses a speech synthesis engine such as Google Text-to-Speech to notify the user audibly. The input is the generated notification message, and the output is the audio notification delivered to the user.

[2289] Specific operation: The device converts the received message into voice and notifies the user through the speaker, "There is an obstacle 2 meters ahead."

[2290] Mapping function for facilities accessible to the visually impaired

[2291] Step 1:

[2292] Scan your current location

[2293] The device periodically scans the signal strength of surrounding Wi-Fi access points. The input is the Wi-Fi signal captured by the Wi-Fi module, and the output is the signal strength data.

[2294] Specific operation: The device activates the Wi-Fi module and collects the surrounding network signal strength.

[2295] Step 2:

[2296] Sending scan results

[2297] The terminal sends the collected signal strength data to the server. The input is the signal strength data collected in step 1, and the output is the upload to the server.

[2298] Specific operation: The terminal compresses the signal data and uploads it to the server via HTTPS protocol.

[2299] Step 3:

[2300] Location analysis

[2301] The server analyzes the location information based on the signal data of the wireless LAN access point and determines the current location. The input is the transmitted signal strength data, and the output is the determined current location information.

[2302] What it does: The server analyzes the signal data and locates it by comparing it with a map database (e.g., Google Maps API).

[2303] Step 4:

[2304] Route calculation

[2305] The server calculates the optimal route from the specified current location to the destination. The input is the current location information and the destination information, and the output is the route information.

[2306] Specific operation: The server executes a route calculation algorithm (e.g., A search algorithm) to calculate the route to the destination.

[2307] Step 5:

[2308] Route guidance

[2309] The server sends route information to the device, which then uses a speech synthesis engine such as Google Text-to-Speech to provide route guidance to the user. The input is the calculated route information, and the output is voice guidance provided to the user.

[2310] Specific operation: The device converts the route information it receives into voice and notifies the user, "Turn left, go 20 meters, then turn right."

[2311] Image Recognition Function

[2312] Step 1:

[2313] Image capture

[2314] The user takes a photo of the product they are holding with the device's camera. The input is image data acquired from the camera, and the output is the captured image file.

[2315] Specific operation: The device automatically launches the camera app, and the user presses the shutter button.

[2316] Step 2:

[2317] Sending images

[2318] The device sends the captured image file to the server. The input is the captured image file, and the output is the upload of the image data to the server.

[2319] Specific operation: The device compresses the image data and uploads it to the server using the HTTPS protocol.

[2320] Step 3:

[2321] Product Identification

[2322] The server uses an image recognition algorithm (e.g., YOLOv3) to identify the product in the image. The input is the submitted image data, and the output is the identified product information.

[2323] Specific operation: The server analyzes the received image and identifies it as "This is a novel called 'The Blue Bird'."

[2324] Step 4:

[2325] Information acquisition

[2326] The server retrieves detailed product information (such as name, price, description, etc.) from the Internet. The input is the identified product information, and the output is the retrieved product detail data.

[2327] Specific operation: The server calls the API to obtain product data.

[2328] Step 5:

[2329] Information Notification

[2330] The server sends the acquired product details to the terminal, which then notifies the user using a speech synthesis engine such as Google Text-to-Speech. The input is the acquired product details data, and the output is a voice notification delivered to the user.

[2331] Specific operation: The device converts the received product information into voice and notifies the user, "This is the 347-page novel 'The Blue Bird', priced at 1,200 yen."

[2332] Registered user recognition function

[2333] Step 1:

[2334] Audio capture

[2335] The device captures the surrounding sound with a microphone. The input is the audio data collected by the microphone, and the output is the captured audio file.

[2336] Specific operation: The device activates the microphone and collects surrounding sounds.

[2337] Step 2:

[2338] Sending Audio

[2339] The device sends the captured audio file to the server. The input is the captured audio file, and the output is the audio data uploaded to the server.

[2340] Specific operation: The device compresses the audio data and uploads it to the server using the HTTPS protocol.

[2341] Step 3:

[2342] Audio analysis

[2343] The server analyzes the voice data using a speech recognition algorithm (e.g., IBM Watson Speech to Text) to identify a specific person. The input is the transmitted voice data, and the output is the identified person's information.

[2344] Specific operation: The server processes the voice data and generates an identification result such as "This is Mr. Tanaka."

[2345] Step 4:

[2346] Generating the identification results

[2347] The server generates notification information based on the identification result. The input is the identified person's information, and the output is the generated notification message.

[2348] Specific operation: The server generates a text message and sends it to the device.

[2349] Step 5:

[2350] Notification of identity

[2351] The device uses a speech synthesis engine such as Google Text-to-Speech to convert the notification message into speech and notify the user. The input is the generated notification message, and the output is the audio notification delivered to the user.

[2352] Specific operation: The device uses the speaker to announce, "This is Tanaka-san."

[2353] (Application example 1)

[2354] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2355] Due to their unique needs, it is difficult for visually impaired people to shop safely and comfortably in brick-and-mortar stores. To solve this problem, a system is needed that allows visually impaired people to independently select and navigate products in brick-and-mortar stores. In addition, since it is difficult for them to accurately grasp product information in real time, appropriate support is required.

[2356] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[2357] In this invention, the server includes an imaging means for capturing an image of the surroundings, an image analysis means for analyzing the image captured by the imaging means and detecting obstacles, a notification means for providing a voice notification based on information detected by the image analysis means, an image recognition means for identifying products, an information acquisition means for acquiring detailed information about the identified products, and a notification means for providing a voice notification based on information acquired by the information acquisition means. This enables visually impaired people to receive product information by voice while safely moving around in a physical store.

[2358] An "imaging means" is a device that captures an image of the surroundings.

[2359] The "image analysis means" is a device that analyzes the image captured by the imaging means and detects objects such as obstacles.

[2360] The "notification means" is a device that notifies the user of information detected by the image analysis means and other information by voice.

[2361] An "image recognition means" is a device for identifying an item in a captured image.

[2362] The "information acquisition means" is a device that acquires detailed information (such as the name and price) of the identified product from the Internet or the like.

[2363] The "location specifying means" is a device that specifies the current location by analyzing the location information performed by the image analysis means.

[2364] An "audio recording means" is a device that captures ambient sounds.

[2365] The "voice analysis means" is a device that analyzes the voice captured by the voice recording means and identifies a specific person.

[2366] The present invention relates to a support system for visually impaired people to shop safely and comfortably in a brick-and-mortar store. Hereinafter, an embodiment of the present invention will be specifically described.

[2367] System Configuration

[2368] The system includes the following main components:

[2369] 1. An imaging means to capture images of the surroundings

[2370] 2. Image analysis means for analyzing images captured by the imaging means and detecting obstacles

[2371] 3. Notification means that notifies by voice based on information detected by image analysis means

[2372] 4. Image recognition methods for identifying products

[2373] 5. Information acquisition means for acquiring detailed information on identified products

[2374] 6. Notification method for notifying information by voice

[2375] Hardware and Software

[2376] Hardware:

[2377] Smartphone, smart glasses, or head-mounted display

[2378] Web camera (imaging means)

[2379] Microphone (audio recording means)

[2380] Speaker (notification means)

[2381] software:

[2382] OpenCV (library for image capture and analysis)

[2383] Google Text-to-Speech (gTTS) (voice notifications)

[2384] Image recognition API (product identification)

[2385] Internet connection (information acquisition)

[2386] Processing steps

[2387] When a user activates the system, the imaging means captures images of the surrounding area at regular intervals. The captured images are sent to a server in real time, and the image analysis means detects obstacles in the images and analyzes their positions. The notification means then notifies the user of the obstacle information by voice.

[2388] At the same time, if the user wants to identify a product, they take a photo of the product using a smartphone or smart glasses. The captured image is sent to a server, and the image recognition means identifies the product. Detailed information about the identified product is obtained from the Internet by an information acquisition means, and the information is notified to the user by voice via a notification means.

[2389] Specific examples

[2390] For example, when a user is walking in a shopping mall, the system will notify them that there is an obstacle two meters ahead, helping them to walk safely. Also, when a user picks up a product and takes a picture of it with the camera, the system will announce the product information in a voice message, saying, "This is a 347-page book, priced at 1,200 yen."

[2391] Prompt Sentence Examples

[2392] Example prompt for image recognition:

[2393] "Please recognize the item in this image and provide the item name and distance."

[2394] Example of a prompt for audio notification:

[2395] "2 meters ahead, there is an obstacle. It is a table."

[2396] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[2397] Step 1:

[2398] The imaging means captures images of the surroundings at regular intervals (e.g., every second). The input is video data from the camera, which is captured and saved as image data. The output is the saved image data.

[2399] Step 2:

[2400] The device sends the captured image data to the server. The input is the image data obtained in step 1, which is uploaded to the server via the Internet. The output is the image data received by the server.

[2401] Step 3:

[2402] The server analyzes the image data it receives and identifies the presence and location of obstacles. The input is the image data received by the server, which is analyzed using an image analysis algorithm (e.g., OpenCV). The output is the location information of obstacles.

[2403] Step 4:

[2404] The server generates the content to be notified to the user based on the analysis results. The input is the obstacle location information obtained in step 3, and an appropriate notification message is created using the generative AI model. The output is the notification message (text data).

[2405] Step 5:

[2406] The server sends a notification message to the device, and the device notifies the user by voice. The input is the notification message created in step 4, which is input to a speech synthesis engine (e.g., gTTS) to generate voice data and play it through the speaker. The output is a voice notification.

[2407] Step 6:

[2408] When a user wants to identify a product, they take a picture of the product with the camera on their device. The input is the video data from the camera, which is captured and saved as image data. The output is the saved image data.

[2409] Step 7:

[2410] The device sends the captured product image to the server. The input is the image data obtained in step 6, which is uploaded to the server via the Internet. The output is the image data received by the server.

[2411] Step 8:

[2412] The server analyzes the product image and identifies the product. The input is the image data obtained in step 7, and the product is identified using an image recognition algorithm (e.g., image recognition API). The output is the identified product data (e.g., product name, price).

[2413] Step 9:

[2414] The server acquires detailed information based on the identified product data. The input is the product data acquired in step 8, and detailed information is collected from the Internet, etc. using an information acquisition means. The output is detailed product information.

[2415] Step 10:

[2416] The server sends the product information to the terminal, which then notifies the user by voice. The input is the detailed product information obtained in step 9, which is input into a speech synthesis engine (e.g., gTTS) to generate voice data and play it back through the speaker. The output is a voice notification of the product information.

[2417] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[2418] The support system for visually impaired people of the present invention has multiple functions for solving various problems that users face in their daily lives. This system incorporates a walking navigation function, a facility mapping function for visually impaired people, an image recognition function, a registered user recognition function, and an emotion engine that recognizes the user's emotions. An embodiment of the present invention will be described in detail below.

[2419] Walking navigation function

[2420] 1. Image capture

[2421] The device captures images of the user's surroundings with its camera at regular intervals (e.g., every second).

[2422] The captured images are sent to a server in real time.

[2423] 2. Image Analysis

[2424] The server decodes the received image and runs an image recognition algorithm on it.

[2425] The server detects obstacles in the image, the distance to the obstacles, the signal status, and the presence or absence of walls.

[2426] 3. Information Generation and Notification

[2427] Based on the image analysis results, the server generates information to be notified to the user (e.g., "There is an obstacle 2 meters ahead").

[2428] The server sends the generated information to the terminal, which notifies the user using a voice synthesis engine.

[2429] Example: While the user is walking, the device announces, "There is an obstacle 2 meters ahead."

[2430] Mapping function for facilities accessible to the visually impaired

[2431] 1. Identifying your current location

[2432] The device periodically scans the signal strength of Wi-Fi access points in its current location.

[2433] The scan results are sent to the server.

[2434] 2. Location analysis and route guidance

[2435] The server compares the user's current location with a map database based on the signal from the wireless LAN access point.

[2436] The server determines the user's current location and calculates a route to the destination.

[2437] The server transmits route information to the terminal, and the terminal uses a voice synthesis engine to provide route guidance to the user.

[2438] Example: When a user enters the voice command "I want to go to the library," the device will provide voice guidance such as "Turn left and go 20 meters, then turn right."

[2439] Image Recognition Function

[2440] 1. Image capture

[2441] The user takes a photo of the product they are holding with the device's camera.

[2442] The captured image is sent to a server.

[2443] 2. Product Identification and Information Notification

[2444] The server uses an image recognition algorithm to identify the product in the image.

[2445] Based on the identification results, the server obtains detailed product information (name, price, description, etc.) from the Internet.

[2446] The server sends the acquired detailed product information to the terminal, which then notifies the user using a voice synthesis engine.

[2447] Example: When a user picks up an item and takes a picture of it with the camera, the device announces in a voice message, "This is the 347-page novel 'The Blue Bird' and costs 1,200 yen."

[2448] Registered user recognition function

[2449] 1. Audio capture

[2450] The device captures surrounding sounds using a microphone.

[2451] The captured audio data is sent to a server.

[2452] 2. Voice analysis and identity notification

[2453] The server uses a voice recognition algorithm to analyze the voice data and identify a specific person.

[2454] The server transmits the identification result to the terminal, and the terminal notifies the user using a voice synthesis engine.

[2455] Example: When a person in front of the user speaks, the device will announce in voice, "This is Tanaka-san."

[2456] Emotion Engine

[2457] 1. Voice and facial expression capture

[2458] The device captures the user's voice and facial expressions using sensors and microphones.

[2459] The captured data is sent to a server.

[2460] 2. Emotion analysis

[2461] The server uses an emotion engine to parse the user's emotional state from the captured data.

[2462] The analysis results are classified as emotional states such as "joy," "sadness," and "stress."

[2463] 3. Emotion-based feedback

[2464] The server generates notification information based on the analyzed emotional state (e.g., "Walk a little more slowly").

[2465] The server sends the generated information to the terminal, which notifies the user using a voice synthesis engine.

[2466] Example: If the device determines that the user is feeling stressed, it will provide voice advice such as "Relax and take a deep breath."

[2467] By integrating the above functions, the support system for visually impaired people of the present invention can provide support for efficiently and effectively resolving the various problems that visually impaired people encounter in their daily lives. In addition, by incorporating an emotion engine, it can provide support that is adapted to the user's emotional state, allowing visually impaired people to live their lives with greater peace of mind.

[2468] The processing flow will be explained below.

[2469] Emotion Engine

[2470] Step 1:

[2471] The device captures the user's voice and facial expressions at regular intervals using sensors and microphones. The sampling period for audio and video data is set appropriately (e.g., every second).

[2472] Step 2:

[2473] The captured voice and facial expression data is transmitted to a server in real time.

[2474] Step 3:

[2475] The server decodes the received voice and facial expression data and runs emotion recognition algorithms.

[2476] Step 4:

[2477] Using an emotion recognition algorithm, the server analyzes the user's emotional state (happiness, sadness, anger, stress, etc.) and assigns a score to each emotion.

[2478] Step 5:

[2479] The server generates a feedback message that corresponds to the emotional state (e.g., "Take a break" or "Relax and take a deep breath").

[2480] Step 6:

[2481] The server sends the generated feedback message to the terminal.

[2482] Step 7:

[2483] The terminal passes the received feedback message to a speech synthesis engine to generate speech data.

[2484] Step 8:

[2485] The audio data generated by the terminal is notified to the user through the speaker.

[2486] Walking navigation function

[2487] Step 1:

[2488] The device captures images of the user's surroundings at regular intervals using a camera.

[2489] Step 2:

[2490] The captured images are compressed and sent to a server in real time.

[2491] Step 3:

[2492] The server decodes the received image and runs an image recognition algorithm on it.

[2493] Step 4:

[2494] The server detects obstacles in the image, the distance to the obstacles, the signal status, and the presence or absence of walls.

[2495] Step 5:

[2496] The server generates the information to be notified based on the results of image analysis (e.g., "There is an obstacle 2 meters ahead").

[2497] Step 6:

[2498] The server transmits the generated information to the terminal.

[2499] Step 7:

[2500] The information received by the terminal is passed to a speech synthesis engine to generate voice data.

[2501] Step 8:

[2502] The audio data generated by the terminal is notified to the user through the speaker.

[2503] Mapping function for facilities accessible to the visually impaired

[2504] Step 1:

[2505] The device scans the signal strength of wireless LAN access points.

[2506] Step 2:

[2507] The scan results are sent to the server.

[2508] Step 3:

[2509] The server identifies the user's current location based on the signal from the wireless LAN access point and compares it with a map database.

[2510] Step 4:

[2511] The server determines the user's current location.

[2512] Step 5:

[2513] The server calculates the route to the destination based on the user's instructions (voice commands).

[2514] Step 6:

[2515] The server sends the route information to the terminal.

[2516] Step 7:

[2517] The route information received by the terminal is passed to a voice synthesis engine to generate voice data.

[2518] Step 8:

[2519] The voice guidance generated by the terminal is notified to the user through the speaker.

[2520] Image Recognition Function

[2521] Step 1:

[2522] The user takes a photo of the product they are holding with the device's camera.

[2523] Step 2:

[2524] The captured image is sent to a server.

[2525] Step 3:

[2526] The server uses an image recognition algorithm to identify the product in the image.

[2527] Step 4:

[2528] Based on the product name identified by the server, detailed product information (such as name, price, and description) is obtained from the Internet.

[2529] Step 5:

[2530] The server organizes the detailed product information it has acquired and generates text data.

[2531] Step 6:

[2532] The server transmits the generated text data to the terminal.

[2533] Step 7:

[2534] The information received by the terminal is passed to a speech synthesis engine to generate voice data.

[2535] Step 8:

[2536] The audio data generated by the terminal is notified to the user through the speaker.

[2537] Registered user recognition function

[2538] Step 1:

[2539] The device captures surrounding sounds using a microphone.

[2540] Step 2:

[2541] The captured audio data is sent to a server.

[2542] Step 3:

[2543] The server analyzes the voice data using a voice recognition algorithm.

[2544] Step 4:

[2545] The analysis results are compared with a pre-registered voice database, and the server identifies the person.

[2546] Step 5:

[2547] The server generates the identified information as text data (e.g., "This is Mr. Tanaka").

[2548] Step 6:

[2549] The server transmits the generated text data to the terminal.

[2550] Step 7:

[2551] The information received by the terminal is passed to a speech synthesis engine to generate voice data.

[2552] Step 8:

[2553] The audio data generated by the terminal is notified to the user through the speaker.

[2554] Example 2

[2555] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2556] There is a need for a system that comprehensively resolves the multiple obstacles and difficulties that visually impaired people face in their daily lives. Conventional systems only individually address a wide range of requirements, such as obstacle detection, location identification, person identification, product information acquisition, and even user emotion analysis, and do not provide a unified solution. Furthermore, it is difficult to provide appropriate feedback based on the user's real-time situation and emotions, which limits the effectiveness of these systems in improving the safety and comfort of visually impaired people.

[2557] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[2558] In this invention, the server includes an imaging means for capturing images of the surroundings, an image analysis means for analyzing the images and detecting obstacles, a notification means for providing audio notification based on the detected information, a location identification means for periodically scanning wireless signals to identify the current location, a route guidance means for analyzing voice commands and providing route guidance, a product identification means for capturing images of products and obtaining identification information, a voice analysis means for capturing surrounding voices and identifying specific people, and an emotion analysis means for analyzing the user's emotional state and providing feedback. This makes it possible to provide a system integrating multiple functions to centrally solve various problems in the daily lives of visually impaired people and significantly improve safety and comfort.

[2559] An "imaging means" is a device that captures an image of the surroundings.

[2560] "Image analysis means" is a technology for analyzing captured images and detecting obstacles and their positions.

[2561] The "notification means" is a technology for providing a voice notification based on the analyzed information.

[2562] "Location determination means" is a technology that periodically scans for radio signals to determine the current location.

[2563] "Route guidance means" is a technology for analyzing voice commands and providing guidance on the route to a destination.

[2564] The "product identification means" is a technology for analyzing the captured image of a product and obtaining its identification information.

[2565] "Audio analysis means" is a technology for capturing surrounding sounds and identifying specific people.

[2566] "Emotion analysis means" is a technology for analyzing the user's emotional state and providing feedback based on that state.

[2567] The visually impaired person support system of the present invention is mainly composed of a server, a terminal, and a user. This system is designed to provide multiple integrated functions to enable visually impaired people to spend their daily lives comfortably and safely. Specific embodiments of the present invention will be described in detail below.

[2568] Device Features

[2569] The device has the function of capturing images of the user's surroundings at regular intervals (e.g., every second) using a camera while the user is walking. The hardware used is a smartphone camera, and the software is a camera app. The captured image data is sent to a server in real time. The device also has the function of periodically scanning the signal strength of wireless LAN access points and sending the scan results to the server. It also has the function of capturing images of products held by the user and sending them to the server. As an audio recording function, it also captures surrounding sounds using a microphone.

[2570] Server Features

[2571] The server decodes the received image data and runs an image recognition algorithm. The software used is TensorFlow. The server detects obstacles in the image, the distance to the obstacles, the signal status, and the presence or absence of walls. It also compares the user's current location with a map database based on the signal from the Wi-Fi access point and calculates a route to the destination. The software used is Google Maps API. The server also uses an image recognition algorithm to identify products and retrieve detailed product information (such as name, price, and description) from the Internet. It also has the ability to analyze voice data using a voice recognition algorithm to identify specific people. Finally, an emotion engine is used to analyze the user's emotional state from the captured data. The software used is Affectiva.

[2572] User Notification

[2573] The server generates the notification information (e.g., "There is an obstacle 2 meters ahead") based on the analysis results and sends it to the device. The device then notifies the user using a speech synthesis engine. The software used is Google Text-to-Speech.

[2574] Specific examples

[2575] For example, if a user's smartphone camera captures images in real time while walking and the server detects an obstacle, the device will announce, "There is an obstacle two meters ahead." If the user enters the voice command, "I want to go to the library," the device will provide voice guidance, saying, "Turn left, walk 20 meters, then turn right." If the user picks up an item and takes a picture of it with the camera, the device will announce, "This is a 347-page novel, priced at 1,200 yen." If a person in front of the user speaks, the device will announce, "This is Mr. Tanaka." Finally, if the device analyzes that the user is feeling stressed, it will provide voice advice, saying, "Relax and take a deep breath."

[2576] Example prompts to input to the generative AI model

[2577] 1. Prompt to explain specific examples of walking navigation features:

[2578] "Regarding navigation features for safe walking for the visually impaired, please explain how the system notifies users of obstacles in real time as they walk."

[2579] 2. Prompt to explain the visually impaired accessibility mapping feature:

[2580] "Please explain with specific examples the voice guidance function that allows users to find directions to their destination."

[2581] 3. Prompt to explain specific examples of image recognition features:

[2582] "Please explain how you can provide audio information about an item when the user takes a photo of it."

[2583] 4. Prompt to explain specific examples of subscriber recognition features:

[2584] "Please provide a concrete example of a feature that uses voice to identify and notify the user of specific people in their vicinity."

[2585] 5. Prompt to explain a specific example of an emotion engine:

[2586] "Please give a concrete example of a system that analyzes a user's emotional state a...

Claims

1. imaging means for capturing images of the surroundings; image analysis means for analyzing the image captured by the imaging means and detecting an obstacle; a notification means for providing a voice notification based on the information detected by the image analysis means; Support systems for the visually impaired, including:

2. 2. The visually impaired support system according to claim 1, wherein said image analysis means further comprises location identification means for analyzing location information and identifying a current location.

3. an audio recording means for capturing ambient audio; a voice analysis means for analyzing the voice captured by the voice recording means and identifying a specific person; a notification means for notifying the user by voice based on the person information identified by the voice analysis means; Support systems for the visually impaired, including:

4. image analysis means for analyzing the image of the object held in the hand captured by the imaging means; 2. The visually impaired support system according to claim 1, further comprising notification means for notifying detailed information by voice based on the information of the object analyzed by said image analysis means.

5. a location determination means for scanning signals from wireless LAN access points and determining a current location within a building; a route guidance means for calculating a route to a destination based on the location specification means and providing voice guidance; Support systems for the visually impaired, including:

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A