system

A system for visually impaired individuals provides voice-guided navigation and shopping assistance through voice input conversion, text analysis, real-time video analysis, and product recognition, addressing challenges in route guidance and product assessment.

JP2026063823APending Publication Date: 2026-04-13SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-01
Publication Date
2026-04-13

AI Technical Summary

Technical Problem

Visually impaired individuals face challenges in independently navigating and shopping due to difficulties in route guidance, recognizing products, and assessing product freshness, with a lack of effective means to recognize signals and obstacles during movement.

Method used

A system that integrates voice input conversion, text analysis, map-based route calculation, real-time video analysis for obstacle and traffic light recognition, and product identification with freshness evaluation, providing voice-guided assistance.

Benefits of technology

Enables visually impaired individuals to navigate and shop independently by receiving voice-guided route directions, product recognition, and freshness assessments, enhancing their mobility and shopping independence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026063823000001_ABST
    Figure 2026063823000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for receiving voice input from the user, A means of converting received voice input into text, A means of identifying a destination based on converted text and calculating a route using map information, A means of conveying calculated route information to the user via voice, A means of capturing and analyzing video footage of the user's surroundings, A means of recognizing obstacles and traffic light status in captured video, A means of notifying the user of the recognition result by voice, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , , , , , ,

[0005] , , , , ,

[0003] , , , ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] When visually impaired people move independently or do daily shopping, there are problems that cannot be fully addressed by current support technologies. For example, there are difficulties in providing route guidance for walking safely in the street and in recognizing products and judging freshness when shopping. In addition, there is a lack of means to appropriately recognize signals and obstacles during walking and give instructions based on them. As a result, visually impaired people often need the help of others, and their independence is restricted. It is required to solve these problems so that visually impaired people can live more independently.

Means for Solving the Problems

[0005] This invention provides a system to assist visually impaired individuals in independently moving around and shopping. Specifically, it includes means for receiving voice input and converting it to text, means for identifying a destination based on the converted text and calculating a route using map information, and means for communicating the calculated route information to the user by voice. It also includes means for capturing and analyzing video of the user's surroundings, means for recognizing obstacles and traffic light conditions in the captured video, and means for notifying the user of the recognition results by voice. Furthermore, it includes means for the user to input product names by voice, analyze that information and convert it to text, means for recognizing products from camera footage based on the converted text, means for evaluating the freshness and quality of the recognized products, and means for notifying the user of the evaluation results by voice. This enables visually impaired individuals to independently move around and shop.

[0006] "Means for receiving voice input" refers to a device or software that captures the voice spoken by the user and inputs it into the system.

[0007] "Means for converting speech to text" refers to a device or algorithm for analyzing an input speech signal and converting it into corresponding text data.

[0008] "Means for identifying a destination" refers to a system that analyzes destination information specified by the user and determines its specific geographical location.

[0009] "Means for calculating routes using map information" refers to a device or algorithm that calculates the optimal route from the current location to a specified destination based on map information.

[0010] "Means of conveying route information by voice" refers to a device or program that transmits calculated route information to the user as voice.

[0011] "Means for capturing surrounding video" refers to a mechanism that uses cameras, sensors, etc., to capture video of the user's surroundings in real time.

[0012] "Means for analyzing video" refers to a device or algorithm for analyzing captured video data and recognizing its elements.

[0013] "Means for recognizing obstacles and traffic light conditions" refers to a system that identifies the condition of obstacles and traffic light colors on the road based on the results of video analysis.

[0014] "Means for notifying recognition results by voice" refers to a device or software that notifies the user by voice information regarding the status of recognized obstacles or traffic signals.

[0015] "A means of inputting product names by voice" refers to a device or software that allows a user to input the name of the product they wish to buy by voice.

[0016] "Methods for analyzing product information" refers to a system that converts voice-inputted product names into text data, and then uses that text data to identify the target product.

[0017] "Means for recognizing products from camera footage" refers to a device or algorithm for analyzing camera footage to identify a specific product.

[0018] "Means for evaluating the freshness and quality of a product" refers to a system for determining the freshness and quality of a product based on its perceived appearance and condition.

[0019] "Means for notifying evaluation results by voice" refers to a device or program for communicating evaluation results regarding the freshness and quality of a product to the user by voice. [Brief explanation of the drawing]

[0020] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Mode for Carrying Out the Invention

[0021] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.​​​​​​​In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), and APU (Accelerated Processing Unit).

[0024] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0025] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0026] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0028] [First Embodiment]

[0029] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0030] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0031] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0032] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0033] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0035] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0036] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0037] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0038] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0039] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0040] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0041] System Overview

[0042] This invention is an assistance system for visually impaired individuals to enable them to move around and perform daily shopping independently. This system combines voice input, speech synthesis, image recognition, GPS, map information, and AI technology. The embodiments are described below.

[0043] Navigation function

[0044] Enter destination

[0045] The user voice-inputs "Take me to the nearest convenience store" into their smartphone. The device uses speech recognition technology to convert the voice into text. This text information is sent to a server and processed as destination information.

[0046] Root calculation

[0047] The server calculates a route using map information based on the received text information (destination). Specifically, it uses an API to calculate the optimal walking route from the current location to the destination. The calculated route information is then returned to the device.

[0048] Voice directions

[0049] The terminal analyzes the calculated route information and provides voice instructions to the user regarding the next direction and distance to travel. For example, it might say, "Go 50 meters and turn right," providing step-by-step guidance.

[0050] Real-time video analysis

[0051] The device uses a camera to capture video of the user's surroundings in real time. This video data is sent to a server, which analyzes the video using deep learning technology. Specifically, it recognizes the status of traffic lights and obstacles. The results of this analysis are returned to the device, and the user is notified of the situation via voice, such as, "The traffic light is red. Please stop."

[0052] Shopping assistance

[0053] Product Selection

[0054] The user voice-inputs, "I want to buy tomatoes." The device then uses speech recognition technology to convert this voice into text and sends it to the server.

[0055] Product recognition

[0056] The server uses a model to recognize products from the terminal's camera footage based on the received text information (product information). When the user points the camera at a tomato shelf, the footage is sent to the server, and image analysis is performed using deep learning technology.

[0057] Freshness evaluation

[0058] The freshness and quality of tomatoes are evaluated based on image analysis. For example, the evaluation is based on color, shape, and gloss. The evaluation results are returned to the device, and the user is notified by voice, "This tomato is very fresh."

[0059] Specific example

[0060] Examples of directions

[0061] User: Inputs "Take me to the nearest convenience store" by voice.

[0062] Terminal: Converts speech to text and sends it to the server.

[0063] Server: Uses the Google Maps API to calculate the route from the current location to the destination and sends it back to the device.

[0064] Terminal: Provides voice guidance to the user regarding the calculated route information (e.g., "Go 50 meters and turn right").

[0065] Terminal: Captures camera footage and sends it to a server to recognize the status of traffic lights.

[0066] Server: Recognizes that the traffic light is red and notifies the terminal with the message, "The traffic light is red. Please stop."

[0067] Terminal: Notifies the user via voice.

[0068] Specific examples of shopping

[0069] User: Inputs "I want to buy tomatoes" by voice.

[0070] Terminal: Converts speech to text and sends it to the server.

[0071] Server: Uses a model for product recognition.

[0072] User: Point the camera at the tomato trellis.

[0073] Terminal: Captures camera footage and sends it to the server.

[0074] Server: Evaluates the freshness and quality of tomatoes and sends the results back to the terminal.

[0075] Device: Notifies the user of the evaluation results via voice (e.g., "This tomato is very fresh").

[0076] In this way, the system of the present invention can effectively provide support to visually impaired persons when they are able to move around and shop independently.

[0077] The following describes the processing flow.

[0078] Processing steps for the navigation function

[0079] Step 1:

[0080] The user uses voice input to tell the terminal, "Take me to the nearest convenience store."

[0081] Step 2:

[0082] The device captures voice input and uses speech recognition technology to convert the voice signal into text.

[0083] Step 3:

[0084] The terminal sends the converted text data to the server.

[0085] Step 4:

[0086] The server extracts destination information from the received text data and uses map information to calculate the route from the current location to the destination.

[0087] Step 5:

[0088] The server returns the calculated route information to the terminal.

[0089] Step 6:

[0090] The terminal analyzes route information and uses speech synthesis technology to instruct the user verbally on the next direction and distance to go. For example, it might say, "Go 50 meters and turn right."

[0091] Step 7:

[0092] The device uses its camera to capture images of its surroundings in real time.

[0093] Step 8:

[0094] The terminal sends the captured video to the server.

[0095] Step 9:

[0096] The server receives the video data and uses a deep learning model to analyze obstacles and traffic light conditions within the video.

[0097] Step 10:

[0098] The server sends the analysis results back to the terminal. For example, if the traffic light is red, it will send back an instruction such as, "The traffic light is red. Please stop."

[0099] Step 11:

[0100] The device notifies the user of the analysis results via voice. The notification will be in the form of, "The traffic light is red. Please stop."

[0101] Processing steps for the shopping support function

[0102] Step 1:

[0103] The user uses voice input to say "I want to buy tomatoes" to the device.

[0104] Step 2:

[0105] The device captures voice input and uses speech recognition technology to convert the voice signal into text.

[0106] Step 3:

[0107] The terminal sends the converted text data to the server.

[0108] Step 4:

[0109] The server extracts product information from the received text data and selects an appropriate image recognition model.

[0110] Step 5:

[0111] The user points the camera at the tomato trellis.

[0112] Step 6:

[0113] The device uses a camera to capture real-time video of the product shelves.

[0114] Step 7:

[0115] The terminal sends the captured video to the server.

[0116] Step 8:

[0117] The server receives the video data and uses an image recognition model to identify tomatoes in the video.

[0118] Step 9:

[0119] The server evaluates the freshness and quality of the identified tomatoes.

[0120] Step 10:

[0121] The server sends the evaluation results back to the terminal. For example, it might return an evaluation result such as, "This tomato is very fresh."

[0122] Step 11:

[0123] The device notifies the user of the evaluation results via voice. The notification may be in the form of, "This tomato is very fresh."

[0124] These detailed processing steps enable the system of the present invention to allow visually impaired individuals to independently receive support in various aspects of their daily lives.

[0125] (Example 1)

[0126] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0127] There is a lack of technical support for visually impaired individuals to move independently, safely, and comfortably, and to perform daily shopping. To address this challenge, a system is needed that recognizes destinations and products through voice input and provides appropriate instructions through voice guidance and video analysis. Furthermore, real-time updated navigation information and product freshness assessments based on this system are also required.

[0128] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0129] In this invention, the server includes means for receiving voice input from a user, means for converting the received voice input into text, means for identifying a destination based on the converted text and calculating a route using map information, means for communicating the calculated route information to the user by voice, means for capturing and analyzing video of the user's surroundings, means for recognizing obstacles and traffic light conditions in the captured video, and means for notifying the user of the recognition results by voice. This enables visually impaired individuals to recognize destinations and products through voice input and to move and shop safely while receiving appropriate guidance and instructions in real time.

[0130] "Means for receiving user voice input" refers to a device or program that allows visually impaired individuals to input voice commands using a smartphone or other device.

[0131] "Means for converting received voice input into text" refers to a device or program that uses speech recognition technology to convert voice data into text data.

[0132] "Means for identifying a destination based on converted text and calculating a route using map information" refers to a device or program that analyzes text data, refers to a map database based on a specified destination, and calculates the optimal route.

[0133] "Means of conveying calculated route information to users via voice" refers to a device or program that provides calculated route information as a voice message using speech synthesis technology.

[0134] "Means for capturing and analyzing video footage of the user's surroundings" refers to a device or program that uses a camera to capture images of the user's surroundings in real time and analyzes that image data.

[0135] "Means for recognizing obstacles and traffic light status in captured video" refers to a device or program that uses deep learning technology to recognize information about obstacles and traffic lights from acquired video data.

[0136] "Means for notifying the user of the recognition results by voice" refers to a device or program that provides the recognized information as a voice message using speech synthesis technology.

[0137] "A means of inputting a product name by voice, analyzing that information, and converting it into text" refers to a device or program that allows a user to input a product name by voice and convert that information into text data.

[0138] "Means for recognizing products from camera footage" refers to a device or program that uses deep learning technology to identify specific products from video data acquired using a camera.

[0139] "Means for evaluating the freshness and quality of recognized goods" refers to a device or program that evaluates the freshness and quality of recognized goods based on criteria such as color, shape, and gloss.

[0140] "Means for notifying users of evaluation results by voice" refers to a device or program that provides evaluation results of freshness and quality as a voice message using speech synthesis technology.

[0141] "Means for updating routes in real time" refers to a device or program that dynamically updates the navigation route for each action based on the user's current location information.

[0142] "Means of providing voice instructions for the direction of travel or the next action" refers to a device or program that uses speech synthesis technology to provide the user with voice messages indicating the direction and action they should take next.

[0143] "Means for recognizing moving objects such as pedestrians and animals in the surrounding area and notifying the user of that information via voice" refers to a device or program that recognizes moving objects from camera footage and provides information for avoiding danger as a voice message.

[0144] System Overview

[0145] This invention is an assistance system for visually impaired individuals to enable them to move around and perform daily shopping independently. This system is comprised of a combination of voice input, speech synthesis, image recognition, GPS, map information, and AI technology.

[0146] Navigation function

[0147] Enter destination

[0148] The user inputs a voice command into their smartphone saying, "Take me to the nearest convenience store."

[0149] The device uses speech recognition technology to convert this speech into text. Specifically, it uses Google Cloud Speech-to-Text.

[0150] The device sends the generated text data ("Please guide me to the nearest convenience store") to the server.

[0151] The server receives the text data and uses it to determine the destination.

[0152] Root calculation

[0153] The server determines the user's current location based on the destination information received.

[0154] The server uses map information and the Google Maps API to calculate the optimal walking route from the current location to the destination.

[0155] The server generates the calculated route information in JSON format and sends it back to the terminal.

[0156] Voice directions

[0157] The terminal parses the received route information in JSON format.

[0158] To guide the user verbally in the next direction and distance, a voice message is generated using Google Cloud Text-to-Speech.

[0159] For example, based on route information, it might instruct you to "Go 50 meters and turn right."

[0160] Real-time video analysis

[0161] The device activates its camera and captures video of the user's surroundings in real time.

[0162] The terminal sends the captured video data to the server.

[0163] The server analyzes the received video using deep learning technology (TENSORFLOW®). It recognizes traffic lights and obstacles, generates that information in text format, and sends it back to the terminal.

[0164] The device generates a voice message using speech synthesis technology (Google Cloud Text-to-Speech) based on the received recognition results and notifies the user.

[0165] For example, it might announce, "The traffic light is red. Please stop."

[0166] Shopping assistance

[0167] Product Selection

[0168] The user voice-inputs "I want to buy tomatoes" into their smartphone.

[0169] The device uses speech recognition technology (Google Cloud Speech-to-Text) to convert this speech into text.

[0170] The terminal sends the generated text data ("I want to buy tomatoes") to the server.

[0171] The server receives the text data and processes it as product information.

[0172] Product recognition

[0173] The server uses a deep learning model to recognize tomatoes based on product information.

[0174] The user points the camera at the tomato trellis.

[0175] The device captures camera footage and sends it to the server.

[0176] The server analyzes the received video using deep learning technology (TensorFlow) to recognize tomatoes.

[0177] Freshness evaluation

[0178] The server evaluates the freshness and quality of the recognized tomatoes. Specifically, it evaluates them based on criteria such as color, shape, and gloss.

[0179] The server generates the evaluation results in text format and sends them back to the terminal.

[0180] The device generates a voice message saying, "These tomatoes are very fresh," and notifies the user.

[0181] Specific example

[0182] Examples of directions

[0183] User: Inputs "Take me to the nearest convenience store" by voice.

[0184] Terminal: Use Google Cloud Speech-to-Text to convert speech to text and send it to the server.

[0185] Server: Uses the Google Maps API to calculate the route from the current location to the destination and returns it to the device in JSON format.

[0186] Terminal: Analyzes the received route information and uses Google Cloud Text-to-Speech to provide voice guidance such as, "Go 50 meters and turn right."

[0187] Terminal: Captures camera footage and sends it to a server to recognize the status of traffic lights.

[0188] Server: Uses deep learning technology (TensorFlow) to recognize that the signal is red and sends a message back to the terminal.

[0189] Terminal: Based on the received information, it notifies the user by voice, "The traffic light is red. Please stop."

[0190] Specific examples of shopping

[0191] User: Inputs "I want to buy tomatoes" by voice.

[0192] Terminal: Use Google Cloud Speech-to-Text to convert speech to text and send it to the server.

[0193] Server: Prepares to recognize tomatoes using deep learning technology.

[0194] User: Point the camera at the tomato trellis.

[0195] Terminal: Captures camera footage and sends it to the server.

[0196] Server: Uses deep learning technology (TensorFlow) to analyze received video and evaluate the freshness and quality of tomatoes.

[0197] Server: Generates evaluation results in text format and sends them back to the terminal.

[0198] Device: Use Google Cloud Text-to-Speech to announce, "These tomatoes are very fresh."

[0199] In this way, the system of the present invention can effectively provide support to visually impaired persons when they are able to move around and shop independently.

[0200] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0201] Specific processing flow of the navigation function

[0202] Enter destination

[0203] Step 1:

[0204] The user uses voice input on their smartphone, saying, "Take me to the nearest convenience store."

[0205] Input: Voice command ("Take me to the nearest convenience store.")

[0206] Output: Audio data

[0207] Step 2:

[0208] The device uses speech recognition technology (Google Cloud Speech-to-Text) to convert the input voice data into text data.

[0209] Specifically, audio data is sent to the server, and the server returns text data.

[0210] Input: Audio data

[0211] Output: Text data ("Please guide me to the nearest convenience store")

[0212] Step 3:

[0213] The terminal sends the generated text data to the server. The server receives this data and processes it as destination information.

[0214] Input: Text data ("Please guide me to the nearest convenience store")

[0215] Output: Destination information

[0216] Root calculation

[0217] Step 4:

[0218] The server determines the user's current location based on the destination information received.

[0219] Input: Destination information

[0220] Output: Current location information

[0221] Step 5:

[0222] The server uses the Google Maps API to calculate the optimal walking route from the current location to the destination.

[0223] Input: Current location information, destination information

[0224] Output: Route information (JSON format)

[0225] Step 6:

[0226] The server generates the calculated route information in JSON format and sends it back to the terminal.

[0227] Input: Route information (in generated JSON format)

[0228] Output: Route information (JSON format)

[0229] Voice directions

[0230] Step 7:

[0231] The terminal analyzes the received route information and calculates the next direction and distance to proceed.

[0232] Input: Route information (JSON format)

[0233] Output: Guidance / Instructions

[0234] Step 8:

[0235] The device uses Google Cloud Text-to-Speech to generate guidance instructions as voice messages and notify the user.

[0236] For example, you might give directions like, "Go 50 meters and turn right."

[0237] Input: Guidance / Instructions

[0238] Output: Voice message

[0239] Real-time video analysis

[0240] Step 9:

[0241] The device activates its camera and captures video of the user's surroundings in real time.

[0242] Input: Camera video

[0243] Output: Video data

[0244] Step 10:

[0245] The terminal sends the captured video data to the server.

[0246] Input: Video data

[0247] Output: Video data sent to the server

[0248] Step 11:

[0249] The server uses deep learning technology (TensorFlow) to analyze the received video and recognize the status of traffic lights and obstacles.

[0250] Input: Video data

[0251] Output: Analysis results (status of traffic lights and obstacles)

[0252] Step 12:

[0253] The server generates the analysis results in text format and sends them back to the terminal.

[0254] Input: Analysis results

[0255] Output: Text data (analysis results)

[0256] Step 13:

[0257] The device generates a voice message using Google Cloud Text-to-Speech based on the received recognition results and notifies the user.

[0258] For example, you might announce, "The traffic light is red. Please stop."

[0259] Input: Text data (analysis results)

[0260] Output: Voice message

[0261] Specific processing flow for shopping assistance

[0262] Product Selection

[0263] Step 1:

[0264] The user voice-inputs "I want to buy tomatoes" into their smartphone.

[0265] Input: Voice command ("I want to buy tomatoes")

[0266] Output: Audio data

[0267] Step 2:

[0268] The device uses speech recognition technology (Google Cloud Speech-to-Text) to convert this speech into text.

[0269] Specifically, audio data is sent to the server, and the server returns text data.

[0270] Input: Audio data

[0271] Output: Text data ("I want to buy tomatoes")

[0272] Step 3:

[0273] The terminal sends the generated text data to the server. The server receives this data and processes it as product information.

[0274] Input: Text data ("I want to buy tomatoes")

[0275] Output: Product Information

[0276] Product recognition

[0277] Step 4:

[0278] The server prepares a product recognition model based on the product information.

[0279] Input: Product Information

[0280] Output: Recognition Model

[0281] Step 5:

[0282] The user holds the camera over the tomato shelf.

[0283] Input: Video of the product shelf

[0284] Output: Camera Data

[0285] Step 6:

[0286] The terminal captures the camera video and sends it to the server.

[0287] Input: Camera Data

[0288] Output: Video Data Sent to the Server

[0289] Step 7:

[0290] The server uses deep learning technology (TensorFlow) to analyze the received video and recognize the products.

[0291] Input: Video Data

[0292] Output: Analysis Result (Product Recognition)

[0293] Step 8:

[0294] The server generates the recognition result in text format and returns it to the terminal.

[0295] Input: Analysis Result

[0296] Output: Text Data (Product Recognition)

[0297] Freshness Evaluation

[0298] Step 9:

[0299] The server evaluates the freshness and quality of the recognized products. Specifically, the evaluation is carried out based on criteria such as color, shape, and gloss.

[0300] Input: Product recognition result

[0301] Output: Freshness evaluation result

[0302] Step 10:

[0303] The server generates the evaluation result in text format and returns it to the terminal.

[0304] Input: Freshness evaluation result

[0305] Output: Text data (evaluation result)

[0306] Step 11:

[0307] The terminal uses Google Cloud Text-to-Speech to generate the evaluation result as an audio message and notify the user.

[0308] As an example, it guides with "This tomato has very good freshness".

[0309] Input: Text data (evaluation result)

[0310] Output: Audio message

[0311] (Application Example 1)

[0312] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart device 14 is referred to as the "terminal".

[0313] Visually impaired individuals face numerous challenges when attempting to travel and shop independently. In particular, accurately perceiving their surroundings is difficult when seeking directions to a destination or purchasing goods. They struggle to recognize traffic signals, pedestrians, animals, and other moving objects. Furthermore, locating products and assessing their freshness and quality is extremely challenging. Effective support systems are needed to address these issues.

[0314] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0315] In this invention, the server includes: means for receiving voice input from the user; means for converting the received voice input into text; means for identifying a destination based on the converted text and calculating a route using map information; means for communicating the calculated route information to the user by voice; means for capturing and analyzing video of the user's surroundings; means for recognizing obstacles and traffic light conditions in the captured video; means for notifying the user of the recognition results by voice; means for the user to input a product name by voice, analyzing that information and converting it into text; means for recognizing a product from the camera image based on the converted text; and means for providing information about the recognized product. The system includes means for evaluating degree and quality; means for notifying the user of the evaluation results by voice; means for identifying the user's current location and updating the route in real time using map information; means for instructing the user on the direction of travel and the next action by voice based on the updated route information; means for recognizing moving objects such as pedestrians and animals in the surroundings other than obstacles and traffic lights, and notifying the user of that information by voice; means for the user to input searched products by voice, and for analyzing that information and converting it into text; means for calculating the evaluation results of product candidates based on the converted text using deep learning technology; and means for guiding the user to the desired product in real time based on the evaluation results. This enables visually impaired people to move and shop independently, safely and effectively.

[0316] A "system" refers to the entire apparatus, including a series of means, designed to support visually impaired individuals in independently navigating and shopping.

[0317] "Users" refers to visually impaired individuals who use this system to receive guidance and shopping assistance.

[0318] "Voice input" refers to the voice signals that a user speaks to the system.

[0319] "Means for receiving voice input" refers to devices or software that acquire the user's voice and allow the system to interpret it.

[0320] "Means of converting speech to text" refers to devices or software that analyze speech input and convert it into textual information.

[0321] "Means of identifying a destination based on text" refers to devices or software that determine a destination based on converted character information.

[0322] "Map information" refers to a database containing geographical information, providing destinations, current location, route information, and more.

[0323] "Means of calculating routes" refers to devices or software that use map information to calculate the optimal route to a destination.

[0324] "Means of conveying route information by voice" refers to devices or software that guide users through calculated routes by voice.

[0325] "Means of capturing video of the user's surroundings" refers to devices that acquire images or videos of the user's surroundings using cameras or similar devices.

[0326] "Means of analyzing video" refers to devices or software that analyze captured video to extract useful information.

[0327] "Means for recognizing obstacles and traffic signal status" refers to devices and software that detect the status of surrounding obstacles and traffic signals through video analysis.

[0328] "Means of notifying the user of recognition results by voice" refers to devices or software that inform the user of the analysis results by voice.

[0329] "Methods for inputting product names by voice" refer to devices or software that allow users to tell the system by voice the products they wish to purchase.

[0330] "Means of recognizing products from camera footage" refers to devices or software that use a camera to acquire images of products and then recognize those products.

[0331] "Means for evaluating freshness and quality" refers to devices or software used to evaluate the condition of recognized products.

[0332] "Means for determining current location" refers to devices or software used to determine the user's current geographical location.

[0333] "Means of updating routes in real time" refers to devices or software that continuously update route guidance in accordance with the user's movements.

[0334] "Means of providing voice instructions for direction of travel or next action" refers to devices or software that provide voice guidance for the next action based on updated route information.

[0335] "Means of recognizing moving objects such as pedestrians and animals" refers to devices and software used to detect moving objects in the surrounding environment.

[0336] "Methods for inputting searched products by voice" refers to devices or software that allow users to tell the system by voice what product they are looking for.

[0337] "Methods for calculating evaluation results of candidate products using deep learning technology" refers to devices or software that analyze and evaluate the characteristics of candidate products using deep learning.

[0338] "A means of guiding users to desired products in real time based on evaluation results" refers to devices or software that guide users to products they are looking for based on evaluation results.

[0339] System Configuration

[0340] To implement this invention, a server, terminals, and users are required. The terminals consist of smartphones or devices with built-in cameras. The server is responsible for data processing and analysis and operates in a cloud computing environment. The system integrates multiple functions, including voice input recognition, text conversion, route guidance using map information, real-time video analysis, product recognition, and quality evaluation.

[0341] Program and Processing Overview

[0342] Voice input acceptance and text conversion:

[0343] The user provides voice input. For example, they might say, "Take me to the nearest convenience store." The device receives this input and uses speech recognition technology to convert it into text. The Google Speech-to-Text API is used for this process.

[0344] Destination identification and route calculation:

[0345] The converted audio data is sent to a server, which identifies the destination based on the converted text. It then uses map information to calculate the optimal route. The Google Maps API is used here. The calculated route information is returned to the device.

[0346] Voice-guided route guidance:

[0347] The terminal analyzes route information calculated using speech synthesis technology and provides voice instructions to the user regarding the next direction and distance to travel. For example, it might say, "Go 50 meters and turn right."

[0348] Real-time video capture and analysis:

[0349] The device's camera captures video of the user's surroundings in real time and sends the data to a server. The server uses deep learning technology to analyze the video data and recognize traffic lights and obstacles. The results of this analysis are returned to the device and notified to the user via voice.

[0350] Product recognition and quality assessment:

[0351] The user voice-inputs "I want to buy tomatoes." This voice is converted back into text, and the server uses a model that recognizes products from camera footage based on the text information. When the user points the camera at a shelf of tomatoes, the image is sent to the server. Deep learning technology is used to evaluate the freshness and quality of the tomatoes, and the results are returned to the terminal. For example, the user might receive a voice notification saying, "These tomatoes are very fresh."

[0352] Current location and real-time route updates:

[0353] The server uses GPS to determine the user's current location and updates the route in real time based on map information. Based on the updated route information, the terminal provides voice instructions to the user regarding the direction to go and the next action.

[0354] Recognition and notification of moving objects:

[0355] The server uses video analysis to recognize moving objects in the surrounding area, such as pedestrians and animals, other than obstacles and traffic lights, and notifies the user of this information via voice.

[0356] Specific example

[0357] 1. Specific Examples of Navigation

[0358] User: "Take me to the nearest convenience store."

[0359] Terminal: Converts speech to text and sends it to the server.

[0360] Server: Route calculation using Google Maps API

[0361] Device: Voice-guided route instructions (e.g., "Go 50 meters and turn right")

[0362] Terminal: Captures camera footage and sends the traffic light status to the server.

[0363] Server: Recognizes the status of the traffic light and notifies, "The traffic light is red. Please stop."

[0364] Device: Notifies the user via voice.

[0365] 2. Specific examples of shopping

[0366] User: "I want to buy tomatoes."

[0367] Terminal: Converts speech to text and sends it to the server.

[0368] Server: Uses a model for product recognition.

[0369] User: Pointing the camera at the tomato trellis.

[0370] Terminal: Captures camera footage and sends it to the server.

[0371] Server: Evaluates the freshness and quality of the tomatoes and notifies, "These tomatoes are very fresh."

[0372] Device: Notifies the user via voice.

[0373] This system assists users in moving around and shopping safely and conveniently. Its purpose is to provide an environment where visually impaired individuals can live independently, utilizing technologies such as voice recognition, deep learning, and GPS.

[0374] Example of a prompt

[0375] 1. Navigation prompt text:

[0376] Please guide me to the nearest convenience store.

[0377] 2. Product purchase prompt message:

[0378] I bought tomatoes.

[0379] As described above, this invention provides comprehensive support for visually impaired individuals to move around and shop independently.

[0380] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0381] Flow of processing steps

[0382] Navigation function

[0383] Step 1:

[0384] Input: The user inputs "Take me to the nearest convenience store" by voice.

[0385] Operation: The device (smartphone) accepts voice input.

[0386] Output: Audio data.

[0387] Step 2:

[0388] Input: Audio data.

[0389] Operation: The device uses speech recognition technology to convert speech data into text. The Google Speech-to-Text API is used here.

[0390] Output: Text data.

[0391] Step 3:

[0392] Input: Text data.

[0393] Operation: The server identifies the destination based on text data and calculates the route using the Google Maps API.

[0394] Output: Route information.

[0395] Step 4:

[0396] Input: Route information.

[0397] Operation: The server sends route information back to the terminal, and the terminal parses this information.

[0398] Output: Route guidance information after analysis.

[0399] Step 5:

[0400] Input: Route guidance information after analysis.

[0401] Operation: The device uses speech synthesis technology to provide instructions to the user. For example, it might say, "Go 50 meters and turn right."

[0402] Output: Voice guidance.

[0403] Step 6:

[0404] Input: Video of the user's surroundings.

[0405] Operation: The device's camera captures video and sends it to the server.

[0406] Output: Video data.

[0407] Step 7:

[0408] Input: Video data.

[0409] Operation: The server uses deep learning technology to analyze video and recognize traffic lights and obstacles.

[0410] Output: Recognition result.

[0411] Step 8:

[0412] Input: Recognition result.

[0413] Operation: The server sends the recognition result back to the terminal, and the terminal notifies the user by voice, "The traffic light is red. Please stop."

[0414] Output: Voice notification.

[0415] Shopping support function

[0416] Step 1:

[0417] Input: The user enters "I want to buy tomatoes" by voice.

[0418] Operation: The device (smartphone) accepts voice input.

[0419] Output: Audio data.

[0420] Step 2:

[0421] Input: Audio data.

[0422] Operation: The device uses speech recognition technology to convert speech data into text. The Google Speech-to-Text API is used here.

[0423] Output: Text data.

[0424] Step 3:

[0425] Input: Text data.

[0426] Operation: The server prepares a product recognition model based on text data.

[0427] Output: Recognition model.

[0428] Step 4:

[0429] Input: Camera footage.

[0430] Operation: The user points their smartphone camera at a product shelf (e.g., a tomato shelf). This video data is sent from the device to the server.

[0431] Output: Video data.

[0432] Step 5:

[0433] Input: Video data.

[0434] Operation: The server uses deep learning technology to recognize products from images and evaluate the freshness and quality of tomatoes.

[0435] Output: Evaluation results.

[0436] Step 6:

[0437] Input: Evaluation results.

[0438] Operation: The server sends the evaluation result back to the terminal, and the terminal notifies the user by voice, "This tomato is very fresh."

[0439] Output: Voice notification.

[0440] Real-time route update function

[0441] Step 1:

[0442] Input: Current location.

[0443] Operation: The device's GPS function determines the user's current location and sends it to the server.

[0444] Output: Current location information.

[0445] Step 2:

[0446] Input: Current location information.

[0447] Operation: The server updates the route in real time using map information.

[0448] Output: Updated route information.

[0449] Step 3:

[0450] Input: Updated route information.

[0451] Operation: The device provides voice instructions to the user regarding the direction of travel and the next action based on updated route information.

[0452] Output: Voice command.

[0453] Recognition function for moving objects

[0454] Step 1:

[0455] Input: Video of the user's surroundings.

[0456] Operation: The device's camera captures the surrounding video and sends it to the server.

[0457] Output: Video data.

[0458] Step 2:

[0459] Input: Video data.

[0460] Operation: The server uses deep learning technology to analyze video and recognize moving objects such as pedestrians and animals.

[0461] Output: Recognition result.

[0462] Step 3:

[0463] Input: Recognition result.

[0464] Operation: The server sends the recognition result back to the terminal, and the terminal notifies the user by voice, "There is a pedestrian ahead. Please be careful."

[0465] Output: Voice notification.

[0466] The above outlines the specific processing steps of this system. This will enable visually impaired individuals to move around and shop independently.

[0467] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0468] System Overview

[0469] This invention is an assistance system for visually impaired individuals to independently navigate and perform daily shopping. By incorporating an emotion engine, it provides optimal support based on the user's current emotional state. The system combines voice input, speech synthesis, image recognition, GPS, map information, AI technology, and emotion recognition technology.

[0470] Navigation function

[0471] Enter destination

[0472] The user speaks into their smartphone and says, "Take me to the nearest convenience store." The device uses speech recognition technology to convert the speech into text. This text information is sent to a server and processed as destination information.

[0473] Root calculation

[0474] The server extracts destination information from the received text information and calculates the route from the current location to the destination using map information. It then sends the calculated route information back to the terminal.

[0475] Voice directions

[0476] The terminal analyzes the calculated route information and uses speech synthesis technology to instruct the user on the next direction and distance to go. For example, it provides step-by-step instructions such as, "Go 50 meters and turn right."

[0477] Real-time video analysis

[0478] The device uses a camera to capture video of the user's surroundings in real time. This video data is sent to a server, which analyzes the video using deep learning technology. Specifically, it recognizes the status of traffic lights and obstacles. The results of this analysis are sent back to the device, and the user is notified of the situation via voice, such as, "The traffic light is red. Please stop."

[0479] Emotion recognition and regulation

[0480] The device captures the user's voice and facial expressions and uses an emotion engine to recognize their current emotional state. For example, if the user is anxious, the guidance voice will slow down and provide detailed instructions to create a sense of security. Conversely, if the user is relaxed, the guidance will be given at a normal speed.

[0481] Shopping assistance

[0482] Product Selection

[0483] The user voice-inputs, "I want to buy tomatoes." The device uses speech recognition technology to convert the voice into text and sends it to the server.

[0484] Product recognition

[0485] The server uses a model to recognize products from the terminal's camera footage based on the received text information. When the user points the camera at a tomato shelf, the footage is sent to the server, and image analysis is performed using deep learning technology.

[0486] Freshness evaluation

[0487] The freshness and quality of tomatoes are evaluated based on image analysis. For example, the evaluation is based on color, shape, and gloss. The evaluation results are sent back to the device, and the user is notified by voice, "This tomato is very fresh."

[0488] Emotion recognition and product recommendations

[0489] The device captures the user's voice and facial expressions and uses an emotion engine to recognize their current emotional state. For example, if the user is undecided, it suggests other items on the same shelf to broaden their options. On the other hand, if the user is in a hurry, it quickly suggests the highest-rated product.

[0490] Specific example

[0491] Examples of directions

[0492] User: Inputs "Take me to the nearest convenience store" by voice.

[0493] Terminal: Converts speech to text and sends it to the server.

[0494] Server: Calculates the route from the current location to the destination and sends it back to the terminal.

[0495] Terminal: Provides route information to the user via voice (e.g., "Go 50 meters and turn right").

[0496] Terminal: Captures camera footage and sends it to the server to analyze the status of the traffic lights.

[0497] Server: Recognizes that the traffic light is red and notifies, "The traffic light is red. Please stop."

[0498] Terminal: Notifies the user via voice.

[0499] Terminal: If the user is anxious, slow down the voice guidance and add more detailed explanations.

[0500] Specific examples of shopping

[0501] User: Inputs "I want to buy tomatoes" by voice.

[0502] Terminal: Converts speech to text and sends it to the server.

[0503] Server: Uses a model for product recognition.

[0504] User: Point the camera at the tomato trellis.

[0505] Terminal: Captures camera footage and sends it to the server.

[0506] Server: Evaluates the freshness and quality of tomatoes and sends the results back to the terminal.

[0507] Device: Notifies the user of the evaluation results via voice (e.g., "This tomato is very fresh").

[0508] Terminal: If the user is undecided, suggest other products.

[0509] In this way, the system of the present invention can provide optimal support tailored to the emotional state of visually impaired individuals when they are independently moving around or shopping.

[0510] The following describes the processing flow.

[0511] Processing steps for the navigation function

[0512] Step 1:

[0513] The user uses voice input to tell the terminal, "Take me to the nearest convenience store."

[0514] Step 2:

[0515] The device captures voice input and uses speech recognition technology to convert the voice signal into text.

[0516] Step 3:

[0517] The terminal sends the converted text data to the server.

[0518] Step 4:

[0519] The server extracts destination information from the received text data and uses map information to calculate the route from the current location to the destination.

[0520] Step 5:

[0521] The server returns the calculated route information to the terminal.

[0522] Step 6:

[0523] The terminal analyzes route information and uses speech synthesis technology to instruct the user verbally on the next direction and distance to go. For example, it might say, "Go 50 meters and turn right."

[0524] Step 7:

[0525] The device uses its camera to capture images of its surroundings in real time.

[0526] Step 8:

[0527] The terminal sends the captured video to the server.

[0528] Step 9:

[0529] The server receives the video data and uses a deep learning model to analyze obstacles and traffic light conditions within the video.

[0530] Step 10:

[0531] The server sends the analysis results back to the terminal. For example, if the traffic light is red, it will send back an instruction such as, "The traffic light is red. Please stop."

[0532] Step 11:

[0533] The device notifies the user of the analysis results via voice. The notification will be in the form of, "The traffic light is red. Please stop."

[0534] Step 12:

[0535] The device captures the user's voice and facial expressions and uses an emotion engine to recognize their current emotional state.

[0536] Step 13:

[0537] The device adjusts the speed and content of voice guidance based on the recognized emotional state. For example, if the user is anxious, the voice guidance will slow down and more detailed instructions will be added.

[0538] Processing steps for the shopping support function

[0539] Step 1:

[0540] The user uses voice input to say "I want to buy tomatoes" to the device.

[0541] Step 2:

[0542] The device captures voice input and uses speech recognition technology to convert the voice signal into text.

[0543] Step 3:

[0544] The terminal sends the converted text data to the server.

[0545] Step 4:

[0546] The server extracts corresponding product information based on the received text data.

[0547] Step 5:

[0548] The user points the camera at the tomato trellis.

[0549] Step 6:

[0550] The terminal captures video of the product shelves in real time and sends that video to the server.

[0551] Step 7:

[0552] The server analyzes the received video data and identifies the recognized products using an image recognition model.

[0553] Step 8:

[0554] The server evaluates the freshness and quality of the products.

[0555] Step 9:

[0556] The server sends the evaluation results back to the terminal. For example, it might return an evaluation result such as, "This tomato is very fresh."

[0557] Step 10:

[0558] The device notifies the user of the evaluation results via voice. The notification may be in the form of, "This tomato is very fresh."

[0559] Step 11:

[0560] The device captures the user's voice and facial expressions and uses an emotion engine to recognize their current emotional state.

[0561] Step 12:

[0562] The device makes product suggestions based on the recognized emotional state. For example, if the user is undecided, it adds suggestions for other products. On the other hand, if the user is in a hurry, it quickly suggests highly-rated products.

[0563] These detailed processing steps enable the system of the present invention to allow visually impaired individuals to independently receive support in various aspects of their daily lives. Furthermore, it detects the user's emotional state and provides optimal support accordingly.

[0564] (Example 2)

[0565] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0566] There is a need to ensure that visually impaired individuals receive appropriate support when they are able to move around independently and do their daily shopping. In particular, flexible responses that take into account the user's emotional state are necessary in areas such as navigation to destinations, product selection, and recognition of obstacles along the way. However, current systems have difficulty meeting all of these requirements simultaneously, which has been a problem as it makes it inconvenient for visually impaired individuals to move around and shop with peace of mind.

[0567] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0568] In this invention, the server includes means for receiving voice input from a user, means for converting the received voice input into text, means for identifying a destination based on the converted text and calculating a route using map information, means for communicating the calculated route information to the user by voice, means for capturing and analyzing video of the user's surroundings, means for recognizing obstacles and traffic light conditions in the captured video, means for notifying the user of the recognition results by voice, and means for recognizing the user's emotional state and adjusting voice guidance based on that emotion. This makes it possible for visually impaired people to receive optimal support in real time, tailored to environmental information and their emotional state, when they are able to move around and shop independently with peace of mind.

[0569] "Means for receiving user voice input" refers to a mechanism for capturing the voice spoken by the user into the device.

[0570] "Means of converting received voice input into text" refers to the process of converting voice data into text data.

[0571] "A means of identifying a destination based on converted text and calculating a route using map information" refers to a process that analyzes text data to determine the destination and then calculates a route using map information.

[0572] "A means of communicating calculated route information to the user via voice" refers to a mechanism that informs the user of calculated route information via voice.

[0573] "Means for capturing and analyzing video footage of the user's surroundings" refers to a mechanism that uses a camera to capture video footage of the user's surroundings and then analyzes that footage.

[0574] "Means for recognizing obstacles and traffic light conditions in captured video" refers to technologies for identifying the condition of obstacles and traffic lights from captured video footage.

[0575] "Means for notifying the user of recognition results by voice" refers to a mechanism that communicates the analysis results to the user by voice.

[0576] "Means for recognizing the user's emotional state and adjusting voice guidance based on those emotions" refers to technology that analyzes the user's emotions and adjusts the content and speed of voice guidance accordingly.

[0577] "A method for inputting product names by voice, analyzing that information, and converting it into text" refers to a method in which a user inputs a product name by voice, and that voice is analyzed and converted into text data.

[0578] "Means for recognizing products from camera footage based on converted text" refers to a technology that uses converted text to identify specific products from camera footage.

[0579] "Means for evaluating the freshness and quality of recognized products" refers to the process of analyzing the freshness and quality of identified products.

[0580] "Means of notifying users of evaluation results by voice" refers to a mechanism that informs users of the product evaluation results by voice.

[0581] "A means of identifying the user's current location and updating the route in real time using map information" refers to a technology that obtains the user's current location and recalculates the route in real time based on that location information.

[0582] "A means of instructing users on their direction of travel and next actions via voice based on updated route information" refers to a voice guidance system that uses recalculated route information to provide instructions to users.

[0583] "A means of recognizing moving objects in the surroundings, such as pedestrians and animals other than obstacles and traffic lights, and notifying the user of that information by voice" refers to a technology that identifies moving objects in the surroundings and transmits that information to the user by voice.

[0584] This invention is an assistance system for visually impaired individuals to independently navigate and perform daily shopping, providing optimal support based on the user's emotional state. This system is realized by combining voice input, speech synthesis, image recognition, GPS, map information, AI technology, and emotion recognition technology.

[0585] System Configuration

[0586] The system mainly consists of the following components:

[0587] Speech recognition technology

[0588] Speech synthesis technology

[0589] Image recognition technology

[0590] GPS

[0591] Map information

[0592] emotion recognition technology

[0593] server

[0594] Devices (smartphones and mobile devices)

[0595] Voice input and conversion

[0596] The user provides voice input. For example, they might say, "Take me to the nearest convenience store." This voice data is captured by the device. The device uses speech recognition technology (e.g., Google Cloud Speech-to-Text) to convert the voice data into text data. The converted text data is sent to a server and processed as destination information.

[0597] Route calculation and communication

[0598] The server analyzes the received text data and extracts destination information. Then, using map information (such as the Google Maps API), the server calculates the optimal route from the user's current location to the destination and sends the calculated route information back to the device. The device then uses speech synthesis technology (for example, Amazon Polly) to communicate the calculated route information to the user verbally.

[0599] Surrounding video analysis

[0600] The device uses its camera to capture video of the user's surroundings in real time and sends the video data to a server. The server analyzes the video data using deep learning technology (such as TensorFlow). This analysis allows for the recognition of traffic lights and obstacles. The analysis results are sent back from the server to the device, and the user is notified of the situation via voice. For example, it might say, "The traffic light is red. Please stop."

[0601] Emotion recognition and regulation

[0602] The device captures the user's voice and facial expressions through its camera and microphone. The captured data is analyzed using emotion recognition technology (e.g., IBM Watson® Tone Analyzer). Based on the analysis results, the user's emotional state is determined. For example, if the user is anxious, the speed of the voice guidance is slowed down, and more detailed instructions are added.

[0603] Shopping assistance

[0604] The user voice-inputs, "I want to buy tomatoes." The terminal uses speech recognition technology to convert the voice into text and sends it to the server. Based on the received text information, the server recognizes the products from the terminal's camera image. When the user points the camera at the tomato shelf, the image is sent to the server and analyzed using deep learning technology. The freshness and quality of the recognized tomatoes are evaluated, and the evaluation results are sent back to the terminal. The terminal then notifies the user by voice, "These tomatoes are very fresh."

[0605] Emotion recognition and product recommendations

[0606] The device then captures the user's voice and facial expressions and analyzes them using emotion recognition technology. For example, if the user is unsure, it suggests other products on the same shelf. On the other hand, if the user is in a hurry, it quickly suggests the highest-rated product.

[0607] Specific example

[0608] Examples of directions

[0609] User: "Take me to the nearest convenience store" (voice input).

[0610] Terminal: Converts audio data to text and sends it to the server.

[0611] Server: Extracts destination information, calculates the optimal route, and sends it back to the terminal.

[0612] Terminal: Provides calculated route information via voice (e.g., "Go 50 meters and turn right").

[0613] Terminal: Captures camera footage and sends it to the server.

[0614] Server: Analyzes video and recognizes the signal status.

[0615] Terminal: Announces analysis results by voice (e.g., "The traffic light is red. Please stop.").

[0616] Terminal: Adjusts the content of voice guidance according to the user's emotional state.

[0617] Specific examples of shopping

[0618] User: "I want to buy tomatoes" (voice input).

[0619] Terminal: Converts speech to text and sends it to the server.

[0620] Server: Uses a product recognition model, and the user points their camera at a tomato shelf.

[0621] Terminal: Captures video and sends it to the server.

[0622] Server: Analyzes video footage to evaluate the freshness of tomatoes.

[0623] Device: Notifies the evaluation result by voice (e.g., "This tomato is very fresh").

[0624] Terminal: If the user is undecided, suggest other products.

[0625] Example of a prompt

[0626] The following are examples of prompts to input into the generative AI model:

[0627] "Please describe the overview, functions, and specific examples of support systems for visually impaired individuals to shop independently."

[0628] "Please describe the specific processing steps and technologies used in a system that combines speech recognition and image recognition technologies."

[0629] The above describes specific embodiments for carrying out the present invention.

[0630] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0631] Step 1: Obtaining user voice input

[0632] The user voice-inputs "Take me to the nearest convenience store" into their smartphone. The device receives this voice input via its microphone and stores it as audio data.

[0633] Input data: Audio data

[0634] Output data: Stored audio data

[0635] Step 2: Convert audio to text

[0636] The device analyzes the accumulated audio data using speech recognition technology (for example, Google Cloud Speech-to-Text) and converts it into text data.

[0637] Input data: Audio data

[0638] Data processing: Analysis using speech recognition technology

[0639] Output data: Text data

[0640] Step 3: Send destination information

[0641] The converted text data is sent from the terminal to the server. The server receives this text data and processes it as destination information.

[0642] Input data: Text data

[0643] Data processing: Data transmission

[0644] Output data: Text data sent to the server

[0645] Step 4: Calculating square roots

[0646] The server parses the received text data and extracts destination information. Next, it uses map information (for example, Google Maps API) and GPS information to calculate the route from the current location to the destination.

[0647] Input data: Text data, map information, GPS information

[0648] Data processing: Extraction of destination information and route calculation.

[0649] Output data: Route information

[0650] Step 5: Return route information

[0651] The calculated route information is sent back from the server to the terminal. The terminal receives this route information.

[0652] Input data: Route information

[0653] Data processing: Data transmission

[0654] Output data: Route information sent to the terminal

[0655] Step 6: Generate audio guide

[0656] The device analyzes the received route information and generates a voice guide using speech synthesis technology (for example, Amazon Polly).

[0657] Input data: Route information

[0658] Data processing: Voice guide generation using speech synthesis technology

[0659] Output data: Audio guide

[0660] Step 7: Providing audio guides

[0661] The device provides the user with generated audio guidance. Specifically, it gives instructions via voice, such as, "Go 50 meters and turn right."

[0662] Input data: Audio guide

[0663] Data processing: Audio guide playback

[0664] Output data: Voice guidance

[0665] Step 8: Video Capture

[0666] The device uses its camera to capture video footage of the user's surroundings in real time. This video data is then sent to a server.

[0667] Input data: Camera video

[0668] Data processing: Video capture and transmission

[0669] Output data: Video data sent to the server

[0670] Step 9: Video Analysis

[0671] The server analyzes the received video data using deep learning technology (for example, TensorFlow). This allows it to recognize traffic lights and obstacles.

[0672] Input data: Video data

[0673] Data processing: Video analysis using deep learning technology

[0674] Output data: Recognition result

[0675] Step 10: Return the recognition results

[0676] The server sends the analysis results back to the terminal. Specifically, it includes information such as, "The signal is red. Please stop." The terminal receives this information.

[0677] Input data: Recognition result

[0678] Data processing: Data transmission

[0679] Output data: Recognition results sent to the terminal

[0680] Step 11: Notification of Recognition Results

[0681] The device notifies the user of the received recognition results via voice.

[0682] Input data: Recognition result

[0683] Data processing: Voice notification

[0684] Output data: Voice notification

[0685] Step 12: Recognizing your emotional state

[0686] The device captures the user's voice and facial expressions and analyzes their emotional state using emotion recognition technology (for example, IBM Watson Tone Analyzer).

[0687] Input data: Voice and facial expression data

[0688] Data processing: Analysis using emotion recognition technology

[0689] Output data: Emotional state

[0690] Step 13: Adjusting the voice guidance

[0691] The device adjusts the content and speed of voice guidance based on the recognized emotional state. For example, if the user is anxious, the guidance voice will slow down and more detailed explanations will be added.

[0692] Input data: Emotional state

[0693] Data processing: Adjustment of the content and speed of voice guidance.

[0694] Output data: Adjusted voice guidance

[0695] (Application Example 2)

[0696] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0697] Conventional support systems for the visually impaired possessed basic functions such as voice guidance and image recognition, but lacked adaptive support that took into account the user's emotional state. As a result, they were unable to provide appropriate guidance in situations where the user felt anxious or unsure, or adequate support when the user was unsure about product selection. Furthermore, they were insufficient in supporting independent shopping in physical stores and in real-time adjustments during route guidance. Therefore, this invention proposes a system that solves these problems and provides optimal support according to the user's emotional state.

[0698] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for capturing the user's voice and facial expressions and recognizing their emotional state; means for adjusting the speed and detail of voice guidance based on the recognized emotional state; means for recognizing the user's emotional state and suggesting the most suitable product according to the emotion; and means for dynamically adjusting the guidance content based on the emotional state. As a result, the user can receive guidance and support optimized for their own emotional state.

[0699] "Users" refers to the users of the system, and specifically to individuals with visual impairments.

[0700] "Voice input" refers to the act of capturing the user's voice into the system via a microphone.

[0701] "Text conversion" is the process of converting voice input into text information.

[0702] "Destination" refers to the specific place the user wishes to travel to.

[0703] "Map information" refers to data that includes the route from the current location to the destination and surrounding geographical information.

[0704] "Route calculation" is the process of using map information to determine the optimal route from the current location to the destination.

[0705] "Voice guidance" refers to the act of conveying calculated route information and other instructions to users via voice.

[0706] "Video capture" is the act of capturing images of the surroundings using devices such as cameras.

[0707] "Video analysis" is the process of processing captured video data to identify specific information or objects.

[0708] An "obstacle" refers to a physical object that hinders the user's progress.

[0709] A "traffic light" is a light signaling device used to control road traffic.

[0710] "Recognition results" refer to information obtained through methods such as video analysis and audio analysis.

[0711] "Emotional state" refers to the user's current psychological and emotional condition.

[0712] "Emotion recognition means" refers to tools and technologies used to identify a user's emotional state by analyzing their voice and facial expressions.

[0713] "Product" refers to a specific item that a customer wishes to purchase within a physical store.

[0714] "Product recognition" is the process of identifying a specific product from camera footage or other images.

[0715] "Freshness" refers to an indicator related to the freshness and quality of a product.

[0716] "Quality evaluation" is the process of analyzing the freshness and condition of a product to determine its value.

[0717] "Voice notification" refers to the act of a system communicating analysis results, instructions, and other information to the user via voice.

[0718] "Guidance adjustment" is the process of appropriately changing the content and speed of guidance according to the user's emotional state and circumstances.

[0719] This invention is a support system for visually impaired individuals to independently navigate and perform daily shopping. By incorporating an emotion engine, the system provides optimal support based on the user's current emotional state. This system is realized by combining voice input, speech synthesis, image recognition, GPS, map information, AI technology, and emotion recognition technology.

[0720] System Configuration

[0721] Voice Input and Speech Synthesis: This system uses a terminal equipped with a microphone to accept voice input. Speech synthesis is performed using a speech synthesis library such as pyttsx3.

[0722] Image Recognition: A device equipped with a camera captures video of the surroundings and sends the data to a server. The server uses deep learning technology and image analysis tools such as OpenCV to analyze the video and recognize obstacles and the status of traffic lights.

[0723] GPS and Map Information: Use GPS modules or APIs (e.g., Google Maps API) to obtain current location information and calculate the optimal route to the destination.

[0724] Emotion Recognition: The `pipeline` function from the `transformers` library is used to recognize the user's emotional state from their voice and facial expressions.

[0725] Main Features

[0726] 1. In-store navigation

[0727] Users use their smartphone's voice input function to give commands such as, "Take me to the nearest milk stand."

[0728] The voice input is converted to text using the speech_recognition library and sent to the server.

[0729] The server uses map information based on the text to calculate the route. It then sends the calculation result back to the terminal and provides the user with voice guidance such as, "Go 20 meters and turn right."

[0730] 2. Product Recognition and Evaluation

[0731] When a user is looking for a specific product, they can use voice input, for example, by saying, "Tell me how fresh this milk is."

[0732] A camera is used to photograph the target product, and the video is sent to a server.

[0733] The server uses a deep learning model to recognize products and evaluate their freshness and quality. The evaluation results are sent back to the terminal, and the user is notified via voice message, "This milk is very fresh."

[0734] 3. Support tailored to emotional state

[0735] The system captures the user's voice and facial expressions, and recognizes their emotional state using the transformers library's pipeline.

[0736] Based on the recognized emotional state, the system automatically adjusts the speed and level of detail of voice guidance. For example, if the user is anxious, the guidance will be slow and detailed; if they are relaxed, the guidance will be at a normal speed.

[0737] Examples of specific cases and prompt statements

[0738] Specific example

[0739] User: "Take me to the nearest milk stand."

[0740] Terminal message: "Go 20 meters and turn right."

[0741] User: While shopping, asked, "Please tell me how fresh this milk is."

[0742] Device: Captures camera footage, analyzes it on the server, and then sends a notification saying, "This milk is very fresh."

[0743] Example of a prompt

[0744] "Take me to the nearest milk stand."

[0745] "Please tell me how fresh this milk is."

[0746] "Please turn right next."

[0747] "There is an obstacle in the direction you are going."

[0748] "This product is extremely fresh."

[0749] This invention enables users to navigate and shop efficiently and safely in physical stores by receiving optimized guidance and support tailored to their emotional state. Furthermore, emotion recognition technology can reduce the psychological burden on users.

[0750] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0751] Step 1:

[0752] Users use their smartphone's voice input function to enter commands regarding destinations or products. This voice input is captured by the smartphone's built-in microphone. The input data consists of the user's voice, specifically commands such as "Take me to the nearby milk section" or "Tell me the freshness of this milk."

[0753] Step 2:

[0754] The voice input data is converted into text data using the terminal's speech_recognition library. At this stage, the input is the user's voice data, and the output is text data obtained by analyzing the voice. The text data will be in the format of "Take me to the nearest milk stand" or "Tell me how fresh this milk is."

[0755] Step 3:

[0756] The data, converted to text, is sent to the server, which analyzes the data to identify destination and product information. The input here is a text-converted instruction, which the server then analyzes. The output includes destination information (milk section) and product information (milk).

[0757] Step 4:

[0758] Once a destination is identified, the server uses map information to calculate the optimal route from the current location to the destination. The input data consists of the current location and destination information, and the server executes a route calculation algorithm based on this data. The output data generates detailed route information, such as "Go 20 meters and turn right."

[0759] Step 5:

[0760] The calculated route information is sent back to the terminal, which then uses a speech synthesis library (such as pyttsx3) to provide voice guidance. The input data is route information obtained from the server, which the terminal converts into voice data. The output data is the voice guidance that the user can hear. Specifically, it generates voice guidance such as, "Go 20 meters and turn right."

[0761] Step 6:

[0762] When a user selects a product, the camera is used to photograph the product, and the video data is sent to the server. The input data for this step is the video data captured by the camera. The video data is an image containing the product, specifically including images of milk, for example.

[0763] Step 7:

[0764] The server analyzes the received video data using deep learning techniques and image recognition tools (such as OpenCV) to identify products. The input data is video data transmitted from the camera, and the output data is information about the analyzed products. Specifically, it recognizes milk in the video and evaluates its quality and freshness.

[0765] Step 8:

[0766] Product information analyzed by the server undergoes a quality evaluation, and the evaluation results are sent back to the terminal. The input data is the product recognition result, and the server performs the evaluation using a quality evaluation algorithm. The output data is the freshness evaluation result, and information such as "This milk is very fresh" can be obtained.

[0767] Step 9:

[0768] The terminal uses speech synthesis technology (such as pyttsx3) to notify the user of the received evaluation results via voice. The input data is the freshness evaluation results returned from the server, which the terminal converts into voice data. The output data is a voice notification of the quality evaluation results that the user can hear.

[0769] Step 10:

[0770] The user's voice and facial expressions are captured and acquired by the device as data to recognize their emotional state. The input data consists of the user's voice and facial expressions, which the device captures.

[0771] Step 11:

[0772] The acquired emotion data is analyzed using emotion recognition technology (the pipeline function in the transformers library). The input data consists of voice and facial capture data, and the output data is the recognized emotional state. Specifically, emotional states such as "anxious" or "relaxed" can be obtained.

[0773] Step 12:

[0774] Based on the recognized emotional state, the speed and detail of the voice guidance are adjusted. The input data is the result of emotion recognition, and the terminal changes the voice guidance settings based on this. The output data is the adjusted voice guidance; for example, if the user is "anxious," slow and detailed guidance will be provided.

[0775] Through this series of steps, users can independently navigate and shop in physical stores, and receive optimal support tailored to their emotional state.

[0776] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0777] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0778] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0779] [Second Embodiment]

[0780] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0781] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0782] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0783] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0784] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0785] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0786] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0787] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0788] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0789] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0790] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0791] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0792] System Overview

[0793] This invention is an assistance system for visually impaired individuals to enable them to move around and perform daily shopping independently. This system combines voice input, speech synthesis, image recognition, GPS, map information, and AI technology. The embodiments are described below.

[0794] Navigation function

[0795] Enter destination

[0796] The user voice-inputs "Take me to the nearest convenience store" into their smartphone. The device uses speech recognition technology to convert the voice into text. This text information is sent to a server and processed as destination information.

[0797] Root calculation

[0798] The server calculates a route using map information based on the received text information (destination). Specifically, it uses an API to calculate the optimal walking route from the current location to the destination. The calculated route information is then returned to the device.

[0799] Voice directions

[0800] The terminal analyzes the calculated route information and provides voice instructions to the user regarding the next direction and distance to travel. For example, it might say, "Go 50 meters and turn right," providing step-by-step guidance.

[0801] Real-time video analysis

[0802] The device uses a camera to capture video of the user's surroundings in real time. This video data is sent to a server, which analyzes the video using deep learning technology. Specifically, it recognizes the status of traffic lights and obstacles. The results of this analysis are returned to the device, and the user is notified of the situation via voice, such as, "The traffic light is red. Please stop."

[0803] Shopping assistance

[0804] Product Selection

[0805] The user voice-inputs, "I want to buy tomatoes." The device then uses speech recognition technology to convert this voice into text and sends it to the server.

[0806] Product recognition

[0807] The server uses a model to recognize products from the terminal's camera footage based on the received text information (product information). When the user points the camera at a tomato shelf, the footage is sent to the server, and image analysis is performed using deep learning technology.

[0808] Freshness evaluation

[0809] The freshness and quality of tomatoes are evaluated based on image analysis. For example, the evaluation is based on color, shape, and gloss. The evaluation results are returned to the device, and the user is notified by voice, "This tomato is very fresh."

[0810] Specific example

[0811] Examples of directions

[0812] User: Inputs "Take me to the nearest convenience store" by voice.

[0813] Terminal: Converts speech to text and sends it to the server.

[0814] Server: Uses the Google Maps API to calculate the route from the current location to the destination and sends it back to the device.

[0815] Terminal: Provides voice guidance to the user regarding the calculated route information (e.g., "Go 50 meters and turn right").

[0816] Terminal: Captures camera footage and sends it to a server to recognize the status of traffic lights.

[0817] Server: Recognizes that the traffic light is red and notifies the terminal with the message, "The traffic light is red. Please stop."

[0818] Terminal: Notifies the user via voice.

[0819] Specific examples of shopping

[0820] User: Inputs "I want to buy tomatoes" by voice.

[0821] Terminal: Converts speech to text and sends it to the server.

[0822] Server: Uses a model for product recognition.

[0823] User: Point the camera at the tomato trellis.

[0824] Terminal: Captures camera footage and sends it to the server.

[0825] Server: Evaluates the freshness and quality of tomatoes and sends the results back to the terminal.

[0826] Device: Notifies the user of the evaluation results via voice (e.g., "This tomato is very fresh").

[0827] In this way, the system of the present invention can effectively provide support to visually impaired persons when they are able to move around and shop independently.

[0828] The following describes the processing flow.

[0829] Processing steps for the navigation function

[0830] Step 1:

[0831] The user uses voice input to tell the terminal, "Take me to the nearest convenience store."

[0832] Step 2:

[0833] The device captures voice input and uses speech recognition technology to convert the voice signal into text.

[0834] Step 3:

[0835] The terminal sends the converted text data to the server.

[0836] Step 4:

[0837] The server extracts destination information from the received text data and uses map information to calculate the route from the current location to the destination.

[0838] Step 5:

[0839] The server returns the calculated route information to the terminal.

[0840] Step 6:

[0841] The terminal analyzes route information and uses speech synthesis technology to instruct the user verbally on the next direction and distance to go. For example, it might say, "Go 50 meters and turn right."

[0842] Step 7:

[0843] The device uses its camera to capture images of its surroundings in real time.

[0844] Step 8:

[0845] The terminal sends the captured video to the server.

[0846] Step 9:

[0847] The server receives the video data and uses a deep learning model to analyze obstacles and traffic light conditions within the video.

[0848] Step 10:

[0849] The server sends the analysis results back to the terminal. For example, if the traffic light is red, it will send back an instruction such as, "The traffic light is red. Please stop."

[0850] Step 11:

[0851] The device notifies the user of the analysis results via voice. The notification will be in the form of, "The traffic light is red. Please stop."

[0852] Processing steps for the shopping support function

[0853] Step 1:

[0854] The user uses voice input to say "I want to buy tomatoes" to the device.

[0855] Step 2:

[0856] The device captures voice input and uses speech recognition technology to convert the voice signal into text.

[0857] Step 3:

[0858] The terminal sends the converted text data to the server.

[0859] Step 4:

[0860] The server extracts product information from the received text data and selects an appropriate image recognition model.

[0861] Step 5:

[0862] The user points the camera at the tomato trellis.

[0863] Step 6:

[0864] The device uses a camera to capture real-time video of the product shelves.

[0865] Step 7:

[0866] The terminal sends the captured video to the server.

[0867] Step 8:

[0868] The server receives the video data and uses an image recognition model to identify tomatoes in the video.

[0869] Step 9:

[0870] The server evaluates the freshness and quality of the identified tomatoes.

[0871] Step 10:

[0872] The server sends the evaluation results back to the terminal. For example, it might return an evaluation result such as, "This tomato is very fresh."

[0873] Step 11:

[0874] The device notifies the user of the evaluation results via voice. The notification may be in the form of, "This tomato is very fresh."

[0875] These detailed processing steps enable the system of the present invention to allow visually impaired individuals to independently receive support in various aspects of their daily lives.

[0876] (Example 1)

[0877] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0878] There is a lack of technical support for visually impaired individuals to move independently, safely, and comfortably, and to perform daily shopping. To address this challenge, a system is needed that recognizes destinations and products through voice input and provides appropriate instructions through voice guidance and video analysis. Furthermore, real-time updated navigation information and product freshness assessments based on this system are also required.

[0879] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0880] In this invention, the server includes means for receiving voice input from a user, means for converting the received voice input into text, means for identifying a destination based on the converted text and calculating a route using map information, means for communicating the calculated route information to the user by voice, means for capturing and analyzing video of the user's surroundings, means for recognizing obstacles and traffic light conditions in the captured video, and means for notifying the user of the recognition results by voice. This enables visually impaired individuals to recognize destinations and products through voice input and to move and shop safely while receiving appropriate guidance and instructions in real time.

[0881] "Means for receiving user voice input" refers to a device or program that allows visually impaired individuals to input voice commands using a smartphone or other device.

[0882] "Means for converting received voice input into text" refers to a device or program that uses speech recognition technology to convert voice data into text data.

[0883] "Means for identifying a destination based on converted text and calculating a route using map information" refers to a device or program that analyzes text data, refers to a map database based on a specified destination, and calculates the optimal route.

[0884] "Means of conveying calculated route information to users via voice" refers to a device or program that provides calculated route information as a voice message using speech synthesis technology.

[0885] "Means for capturing and analyzing video footage of the user's surroundings" refers to a device or program that uses a camera to capture images of the user's surroundings in real time and analyzes that image data.

[0886] "Means for recognizing obstacles and traffic light status in captured video" refers to a device or program that uses deep learning technology to recognize information about obstacles and traffic lights from acquired video data.

[0887] "Means for notifying the user of the recognition results by voice" refers to a device or program that provides the recognized information as a voice message using speech synthesis technology.

[0888] "A means of inputting a product name by voice, analyzing that information, and converting it into text" refers to a device or program that allows a user to input a product name by voice and convert that information into text data.

[0889] "Means for recognizing products from camera footage" refers to a device or program that uses deep learning technology to identify specific products from video data acquired using a camera.

[0890] "Means for evaluating the freshness and quality of recognized goods" refers to a device or program that evaluates the freshness and quality of recognized goods based on criteria such as color, shape, and gloss.

[0891] "Means for notifying users of evaluation results by voice" refers to a device or program that provides evaluation results of freshness and quality as a voice message using speech synthesis technology.

[0892] "Means for updating routes in real time" refers to a device or program that dynamically updates the navigation route for each action based on the user's current location information.

[0893] "Means of providing voice instructions for the direction of travel or the next action" refers to a device or program that uses speech synthesis technology to provide the user with voice messages indicating the direction and action they should take next.

[0894] "Means for recognizing moving objects such as pedestrians and animals in the surrounding area and notifying the user of that information via voice" refers to a device or program that recognizes moving objects from camera footage and provides information for avoiding danger as a voice message.

[0895] System Overview

[0896] This invention is an assistance system for visually impaired individuals to enable them to move around and perform daily shopping independently. This system is comprised of a combination of voice input, speech synthesis, image recognition, GPS, map information, and AI technology.

[0897] Navigation function

[0898] Enter destination

[0899] The user inputs a voice command into their smartphone saying, "Take me to the nearest convenience store."

[0900] The device uses speech recognition technology to convert this speech into text. Specifically, it uses Google Cloud Speech-to-Text.

[0901] The device sends the generated text data ("Please guide me to the nearest convenience store") to the server.

[0902] The server receives the text data and uses it to determine the destination.

[0903] Root calculation

[0904] The server determines the user's current location based on the destination information received.

[0905] The server uses map information and the Google Maps API to calculate the optimal walking route from the current location to the destination.

[0906] The server generates the calculated route information in JSON format and sends it back to the terminal.

[0907] Voice directions

[0908] The terminal parses the received route information in JSON format.

[0909] To guide the user verbally in the next direction and distance, a voice message is generated using Google Cloud Text-to-Speech.

[0910] For example, based on route information, it might instruct you to "Go 50 meters and turn right."

[0911] Real-time video analysis

[0912] The device activates its camera and captures video of the user's surroundings in real time.

[0913] The terminal sends the captured video data to the server.

[0914] The server analyzes the received video using deep learning technology (TensorFlow). It recognizes traffic lights and obstacles, generates that information in text format, and sends it back to the terminal.

[0915] The device generates a voice message using speech synthesis technology (Google Cloud Text-to-Speech) based on the received recognition results and notifies the user.

[0916] For example, it might announce, "The traffic light is red. Please stop."

[0917] Shopping assistance

[0918] Product Selection

[0919] The user voice-inputs "I want to buy tomatoes" into their smartphone.

[0920] The device uses speech recognition technology (Google Cloud Speech-to-Text) to convert this speech into text.

[0921] The terminal sends the generated text data ("I want to buy tomatoes") to the server.

[0922] The server receives the text data and processes it as product information.

[0923] Product recognition

[0924] The server uses a deep learning model to recognize tomatoes based on product information.

[0925] The user points the camera at the tomato trellis.

[0926] The device captures camera footage and sends it to the server.

[0927] The server analyzes the received video using deep learning technology (TensorFlow) to recognize tomatoes.

[0928] Freshness evaluation

[0929] The server evaluates the freshness and quality of the recognized tomatoes. Specifically, it evaluates them based on criteria such as color, shape, and gloss.

[0930] The server generates the evaluation results in text format and sends them back to the terminal.

[0931] The device generates a voice message saying, "These tomatoes are very fresh," and notifies the user.

[0932] Specific example

[0933] Examples of directions

[0934] User: Inputs "Take me to the nearest convenience store" by voice.

[0935] Terminal: Use Google Cloud Speech-to-Text to convert speech to text and send it to the server.

[0936] Server: Uses the Google Maps API to calculate the route from the current location to the destination and returns it to the device in JSON format.

[0937] Terminal: Analyzes the received route information and uses Google Cloud Text-to-Speech to provide voice guidance such as, "Go 50 meters and turn right."

[0938] Terminal: Captures camera footage and sends it to a server to recognize the status of traffic lights.

[0939] Server: Uses deep learning technology (TensorFlow) to recognize that the signal is red and sends a message back to the terminal.

[0940] Terminal: Based on the received information, it notifies the user by voice, "The traffic light is red. Please stop."

[0941] Specific examples of shopping

[0942] User: Inputs "I want to buy tomatoes" by voice.

[0943] Terminal: Use Google Cloud Speech-to-Text to convert speech to text and send it to the server.

[0944] Server: Prepares to recognize tomatoes using deep learning technology.

[0945] User: Point the camera at the tomato trellis.

[0946] Terminal: Captures camera footage and sends it to the server.

[0947] Server: Uses deep learning technology (TensorFlow) to analyze received video and evaluate the freshness and quality of tomatoes.

[0948] Server: Generates evaluation results in text format and sends them back to the terminal.

[0949] Device: Use Google Cloud Text-to-Speech to announce, "These tomatoes are very fresh."

[0950] In this way, the system of the present invention can effectively provide support to visually impaired persons when they are able to move around and shop independently.

[0951] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0952] Specific processing flow of the navigation function

[0953] Enter destination

[0954] Step 1:

[0955] The user uses voice input on their smartphone, saying, "Take me to the nearest convenience store."

[0956] Input: Voice command ("Take me to the nearest convenience store.")

[0957] Output: Audio data

[0958] Step 2:

[0959] The device uses speech recognition technology (Google Cloud Speech-to-Text) to convert the input voice data into text data.

[0960] Specifically, audio data is sent to the server, and the server returns text data.

[0961] Input: Audio data

[0962] Output: Text data ("Please guide me to the nearest convenience store")

[0963] Step 3:

[0964] The terminal sends the generated text data to the server. The server receives this data and processes it as destination information.

[0965] Input: Text data ("Please guide me to the nearest convenience store")

[0966] Output: Destination information

[0967] Root calculation

[0968] Step 4:

[0969] The server determines the user's current location based on the destination information received.

[0970] Input: Destination information

[0971] Output: Current location information

[0972] Step 5:

[0973] The server uses the Google Maps API to calculate the optimal walking route from the current location to the destination.

[0974] Input: Current location information, destination information

[0975] Output: Route information (JSON format)

[0976] Step 6:

[0977] The server generates the calculated route information in JSON format and sends it back to the terminal.

[0978] Input: Route information (in generated JSON format)

[0979] Output: Route information (JSON format)

[0980] Voice directions

[0981] Step 7:

[0982] The terminal analyzes the received route information and calculates the next direction and distance to proceed.

[0983] Input: Route information (JSON format)

[0984] Output: Guidance / Instructions

[0985] Step 8:

[0986] The device uses Google Cloud Text-to-Speech to generate guidance instructions as voice messages and notify the user.

[0987] For example, you might give directions like, "Go 50 meters and turn right."

[0988] Input: Guidance / Instructions

[0989] Output: Voice message

[0990] Real-time video analysis

[0991] Step 9:

[0992] The device activates its camera and captures video of the user's surroundings in real time.

[0993] Input: Camera video

[0994] Output: Video data

[0995] Step 10:

[0996] The terminal sends the captured video data to the server.

[0997] Input: Video data

[0998] Output: Video data sent to the server

[0999] Step 11:

[1000] The server uses deep learning technology (TensorFlow) to analyze the received video and recognize the status of traffic lights and obstacles.

[1001] Input: Video data

[1002] Output: Analysis results (status of traffic lights and obstacles)

[1003] Step 12:

[1004] The server generates the analysis results in text format and sends them back to the terminal.

[1005] Input: Analysis results

[1006] Output: Text data (analysis results)

[1007] Step 13:

[1008] The device generates a voice message using Google Cloud Text-to-Speech based on the received recognition results and notifies the user.

[1009] For example, you might announce, "The traffic light is red. Please stop."

[1010] Input: Text data (analysis results)

[1011] Output: Voice message

[1012] Specific processing flow for shopping assistance

[1013] Product Selection

[1014] Step 1:

[1015] The user voice-inputs "I want to buy tomatoes" into their smartphone.

[1016] Input: Voice command ("I want to buy tomatoes")

[1017] Output: Audio data

[1018] Step 2:

[1019] The device uses speech recognition technology (Google Cloud Speech-to-Text) to convert this speech into text.

[1020] Specifically, audio data is sent to the server, and the server returns text data.

[1021] Input: Audio data

[1022] Output: Text data ("I want to buy tomatoes")

[1023] Step 3:

[1024] The terminal sends the generated text data to the server. The server receives this data and processes it as product information.

[1025] Input: Text data ("I want to buy tomatoes")

[1026] Output: Product Information

[1027] Product recognition

[1028] Step 4:

[1029] The server prepares a product recognition model based on the product information.

[1030] Input: Product Information

[1031] Output: Recognition Model

[1032] Step 5:

[1033] The user points the camera at the tomato trellis.

[1034] Input: Image of a product shelf

[1035] Output: Camera data

[1036] Step 6:

[1037] The device captures camera footage and sends it to the server.

[1038] Input: Camera data

[1039] Output: Video data sent to the server

[1040] Step 7:

[1041] The server uses deep learning technology (TensorFlow) to analyze the received video and recognize the product.

[1042] Input: Video data

[1043] Output: Analysis results (product recognition)

[1044] Step 8:

[1045] The server generates the recognition results in text format and sends them back to the terminal.

[1046] Input: Analysis results

[1047] Output: Text data (product recognition)

[1048] Freshness evaluation

[1049] Step 9:

[1050] The server evaluates the freshness and quality of the recognized products. Specifically, it evaluates them based on criteria such as color, shape, and gloss.

[1051] Input: Product recognition result

[1052] Output: Freshness evaluation results

[1053] Step 10:

[1054] The server generates the evaluation results in text format and sends them back to the terminal.

[1055] Input: Freshness evaluation results

[1056] Output: Text data (evaluation results)

[1057] Step 11:

[1058] The device uses Google Cloud Text-to-Speech to generate the evaluation results as an audio message and notify the user.

[1059] For example, you might say, "These tomatoes are extremely fresh."

[1060] Input: Text data (evaluation results)

[1061] Output: Voice message

[1062] (Application Example 1)

[1063] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[1064] Visually impaired individuals face numerous challenges when attempting to travel and shop independently. In particular, accurately perceiving their surroundings is difficult when seeking directions to a destination or purchasing goods. They struggle to recognize traffic signals, pedestrians, animals, and other moving objects. Furthermore, locating products and assessing their freshness and quality is extremely challenging. Effective support systems are needed to address these issues.

[1065] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1066] In this invention, the server includes: means for receiving voice input from the user; means for converting the received voice input into text; means for identifying a destination based on the converted text and calculating a route using map information; means for communicating the calculated route information to the user by voice; means for capturing and analyzing video of the user's surroundings; means for recognizing obstacles and traffic light conditions in the captured video; means for notifying the user of the recognition results by voice; means for the user to input a product name by voice, analyzing that information and converting it into text; means for recognizing a product from the camera image based on the converted text; and means for providing information about the recognized product. The system includes means for evaluating degree and quality; means for notifying the user of the evaluation results by voice; means for identifying the user's current location and updating the route in real time using map information; means for instructing the user on the direction of travel and the next action by voice based on the updated route information; means for recognizing moving objects such as pedestrians and animals in the surroundings other than obstacles and traffic lights, and notifying the user of that information by voice; means for the user to input searched products by voice, and for analyzing that information and converting it into text; means for calculating the evaluation results of product candidates based on the converted text using deep learning technology; and means for guiding the user to the desired product in real time based on the evaluation results. This enables visually impaired people to move and shop independently, safely and effectively.

[1067] A "system" refers to the entire apparatus, including a series of means, designed to support visually impaired individuals in independently navigating and shopping.

[1068] "Users" refers to visually impaired individuals who use this system to receive guidance and shopping assistance.

[1069] "Voice input" refers to the voice signals that a user speaks to the system.

[1070] "Means for receiving voice input" refers to devices or software that acquire the user's voice and allow the system to interpret it.

[1071] "Means of converting speech to text" refers to devices or software that analyze speech input and convert it into textual information.

[1072] "Means of identifying a destination based on text" refers to devices or software that determine a destination based on converted character information.

[1073] "Map information" refers to a database containing geographical information, providing destinations, current location, route information, and more.

[1074] "Means of calculating routes" refers to devices or software that use map information to calculate the optimal route to a destination.

[1075] "Means of conveying route information by voice" refers to devices or software that guide users through calculated routes by voice.

[1076] "Means of capturing video of the user's surroundings" refers to devices that acquire images or videos of the user's surroundings using cameras or similar devices.

[1077] "Means of analyzing video" refers to devices or software that analyze captured video to extract useful information.

[1078] "Means for recognizing obstacles and traffic signal status" refers to devices and software that detect the status of surrounding obstacles and traffic signals through video analysis.

[1079] "Means of notifying the user of recognition results by voice" refers to devices or software that inform the user of the analysis results by voice.

[1080] "Methods for inputting product names by voice" refer to devices or software that allow users to tell the system by voice the products they wish to purchase.

[1081] "Means of recognizing products from camera footage" refers to devices or software that use a camera to acquire images of products and then recognize those products.

[1082] "Means for evaluating freshness and quality" refers to devices or software used to evaluate the condition of recognized products.

[1083] "Means for determining current location" refers to devices or software used to determine the user's current geographical location.

[1084] "Means of updating routes in real time" refers to devices or software that continuously update route guidance in accordance with the user's movements.

[1085] "Means of providing voice instructions for direction of travel or next action" refers to devices or software that provide voice guidance for the next action based on updated route information.

[1086] "Means of recognizing moving objects such as pedestrians and animals" refers to devices and software used to detect moving objects in the surrounding environment.

[1087] "Methods for inputting searched products by voice" refers to devices or software that allow users to tell the system by voice what product they are looking for.

[1088] "Methods for calculating evaluation results of candidate products using deep learning technology" refers to devices or software that analyze and evaluate the characteristics of candidate products using deep learning.

[1089] "A means of guiding users to desired products in real time based on evaluation results" refers to devices or software that guide users to products they are looking for based on evaluation results.

[1090] System Configuration

[1091] To implement this invention, a server, terminals, and users are required. The terminals consist of smartphones or devices with built-in cameras. The server is responsible for data processing and analysis and operates in a cloud computing environment. The system integrates multiple functions, including voice input recognition, text conversion, route guidance using map information, real-time video analysis, product recognition, and quality evaluation.

[1092] Program and Processing Overview

[1093] Voice input acceptance and text conversion:

[1094] The user provides voice input. For example, they might say, "Take me to the nearest convenience store." The device receives this input and uses speech recognition technology to convert it into text. The Google Speech-to-Text API is used for this process.

[1095] Destination identification and route calculation:

[1096] The converted audio data is sent to a server, which identifies the destination based on the converted text. It then uses map information to calculate the optimal route. The Google Maps API is used here. The calculated route information is returned to the device.

[1097] Voice-guided route guidance:

[1098] The terminal analyzes route information calculated using speech synthesis technology and provides voice instructions to the user regarding the next direction and distance to travel. For example, it might say, "Go 50 meters and turn right."

[1099] Real-time video capture and analysis:

[1100] The device's camera captures video of the user's surroundings in real time and sends the data to a server. The server uses deep learning technology to analyze the video data and recognize traffic lights and obstacles. The results of this analysis are returned to the device and notified to the user via voice.

[1101] Product recognition and quality assessment:

[1102] The user voice-inputs "I want to buy tomatoes." This voice is converted back into text, and the server uses a model that recognizes products from camera footage based on the text information. When the user points the camera at a shelf of tomatoes, the image is sent to the server. Deep learning technology is used to evaluate the freshness and quality of the tomatoes, and the results are returned to the terminal. For example, the user might receive a voice notification saying, "These tomatoes are very fresh."

[1103] Current location and real-time route updates:

[1104] The server uses GPS to determine the user's current location and updates the route in real time based on map information. Based on the updated route information, the terminal provides voice instructions to the user regarding the direction to go and the next action.

[1105] Recognition and notification of moving objects:

[1106] The server uses video analysis to recognize moving objects in the surrounding area, such as pedestrians and animals, other than obstacles and traffic lights, and notifies the user of this information via voice.

[1107] Specific example

[1108] 1. Specific Examples of Navigation

[1109] User: "Take me to the nearest convenience store."

[1110] Terminal: Converts speech to text and sends it to the server.

[1111] Server: Route calculation using Google Maps API

[1112] Device: Voice-guided route instructions (e.g., "Go 50 meters and turn right")

[1113] Terminal: Captures camera footage and sends the traffic light status to the server.

[1114] Server: Recognizes the status of the traffic light and notifies, "The traffic light is red. Please stop."

[1115] Device: Notifies the user via voice.

[1116] 2. Specific examples of shopping

[1117] User: "I want to buy tomatoes."

[1118] Terminal: Converts speech to text and sends it to the server.

[1119] Server: Uses a model for product recognition.

[1120] User: Pointing the camera at the tomato trellis.

[1121] Terminal: Captures camera footage and sends it to the server.

[1122] Server: Evaluates the freshness and quality of the tomatoes and notifies, "These tomatoes are very fresh."

[1123] Device: Notifies the user via voice.

[1124] This system assists users in moving around and shopping safely and conveniently. Its purpose is to provide an environment where visually impaired individuals can live independently, utilizing technologies such as voice recognition, deep learning, and GPS.

[1125] Example of a prompt

[1126] 1. Navigation prompt text:

[1127] Please guide me to the nearest convenience store.

[1128] 2. Product purchase prompt message:

[1129] I bought tomatoes.

[1130] As described above, this invention provides comprehensive support for visually impaired individuals to move around and shop independently.

[1131] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1132] Flow of processing steps

[1133] Navigation function

[1134] Step 1:

[1135] Input: The user inputs "Take me to the nearest convenience store" by voice.

[1136] Operation: The device (smartphone) accepts voice input.

[1137] Output: Audio data.

[1138] Step 2:

[1139] Input: Audio data.

[1140] Operation: The device uses speech recognition technology to convert speech data into text. The Google Speech-to-Text API is used here.

[1141] Output: Text data.

[1142] Step 3:

[1143] Input: Text data.

[1144] Operation: The server identifies the destination based on text data and calculates the route using the Google Maps API.

[1145] Output: Route information.

[1146] Step 4:

[1147] Input: Route information.

[1148] Operation: The server sends route information back to the terminal, and the terminal parses this information.

[1149] Output: Route guidance information after analysis.

[1150] Step 5:

[1151] Input: Route guidance information after analysis.

[1152] Operation: The device uses speech synthesis technology to provide instructions to the user. For example, it might say, "Go 50 meters and turn right."

[1153] Output: Voice guidance.

[1154] Step 6:

[1155] Input: Video of the user's surroundings.

[1156] Operation: The device's camera captures video and sends it to the server.

[1157] Output: Video data.

[1158] Step 7:

[1159] Input: Video data.

[1160] Operation: The server uses deep learning technology to analyze video and recognize traffic lights and obstacles.

[1161] Output: Recognition result.

[1162] Step 8:

[1163] Input: Recognition result.

[1164] Operation: The server sends the recognition result back to the terminal, and the terminal notifies the user by voice, "The traffic light is red. Please stop."

[1165] Output: Voice notification.

[1166] Shopping support function

[1167] Step 1:

[1168] Input: The user enters "I want to buy tomatoes" by voice.

[1169] Operation: The device (smartphone) accepts voice input.

[1170] Output: Audio data.

[1171] Step 2:

[1172] Input: Audio data.

[1173] Operation: The device uses speech recognition technology to convert speech data into text. The Google Speech-to-Text API is used here.

[1174] Output: Text data.

[1175] Step 3:

[1176] Input: Text data.

[1177] Operation: The server prepares a product recognition model based on text data.

[1178] Output: Recognition model.

[1179] Step 4:

[1180] Input: Camera footage.

[1181] Operation: The user points their smartphone camera at a product shelf (e.g., a tomato shelf). This video data is sent from the device to the server.

[1182] Output: Video data.

[1183] Step 5:

[1184] Input: Video data.

[1185] Operation: The server uses deep learning technology to recognize products from images and evaluate the freshness and quality of tomatoes.

[1186] Output: Evaluation results.

[1187] Step 6:

[1188] Input: Evaluation results.

[1189] Operation: The server sends the evaluation result back to the terminal, and the terminal notifies the user by voice, "This tomato is very fresh."

[1190] Output: Voice notification.

[1191] Real-time route update function

[1192] Step 1:

[1193] Input: Current location.

[1194] Operation: The device's GPS function determines the user's current location and sends it to the server.

[1195] Output: Current location information.

[1196] Step 2:

[1197] Input: Current location information.

[1198] Operation: The server updates the route in real time using map information.

[1199] Output: Updated route information.

[1200] Step 3:

[1201] Input: Updated route information.

[1202] Operation: The device provides voice instructions to the user regarding the direction of travel and the next action based on updated route information.

[1203] Output: Voice command.

[1204] Recognition function for moving objects

[1205] Step 1:

[1206] Input: Video of the user's surroundings.

[1207] Operation: The device's camera captures the surrounding video and sends it to the server.

[1208] Output: Video data.

[1209] Step 2:

[1210] Input: Video data.

[1211] Operation: The server uses deep learning technology to analyze video and recognize moving objects such as pedestrians and animals.

[1212] Output: Recognition result.

[1213] Step 3:

[1214] Input: Recognition result.

[1215] Operation: The server sends the recognition result back to the terminal, and the terminal notifies the user by voice, "There is a pedestrian ahead. Please be careful."

[1216] Output: Voice notification.

[1217] The above outlines the specific processing steps of this system. This will enable visually impaired individuals to move around and shop independently.

[1218] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1219] System Overview

[1220] This invention is an assistance system for visually impaired individuals to independently navigate and perform daily shopping. By incorporating an emotion engine, it provides optimal support based on the user's current emotional state. The system combines voice input, speech synthesis, image recognition, GPS, map information, AI technology, and emotion recognition technology.

[1221] Navigation function

[1222] Enter destination

[1223] The user speaks into their smartphone and says, "Take me to the nearest convenience store." The device uses speech recognition technology to convert the speech into text. This text information is sent to a server and processed as destination information.

[1224] Root calculation

[1225] The server extracts destination information from the received text information and calculates the route from the current location to the destination using map information. It then sends the calculated route information back to the terminal.

[1226] Voice directions

[1227] The terminal analyzes the calculated route information and uses speech synthesis technology to instruct the user on the next direction and distance to go. For example, it provides step-by-step instructions such as, "Go 50 meters and turn right."

[1228] Real-time video analysis

[1229] The device uses a camera to capture video of the user's surroundings in real time. This video data is sent to a server, which analyzes the video using deep learning technology. Specifically, it recognizes the status of traffic lights and obstacles. The results of this analysis are sent back to the device, and the user is notified of the situation via voice, such as, "The traffic light is red. Please stop."

[1230] Emotion recognition and regulation

[1231] The device captures the user's voice and facial expressions and uses an emotion engine to recognize their current emotional state. For example, if the user is anxious, the guidance voice will slow down and provide detailed instructions to create a sense of security. Conversely, if the user is relaxed, the guidance will be given at a normal speed.

[1232] Shopping assistance

[1233] Product Selection

[1234] The user voice-inputs, "I want to buy tomatoes." The device uses speech recognition technology to convert the voice into text and sends it to the server.

[1235] Product recognition

[1236] The server uses a model to recognize products from the terminal's camera footage based on the received text information. When the user points the camera at a tomato shelf, the footage is sent to the server, and image analysis is performed using deep learning technology.

[1237] Freshness evaluation

[1238] The freshness and quality of tomatoes are evaluated based on image analysis. For example, the evaluation is based on color, shape, and gloss. The evaluation results are sent back to the device, and the user is notified by voice, "This tomato is very fresh."

[1239] Emotion recognition and product recommendations

[1240] The device captures the user's voice and facial expressions and uses an emotion engine to recognize their current emotional state. For example, if the user is undecided, it suggests other items on the same shelf to broaden their options. On the other hand, if the user is in a hurry, it quickly suggests the highest-rated product.

[1241] Specific example

[1242] Examples of directions

[1243] User: Inputs "Take me to the nearest convenience store" by voice.

[1244] Terminal: Converts speech to text and sends it to the server.

[1245] Server: Calculates the route from the current location to the destination and sends it back to the terminal.

[1246] Terminal: Provides route information to the user via voice (e.g., "Go 50 meters and turn right").

[1247] Terminal: Captures camera footage and sends it to the server to analyze the status of the traffic lights.

[1248] Server: Recognizes that the traffic light is red and notifies, "The traffic light is red. Please stop."

[1249] Terminal: Notifies the user via voice.

[1250] Terminal: If the user is anxious, slow down the voice guidance and add more detailed explanations.

[1251] Specific examples of shopping

[1252] User: Inputs "I want to buy tomatoes" by voice.

[1253] Terminal: Converts speech to text and sends it to the server.

[1254] Server: Uses a model for product recognition.

[1255] User: Point the camera at the tomato trellis.

[1256] Terminal: Captures camera footage and sends it to the server.

[1257] Server: Evaluates the freshness and quality of tomatoes and sends the results back to the terminal.

[1258] Device: Notifies the user of the evaluation results via voice (e.g., "This tomato is very fresh").

[1259] Terminal: If the user is undecided, suggest other products.

[1260] In this way, the system of the present invention can provide optimal support tailored to the emotional state of visually impaired individuals when they are independently moving around or shopping.

[1261] The following describes the processing flow.

[1262] Processing steps for the navigation function

[1263] Step 1:

[1264] The user uses voice input to tell the terminal, "Take me to the nearest convenience store."

[1265] Step 2:

[1266] The device captures voice input and uses speech recognition technology to convert the voice signal into text.

[1267] Step 3:

[1268] The terminal sends the converted text data to the server.

[1269] Step 4:

[1270] The server extracts destination information from the received text data and uses map information to calculate the route from the current location to the destination.

[1271] Step 5:

[1272] The server returns the calculated route information to the terminal.

[1273] Step 6:

[1274] The terminal analyzes route information and uses speech synthesis technology to instruct the user verbally on the next direction and distance to go. For example, it might say, "Go 50 meters and turn right."

[1275] Step 7:

[1276] The device uses its camera to capture images of its surroundings in real time.

[1277] Step 8:

[1278] The terminal sends the captured video to the server.

[1279] Step 9:

[1280] The server receives the video data and uses a deep learning model to analyze obstacles and traffic light conditions within the video.

[1281] Step 10:

[1282] The server sends the analysis results back to the terminal. For example, if the traffic light is red, it will send back an instruction such as, "The traffic light is red. Please stop."

[1283] Step 11:

[1284] The device notifies the user of the analysis results via voice. The notification will be in the form of, "The traffic light is red. Please stop."

[1285] Step 12:

[1286] The device captures the user's voice and facial expressions and uses an emotion engine to recognize their current emotional state.

[1287] Step 13:

[1288] The device adjusts the speed and content of voice guidance based on the recognized emotional state. For example, if the user is anxious, the voice guidance will slow down and more detailed instructions will be added.

[1289] Processing steps for the shopping support function

[1290] Step 1:

[1291] The user uses voice input to say "I want to buy tomatoes" to the device.

[1292] Step 2:

[1293] The device captures voice input and uses speech recognition technology to convert the voice signal into text.

[1294] Step 3:

[1295] The terminal sends the converted text data to the server.

[1296] Step 4:

[1297] The server extracts corresponding product information based on the received text data.

[1298] Step 5:

[1299] The user points the camera at the tomato trellis.

[1300] Step 6:

[1301] The terminal captures video of the product shelves in real time and sends that video to the server.

[1302] Step 7:

[1303] The server analyzes the received video data and identifies the recognized products using an image recognition model.

[1304] Step 8:

[1305] The server evaluates the freshness and quality of the products.

[1306] Step 9:

[1307] The server sends the evaluation results back to the terminal. For example, it might return an evaluation result such as, "This tomato is very fresh."

[1308] Step 10:

[1309] The device notifies the user of the evaluation results via voice. The notification may be in the form of, "This tomato is very fresh."

[1310] Step 11:

[1311] The device captures the user's voice and facial expressions and uses an emotion engine to recognize their current emotional state.

[1312] Step 12:

[1313] The device makes product suggestions based on the recognized emotional state. For example, if the user is undecided, it adds suggestions for other products. On the other hand, if the user is in a hurry, it quickly suggests highly-rated products.

[1314] These detailed processing steps enable the system of the present invention to allow visually impaired individuals to independently receive support in various aspects of their daily lives. Furthermore, it detects the user's emotional state and provides optimal support accordingly.

[1315] (Example 2)

[1316] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[1317] There is a need to ensure that visually impaired individuals receive appropriate support when they are able to move around independently and do their daily shopping. In particular, flexible responses that take into account the user's emotional state are necessary in areas such as navigation to destinations, product selection, and recognition of obstacles along the way. However, current systems have difficulty meeting all of these requirements simultaneously, which has been a problem as it makes it inconvenient for visually impaired individuals to move around and shop with peace of mind.

[1318] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1319] In this invention, the server includes means for receiving voice input from a user, means for converting the received voice input into text, means for identifying a destination based on the converted text and calculating a route using map information, means for communicating the calculated route information to the user by voice, means for capturing and analyzing video of the user's surroundings, means for recognizing obstacles and traffic light conditions in the captured video, means for notifying the user of the recognition results by voice, and means for recognizing the user's emotional state and adjusting voice guidance based on that emotion. This makes it possible for visually impaired people to receive optimal support in real time, tailored to environmental information and their emotional state, when they are able to move around and shop independently with peace of mind.

[1320] "Means for receiving user voice input" refers to a mechanism for capturing the voice spoken by the user into the device.

[1321] "Means of converting received voice input into text" refers to the process of converting voice data into text data.

[1322] "A means of identifying a destination based on converted text and calculating a route using map information" refers to a process that analyzes text data to determine the destination and then calculates a route using map information.

[1323] "A means of communicating calculated route information to the user via voice" refers to a mechanism that informs the user of calculated route information via voice.

[1324] "Means for capturing and analyzing video footage of the user's surroundings" refers to a mechanism that uses a camera to capture video footage of the user's surroundings and then analyzes that footage.

[1325] "Means for recognizing obstacles and traffic light conditions in captured video" refers to technologies for identifying the condition of obstacles and traffic lights from captured video footage.

[1326] "Means for notifying the user of recognition results by voice" refers to a mechanism that communicates the analysis results to the user by voice.

[1327] "Means for recognizing the user's emotional state and adjusting voice guidance based on those emotions" refers to technology that analyzes the user's emotions and adjusts the content and speed of voice guidance accordingly.

[1328] "A method for inputting product names by voice, analyzing that information, and converting it into text" refers to a method in which a user inputs a product name by voice, and that voice is analyzed and converted into text data.

[1329] "Means for recognizing products from camera footage based on converted text" refers to a technology that uses converted text to identify specific products from camera footage.

[1330] "Means for evaluating the freshness and quality of recognized products" refers to the process of analyzing the freshness and quality of identified products.

[1331] "Means of notifying users of evaluation results by voice" refers to a mechanism that informs users of the product evaluation results by voice.

[1332] "A means of identifying the user's current location and updating the route in real time using map information" refers to a technology that obtains the user's current location and recalculates the route in real time based on that location information.

[1333] "A means of instructing users on their direction of travel and next actions via voice based on updated route information" refers to a voice guidance system that uses recalculated route information to provide instructions to users.

[1334] "A means of recognizing moving objects in the surroundings, such as pedestrians and animals other than obstacles and traffic lights, and notifying the user of that information by voice" refers to a technology that identifies moving objects in the surroundings and transmits that information to the user by voice.

[1335] This invention is an assistance system for visually impaired individuals to independently navigate and perform daily shopping, providing optimal support based on the user's emotional state. This system is realized by combining voice input, speech synthesis, image recognition, GPS, map information, AI technology, and emotion recognition technology.

[1336] System Configuration

[1337] The system mainly consists of the following components:

[1338] Speech recognition technology

[1339] Speech synthesis technology

[1340] Image recognition technology

[1341] GPS

[1342] Map information

[1343] emotion recognition technology

[1344] server

[1345] Devices (smartphones and mobile devices)

[1346] Voice input and conversion

[1347] The user provides voice input. For example, they might say, "Take me to the nearest convenience store." This voice data is captured by the device. The device uses speech recognition technology (e.g., Google Cloud Speech-to-Text) to convert the voice data into text data. The converted text data is sent to a server and processed as destination information.

[1348] Route calculation and communication

[1349] The server analyzes the received text data and extracts destination information. Then, using map information (such as the Google Maps API), the server calculates the optimal route from the user's current location to the destination and sends the calculated route information back to the device. The device then uses speech synthesis technology (for example, Amazon Polly) to communicate the calculated route information to the user verbally.

[1350] Surrounding video analysis

[1351] The device uses its camera to capture video of the user's surroundings in real time and sends the video data to a server. The server analyzes the video data using deep learning technology (such as TensorFlow). This analysis allows for the recognition of traffic lights and obstacles. The analysis results are sent back from the server to the device, and the user is notified of the situation via voice. For example, it might say, "The traffic light is red. Please stop."

[1352] Emotion recognition and regulation

[1353] The device captures the user's voice and facial expressions through its camera and microphone. The captured data is analyzed using emotion recognition technology (e.g., IBM Watson Tone Analyzer). Based on the analysis results, the user's emotional state is determined. For example, if the user is anxious, the speed of the voice guidance is slowed down, and more detailed instructions are added.

[1354] Shopping assistance

[1355] The user voice-inputs, "I want to buy tomatoes." The terminal uses speech recognition technology to convert the voice into text and sends it to the server. Based on the received text information, the server recognizes the products from the terminal's camera image. When the user points the camera at the tomato shelf, the image is sent to the server and analyzed using deep learning technology. The freshness and quality of the recognized tomatoes are evaluated, and the evaluation results are sent back to the terminal. The terminal then notifies the user by voice, "These tomatoes are very fresh."

[1356] Emotion recognition and product recommendations

[1357] The device then captures the user's voice and facial expressions and analyzes them using emotion recognition technology. For example, if the user is unsure, it suggests other products on the same shelf. On the other hand, if the user is in a hurry, it quickly suggests the highest-rated product.

[1358] Specific example

[1359] Examples of directions

[1360] User: "Take me to the nearest convenience store" (voice input).

[1361] Terminal: Converts audio data to text and sends it to the server.

[1362] Server: Extracts destination information, calculates the optimal route, and sends it back to the terminal.

[1363] Terminal: Provides calculated route information via voice (e.g., "Go 50 meters and turn right").

[1364] Terminal: Captures camera footage and sends it to the server.

[1365] Server: Analyzes video and recognizes the signal status.

[1366] Terminal: Announces analysis results by voice (e.g., "The traffic light is red. Please stop.").

[1367] Terminal: Adjusts the content of voice guidance according to the user's emotional state.

[1368] Specific examples of shopping

[1369] User: "I want to buy tomatoes" (voice input).

[1370] Terminal: Converts speech to text and sends it to the server.

[1371] Server: Uses a product recognition model, and the user points their camera at a tomato shelf.

[1372] Terminal: Captures video and sends it to the server.

[1373] Server: Analyzes video footage to evaluate the freshness of tomatoes.

[1374] Device: Notifies the evaluation result by voice (e.g., "This tomato is very fresh").

[1375] Terminal: If the user is undecided, suggest other products.

[1376] Example of a prompt

[1377] The following are examples of prompts to input into the generative AI model:

[1378] "Please describe the overview, functions, and specific examples of support systems for visually impaired individuals to shop independently."

[1379] "Please describe the specific processing steps and technologies used in a system that combines speech recognition and image recognition technologies."

[1380] The above describes specific embodiments for carrying out the present invention.

[1381] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1382] Step 1: Obtaining user voice input

[1383] The user voice-inputs "Take me to the nearest convenience store" into their smartphone. The device receives this voice input via its microphone and stores it as audio data.

[1384] Input data: Audio data

[1385] Output data: Stored audio data

[1386] Step 2: Convert audio to text

[1387] The device analyzes the accumulated audio data using speech recognition technology (for example, Google Cloud Speech-to-Text) and converts it into text data.

[1388] Input data: Audio data

[1389] Data processing: Analysis using speech recognition technology

[1390] Output data: Text data

[1391] Step 3: Send destination information

[1392] The converted text data is sent from the terminal to the server. The server receives this text data and processes it as destination information.

[1393] Input data: Text data

[1394] Data processing: Data transmission

[1395] Output data: Text data sent to the server

[1396] Step 4: Calculating square roots

[1397] The server parses the received text data and extracts destination information. Next, it uses map information (for example, Google Maps API) and GPS information to calculate the route from the current location to the destination.

[1398] Input data: Text data, map information, GPS information

[1399] Data processing: Extraction of destination information and route calculation.

[1400] Output data: Route information

[1401] Step 5: Return route information

[1402] The calculated route information is sent back from the server to the terminal. The terminal receives this route information.

[1403] Input data: Route information

[1404] Data processing: Data transmission

[1405] Output data: Route information sent to the terminal

[1406] Step 6: Generate audio guide

[1407] The device analyzes the received route information and generates a voice guide using speech synthesis technology (for example, Amazon Polly).

[1408] Input data: Route information

[1409] Data processing: Voice guide generation using speech synthesis technology

[1410] Output data: Audio guide

[1411] Step 7: Providing audio guides

[1412] The device provides the user with generated audio guidance. Specifically, it gives instructions via voice, such as, "Go 50 meters and turn right."

[1413] Input data: Audio guide

[1414] Data processing: Audio guide playback

[1415] Output data: Voice guidance

[1416] Step 8: Video Capture

[1417] The device uses its camera to capture video footage of the user's surroundings in real time. This video data is then sent to a server.

[1418] Input data: Camera video

[1419] Data processing: Video capture and transmission

[1420] Output data: Video data sent to the server

[1421] Step 9: Video Analysis

[1422] The server analyzes the received video data using deep learning technology (for example, TensorFlow). This allows it to recognize traffic lights and obstacles.

[1423] Input data: Video data

[1424] Data processing: Video analysis using deep learning technology

[1425] Output data: Recognition result

[1426] Step 10: Return the recognition results

[1427] The server sends the analysis results back to the terminal. Specifically, it includes information such as, "The signal is red. Please stop." The terminal receives this information.

[1428] Input data: Recognition result

[1429] Data processing: Data transmission

[1430] Output data: Recognition results sent to the terminal

[1431] Step 11: Notification of Recognition Results

[1432] The device notifies the user of the received recognition results via voice.

[1433] Input data: Recognition result

[1434] Data processing: Voice notification

[1435] Output data: Voice notification

[1436] Step 12: Recognizing your emotional state

[1437] The device captures the user's voice and facial expressions and analyzes their emotional state using emotion recognition technology (for example, IBM Watson Tone Analyzer).

[1438] Input data: Voice and facial expression data

[1439] Data processing: Analysis using emotion recognition technology

[1440] Output data: Emotional state

[1441] Step 13: Adjusting the voice guidance

[1442] The device adjusts the content and speed of voice guidance based on the recognized emotional state. For example, if the user is anxious, the guidance voice will slow down and more detailed explanations will be added.

[1443] Input data: Emotional state

[1444] Data processing: Adjustment of the content and speed of voice guidance.

[1445] Output data: Adjusted voice guidance

[1446] (Application Example 2)

[1447] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[1448] Conventional support systems for the visually impaired possessed basic functions such as voice guidance and image recognition, but lacked adaptive support that took into account the user's emotional state. As a result, they were unable to provide appropriate guidance in situations where the user felt anxious or unsure, or adequate support when the user was unsure about product selection. Furthermore, they were insufficient in supporting independent shopping in physical stores and in real-time adjustments during route guidance. Therefore, this invention proposes a system that solves these problems and provides optimal support according to the user's emotional state.

[1449] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for capturing the user's voice and facial expressions and recognizing their emotional state; means for adjusting the speed and detail of voice guidance based on the recognized emotional state; means for recognizing the user's emotional state and suggesting the most suitable product according to the emotion; and means for dynamically adjusting the guidance content based on the emotional state. As a result, the user can receive guidance and support optimized for their own emotional state.

[1450] "Users" refers to the users of the system, and specifically to individuals with visual impairments.

[1451] "Voice input" refers to the act of capturing the user's voice into the system via a microphone.

[1452] "Text conversion" is the process of converting voice input into text information.

[1453] "Destination" refers to the specific place the user wishes to travel to.

[1454] "Map information" refers to data that includes the route from the current location to the destination and surrounding geographical information.

[1455] "Route calculation" is the process of using map information to determine the optimal route from the current location to the destination.

[1456] "Voice guidance" refers to the act of conveying calculated route information and other instructions to users via voice.

[1457] "Video capture" is the act of capturing images of the surroundings using devices such as cameras.

[1458] "Video analysis" is the process of processing captured video data to identify specific information or objects.

[1459] An "obstacle" refers to a physical object that hinders the user's progress.

[1460] A "traffic light" is a light signaling device used to control road traffic.

[1461] "Recognition results" refer to information obtained through methods such as video analysis and audio analysis.

[1462] "Emotional state" refers to the user's current psychological and emotional condition.

[1463] "Emotion recognition means" refers to tools and technologies used to identify a user's emotional state by analyzing their voice and facial expressions.

[1464] "Product" refers to a specific item that a customer wishes to purchase within a physical store.

[1465] "Product recognition" is the process of identifying a specific product from camera footage or other images.

[1466] "Freshness" refers to an indicator related to the freshness and quality of a product.

[1467] "Quality evaluation" is the process of analyzing the freshness and condition of a product to determine its value.

[1468] "Voice notification" refers to the act of a system communicating analysis results, instructions, and other information to the user via voice.

[1469] "Guidance adjustment" is the process of appropriately changing the content and speed of guidance according to the user's emotional state and circumstances.

[1470] This invention is a support system for visually impaired individuals to independently navigate and perform daily shopping. By incorporating an emotion engine, the system provides optimal support based on the user's current emotional state. This system is realized by combining voice input, speech synthesis, image recognition, GPS, map information, AI technology, and emotion recognition technology.

[1471] System Configuration

[1472] Voice Input and Speech Synthesis: This system uses a terminal equipped with a microphone to accept voice input. Speech synthesis is performed using a speech synthesis library such as pyttsx3.

[1473] Image Recognition: A device equipped with a camera captures video of the surroundings and sends the data to a server. The server uses deep learning technology and image analysis tools such as OpenCV to analyze the video and recognize obstacles and the status of traffic lights.

[1474] GPS and Map Information: Use GPS modules or APIs (e.g., Google Maps API) to obtain current location information and calculate the optimal route to the destination.

[1475] Emotion Recognition: The `pipeline` function from the `transformers` library is used to recognize the user's emotional state from their voice and facial expressions.

[1476] Main Features

[1477] 1. In-store navigation

[1478] Users use their smartphone's voice input function to give commands such as, "Take me to the nearest milk stand."

[1479] The voice input is converted to text using the speech_recognition library and sent to the server.

[1480] The server uses map information based on the text to calculate the route. It then sends the calculation result back to the terminal and provides the user with voice guidance such as, "Go 20 meters and turn right."

[1481] 2. Product Recognition and Evaluation

[1482] When a user is looking for a specific product, they can use voice input, for example, by saying, "Tell me how fresh this milk is."

[1483] A camera is used to photograph the target product, and the video is sent to a server.

[1484] The server uses a deep learning model to recognize products and evaluate their freshness and quality. The evaluation results are sent back to the terminal, and the user is notified via voice message, "This milk is very fresh."

[1485] 3. Support tailored to emotional state

[1486] The system captures the user's voice and facial expressions, and recognizes their emotional state using the transformers library's pipeline.

[1487] Based on the recognized emotional state, the system automatically adjusts the speed and level of detail of voice guidance. For example, if the user is anxious, the guidance will be slow and detailed; if they are relaxed, the guidance will be at a normal speed.

[1488] Examples of specific cases and prompt statements

[1489] Specific example

[1490] User: "Take me to the nearest milk stand."

[1491] Terminal message: "Go 20 meters and turn right."

[1492] User: While shopping, asked, "Please tell me how fresh this milk is."

[1493] Device: Captures camera footage, analyzes it on the server, and then sends a notification saying, "This milk is very fresh."

[1494] Example of a prompt

[1495] "Take me to the nearest milk stand."

[1496] "Please tell me how fresh this milk is."

[1497] "Please turn right next."

[1498] "There is an obstacle in the direction you are going."

[1499] "This product is extremely fresh."

[1500] This invention enables users to navigate and shop efficiently and safely in physical stores by receiving optimized guidance and support tailored to their emotional state. Furthermore, emotion recognition technology can reduce the psychological burden on users.

[1501] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1502] Step 1:

[1503] Users use their smartphone's voice input function to enter commands regarding destinations or products. This voice input is captured by the smartphone's built-in microphone. The input data consists of the user's voice, specifically commands such as "Take me to the nearby milk section" or "Tell me the freshness of this milk."

[1504] Step 2:

[1505] The voice input data is converted into text data using the terminal's speech_recognition library. At this stage, the input is the user's voice data, and the output is text data obtained by analyzing the voice. The text data will be in the format of "Take me to the nearest milk stand" or "Tell me how fresh this milk is."

[1506] Step 3:

[1507] The data, converted to text, is sent to the server, which analyzes the data to identify destination and product information. The input here is a text-converted instruction, which the server then analyzes. The output includes destination information (milk section) and product information (milk).

[1508] Step 4:

[1509] Once a destination is identified, the server uses map information to calculate the optimal route from the current location to the destination. The input data consists of the current location and destination information, and the server executes a route calculation algorithm based on this data. The output data generates detailed route information, such as "Go 20 meters and turn right."

[1510] Step 5:

[1511] The calculated route information is sent back to the terminal, which then uses a speech synthesis library (such as pyttsx3) to provide voice guidance. The input data is route information obtained from the server, which the terminal converts into voice data. The output data is the voice guidance that the user can hear. Specifically, it generates voice guidance such as, "Go 20 meters and turn right."

[1512] Step 6:

[1513] When a user selects a product, the camera is used to photograph the product, and the video data is sent to the server. The input data for this step is the video data captured by the camera. The video data is an image containing the product, specifically including images of milk, for example.

[1514] Step 7:

[1515] The server analyzes the received video data using deep learning techniques and image recognition tools (such as OpenCV) to identify products. The input data is video data transmitted from the camera, and the output data is information about the analyzed products. Specifically, it recognizes milk in the video and evaluates its quality and freshness.

[1516] Step 8:

[1517] Product information analyzed by the server undergoes a quality evaluation, and the evaluation results are sent back to the terminal. The input data is the product recognition result, and the server performs the evaluation using a quality evaluation algorithm. The output data is the freshness evaluation result, and information such as "This milk is very fresh" can be obtained.

[1518] Step 9:

[1519] The terminal uses speech synthesis technology (such as pyttsx3) to notify the user of the received evaluation results via voice. The input data is the freshness evaluation results returned from the server, which the terminal converts into voice data. The output data is a voice notification of the quality evaluation results that the user can hear.

[1520] Step 10:

[1521] The user's voice and facial expressions are captured and acquired by the device as data to recognize their emotional state. The input data consists of the user's voice and facial expressions, which the device captures.

[1522] Step 11:

[1523] The acquired emotion data is analyzed using emotion recognition technology (the pipeline function in the transformers library). The input data consists of voice and facial capture data, and the output data is the recognized emotional state. Specifically, emotional states such as "anxious" or "relaxed" can be obtained.

[1524] Step 12:

[1525] Based on the recognized emotional state, the speed and detail of the voice guidance are adjusted. The input data is the result of emotion recognition, and the terminal changes the voice guidance settings based on this. The output data is the adjusted voice guidance; for example, if the user is "anxious," slow and detailed guidance will be provided.

[1526] Through this series of steps, users can independently navigate and shop in physical stores, and receive optimal support tailored to their emotional state.

[1527] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1528] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1529] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[1530] [Third Embodiment]

[1531] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[1532] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1533] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1534] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[1535] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1536] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1537] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1538] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1539] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1540] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1541] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1542] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[1543] System Overview

[1544] This invention is an assistance system for visually impaired individuals to enable them to move around and perform daily shopping independently. This system combines voice input, speech synthesis, image recognition, GPS, map information, and AI technology. The embodiments are described below.

[1545] Navigation function

[1546] Enter destination

[1547] The user voice-inputs "Take me to the nearest convenience store" into their smartphone. The device uses speech recognition technology to convert the voice into text. This text information is sent to a server and processed as destination information.

[1548] Root calculation

[1549] The server calculates a route using map information based on the received text information (destination). Specifically, it uses an API to calculate the optimal walking route from the current location to the destination. The calculated route information is then returned to the device.

[1550] Voice directions

[1551] The terminal analyzes the calculated route information and provides voice instructions to the user regarding the next direction and distance to travel. For example, it might say, "Go 50 meters and turn right," providing step-by-step guidance.

[1552] Real-time video analysis

[1553] The device uses a camera to capture video of the user's surroundings in real time. This video data is sent to a server, which analyzes the video using deep learning technology. Specifically, it recognizes the status of traffic lights and obstacles. The results of this analysis are returned to the device, and the user is notified of the situation via voice, such as, "The traffic light is red. Please stop."

[1554] Shopping assistance

[1555] Product Selection

[1556] The user voice-inputs, "I want to buy tomatoes." The device then uses speech recognition technology to convert this voice into text and sends it to the server.

[1557] Product recognition

[1558] The server uses a model to recognize products from the terminal's camera footage based on the received text information (product information). When the user points the camera at a tomato shelf, the footage is sent to the server, and image analysis is performed using deep learning technology.

[1559] Freshness evaluation

[1560] The freshness and quality of tomatoes are evaluated based on image analysis. For example, the evaluation is based on color, shape, and gloss. The evaluation results are returned to the device, and the user is notified by voice, "This tomato is very fresh."

[1561] Specific example

[1562] Examples of directions

[1563] User: Inputs "Take me to the nearest convenience store" by voice.

[1564] Terminal: Converts speech to text and sends it to the server.

[1565] Server: Uses the Google Maps API to calculate the route from the current location to the destination and sends it back to the device.

[1566] Terminal: Provides voice guidance to the user regarding the calculated route information (e.g., "Go 50 meters and turn right").

[1567] Terminal: Captures camera footage and sends it to a server to recognize the status of traffic lights.

[1568] Server: Recognizes that the traffic light is red and notifies the terminal with the message, "The traffic light is red. Please stop."

[1569] Terminal: Notifies the user via voice.

[1570] Specific examples of shopping

[1571] User: Inputs "I want to buy tomatoes" by voice.

[1572] Terminal: Converts speech to text and sends it to the server.

[1573] Server: Uses a model for product recognition.

[1574] User: Point the camera at the tomato trellis.

[1575] Terminal: Captures camera footage and sends it to the server.

[1576] Server: Evaluates the freshness and quality of tomatoes and sends the results back to the terminal.

[1577] Device: Notifies the user of the evaluation results via voice (e.g., "This tomato is very fresh").

[1578] In this way, the system of the present invention can effectively provide support to visually impaired persons when they are able to move around and shop independently.

[1579] The following describes the processing flow.

[1580] Processing steps for the navigation function

[1581] Step 1:

[1582] The user uses voice input to tell the terminal, "Take me to the nearest convenience store."

[1583] Step 2:

[1584] The device captures voice input and uses speech recognition technology to convert the voice signal into text.

[1585] Step 3:

[1586] The terminal sends the converted text data to the server.

[1587] Step 4:

[1588] The server extracts destination information from the received text data and uses map information to calculate the route from the current location to the destination.

[1589] Step 5:

[1590] The server returns the calculated route information to the terminal.

[1591] Step 6:

[1592] The terminal analyzes route information and uses speech synthesis technology to instruct the user verbally on the next direction and distance to go. For example, it might say, "Go 50 meters and turn right."

[1593] Step 7:

[1594] The device uses its camera to capture images of its surroundings in real time.

[1595] Step 8:

[1596] The terminal sends the captured video to the server.

[1597] Step 9:

[1598] The server receives the video data and uses a deep learning model to analyze obstacles and traffic light conditions within the video.

[1599] Step 10:

[1600] The server sends the analysis results back to the terminal. For example, if the traffic light is red, it will send back an instruction such as, "The traffic light is red. Please stop."

[1601] Step 11:

[1602] The device notifies the user of the analysis results via voice. The notification will be in the form of, "The traffic light is red. Please stop."

[1603] Processing steps for the shopping support function

[1604] Step 1:

[1605] The user uses voice input to say "I want to buy tomatoes" to the device.

[1606] Step 2:

[1607] The device captures voice input and uses speech recognition technology to convert the voice signal into text.

[1608] Step 3:

[1609] The terminal sends the converted text data to the server.

[1610] Step 4:

[1611] The server extracts product information from the received text data and selects an appropriate image recognition model.

[1612] Step 5:

[1613] The user points the camera at the tomato trellis.

[1614] Step 6:

[1615] The device uses a camera to capture real-time video of the product shelves.

[1616] Step 7:

[1617] The terminal sends the captured video to the server.

[1618] Step 8:

[1619] The server receives the video data and uses an image recognition model to identify tomatoes in the video.

[1620] Step 9:

[1621] The server evaluates the freshness and quality of the identified tomatoes.

[1622] Step 10:

[1623] The server sends the evaluation results back to the terminal. For example, it might return an evaluation result such as, "This tomato is very fresh."

[1624] Step 11:

[1625] The device notifies the user of the evaluation results via voice. The notification may be in the form of, "This tomato is very fresh."

[1626] These detailed processing steps enable the system of the present invention to allow visually impaired individuals to independently receive support in various aspects of their daily lives.

[1627] (Example 1)

[1628] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1629] There is a lack of technical support for visually impaired individuals to move independently, safely, and comfortably, and to perform daily shopping. To address this challenge, a system is needed that recognizes destinations and products through voice input and provides appropriate instructions through voice guidance and video analysis. Furthermore, real-time updated navigation information and product freshness assessments based on this system are also required.

[1630] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1631] In this invention, the server includes means for receiving voice input from a user, means for converting the received voice input into text, means for identifying a destination based on the converted text and calculating a route using map information, means for communicating the calculated route information to the user by voice, means for capturing and analyzing video of the user's surroundings, means for recognizing obstacles and traffic light conditions in the captured video, and means for notifying the user of the recognition results by voice. This enables visually impaired individuals to recognize destinations and products through voice input and to move and shop safely while receiving appropriate guidance and instructions in real time.

[1632] "Means for receiving user voice input" refers to a device or program that allows visually impaired individuals to input voice commands using a smartphone or other device.

[1633] "Means for converting received voice input into text" refers to a device or program that uses speech recognition technology to convert voice data into text data.

[1634] "Means for identifying a destination based on converted text and calculating a route using map information" refers to a device or program that analyzes text data, refers to a map database based on a specified destination, and calculates the optimal route.

[1635] "Means of conveying calculated route information to users via voice" refers to a device or program that provides calculated route information as a voice message using speech synthesis technology.

[1636] "Means for capturing and analyzing video footage of the user's surroundings" refers to a device or program that uses a camera to capture images of the user's surroundings in real time and analyzes that image data.

[1637] "Means for recognizing obstacles and traffic light status in captured video" refers to a device or program that uses deep learning technology to recognize information about obstacles and traffic lights from acquired video data.

[1638] "Means for notifying the user of the recognition results by voice" refers to a device or program that provides the recognized information as a voice message using speech synthesis technology.

[1639] "A means of inputting a product name by voice, analyzing that information, and converting it into text" refers to a device or program that allows a user to input a product name by voice and convert that information into text data.

[1640] "Means for recognizing products from camera footage" refers to a device or program that uses deep learning technology to identify specific products from video data acquired using a camera.

[1641] "Means for evaluating the freshness and quality of recognized goods" refers to a device or program that evaluates the freshness and quality of recognized goods based on criteria such as color, shape, and gloss.

[1642] "Means for notifying users of evaluation results by voice" refers to a device or program that provides evaluation results of freshness and quality as a voice message using speech synthesis technology.

[1643] "Means for updating routes in real time" refers to a device or program that dynamically updates the navigation route for each action based on the user's current location information.

[1644] "Means of providing voice instructions for the direction of travel or the next action" refers to a device or program that uses speech synthesis technology to provide the user with voice messages indicating the direction and action they should take next.

[1645] "Means for recognizing moving objects such as pedestrians and animals in the surrounding area and notifying the user of that information via voice" refers to a device or program that recognizes moving objects from camera footage and provides information for avoiding danger as a voice message.

[1646] System Overview

[1647] This invention is an assistance system for visually impaired individuals to enable them to move around and perform daily shopping independently. This system is comprised of a combination of voice input, speech synthesis, image recognition, GPS, map information, and AI technology.

[1648] Navigation function

[1649] Enter destination

[1650] The user inputs a voice command into their smartphone saying, "Take me to the nearest convenience store."

[1651] The device uses speech recognition technology to convert this speech into text. Specifically, it uses Google Cloud Speech-to-Text.

[1652] The device sends the generated text data ("Please guide me to the nearest convenience store") to the server.

[1653] The server receives the text data and uses it to determine the destination.

[1654] Root calculation

[1655] The server determines the user's current location based on the destination information received.

[1656] The server uses map information and the Google Maps API to calculate the optimal walking route from the current location to the destination.

[1657] The server generates the calculated route information in JSON format and sends it back to the terminal.

[1658] Voice directions

[1659] The terminal parses the received route information in JSON format.

[1660] To guide the user verbally in the next direction and distance, a voice message is generated using Google Cloud Text-to-Speech.

[1661] For example, based on route information, it might instruct you to "Go 50 meters and turn right."

[1662] Real-time video analysis

[1663] The device activates its camera and captures video of the user's surroundings in real time.

[1664] The terminal sends the captured video data to the server.

[1665] The server analyzes the received video using deep learning technology (TensorFlow). It recognizes traffic lights and obstacles, generates that information in text format, and sends it back to the terminal.

[1666] The device generates a voice message using speech synthesis technology (Google Cloud Text-to-Speech) based on the received recognition results and notifies the user.

[1667] For example, it might announce, "The traffic light is red. Please stop."

[1668] Shopping assistance

[1669] Product Selection

[1670] The user voice-inputs "I want to buy tomatoes" into their smartphone.

[1671] The device uses speech recognition technology (Google Cloud Speech-to-Text) to convert this speech into text.

[1672] The terminal sends the generated text data ("I want to buy tomatoes") to the server.

[1673] The server receives the text data and processes it as product information.

[1674] Product recognition

[1675] The server uses a deep learning model to recognize tomatoes based on product information.

[1676] The user points the camera at the tomato trellis.

[1677] The device captures camera footage and sends it to the server.

[1678] The server analyzes the received video using deep learning technology (TensorFlow) to recognize tomatoes.

[1679] Freshness evaluation

[1680] The server evaluates the freshness and quality of the recognized tomatoes. Specifically, it evaluates them based on criteria such as color, shape, and gloss.

[1681] The server generates the evaluation results in text format and sends them back to the terminal.

[1682] The device generates a voice message saying, "These tomatoes are very fresh," and notifies the user.

[1683] Specific example

[1684] Examples of directions

[1685] User: Inputs "Take me to the nearest convenience store" by voice.

[1686] Terminal: Use Google Cloud Speech-to-Text to convert speech to text and send it to the server.

[1687] Server: Uses the Google Maps API to calculate the route from the current location to the destination and returns it to the device in JSON format.

[1688] Terminal: Analyzes the received route information and uses Google Cloud Text-to-Speech to provide voice guidance such as, "Go 50 meters and turn right."

[1689] Terminal: Captures camera footage and sends it to a server to recognize the status of traffic lights.

[1690] Server: Uses deep learning technology (TensorFlow) to recognize that the signal is red and sends a message back to the terminal.

[1691] Terminal: Based on the received information, it notifies the user by voice, "The traffic light is red. Please stop."

[1692] Specific examples of shopping

[1693] User: Inputs "I want to buy tomatoes" by voice.

[1694] Terminal: Use Google Cloud Speech-to-Text to convert speech to text and send it to the server.

[1695] Server: Prepares to recognize tomatoes using deep learning technology.

[1696] User: Point the camera at the tomato trellis.

[1697] Terminal: Captures camera footage and sends it to the server.

[1698] Server: Uses deep learning technology (TensorFlow) to analyze received video and evaluate the freshness and quality of tomatoes.

[1699] Server: Generates evaluation results in text format and sends them back to the terminal.

[1700] Device: Use Google Cloud Text-to-Speech to announce, "These tomatoes are very fresh."

[1701] In this way, the system of the present invention can effectively provide support to visually impaired persons when they are able to move around and shop independently.

[1702] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1703] Specific processing flow of the navigation function

[1704] Enter destination

[1705] Step 1:

[1706] The user uses voice input on their smartphone, saying, "Take me to the nearest convenience store."

[1707] Input: Voice command ("Take me to the nearest convenience store.")

[1708] Output: Audio data

[1709] Step 2:

[1710] The device uses speech recognition technology (Google Cloud Speech-to-Text) to convert the input voice data into text data.

[1711] Specifically, audio data is sent to the server, and the server returns text data.

[1712] Input: Audio data

[1713] Output: Text data ("Please guide me to the nearest convenience store")

[1714] Step 3:

[1715] The terminal sends the generated text data to the server. The server receives this data and processes it as destination information.

[1716] Input: Text data ("Please guide me to the nearest convenience store")

[1717] Output: Destination information

[1718] Root calculation

[1719] Step 4:

[1720] The server determines the user's current location based on the destination information received.

[1721] Input: Destination information

[1722] Output: Current location information

[1723] Step 5:

[1724] The server uses the Google Maps API to calculate the optimal walking route from the current location to the destination.

[1725] Input: Current location information, destination information

[1726] Output: Route information (JSON format)

[1727] Step 6:

[1728] The server generates the calculated route information in JSON format and sends it back to the terminal.

[1729] Input: Route information (in generated JSON format)

[1730] Output: Route information (JSON format)

[1731] Voice directions

[1732] Step 7:

[1733] The terminal analyzes the received route information and calculates the next direction and distance to proceed.

[1734] Input: Route information (JSON format)

[1735] Output: Guidance / Instructions

[1736] Step 8:

[1737] The device uses Google Cloud Text-to-Speech to generate guidance instructions as voice messages and notify the user.

[1738] For example, you might give directions like, "Go 50 meters and turn right."

[1739] Input: Guidance / Instructions

[1740] Output: Voice message

[1741] Real-time video analysis

[1742] Step 9:

[1743] The device activates its camera and captures video of the user's surroundings in real time.

[1744] Input: Camera video

[1745] Output: Video data

[1746] Step 10:

[1747] The terminal sends the captured video data to the server.

[1748] Input: Video data

[1749] Output: Video data sent to the server

[1750] Step 11:

[1751] The server uses deep learning technology (TensorFlow) to analyze the received video and recognize the status of traffic lights and obstacles.

[1752] Input: Video data

[1753] Output: Analysis results (status of traffic lights and obstacles)

[1754] Step 12:

[1755] The server generates the analysis results in text format and sends them back to the terminal.

[1756] Input: Analysis results

[1757] Output: Text data (analysis results)

[1758] Step 13:

[1759] The device generates a voice message using Google Cloud Text-to-Speech based on the received recognition results and notifies the user.

[1760] For example, you might announce, "The traffic light is red. Please stop."

[1761] Input: Text data (analysis results)

[1762] Output: Voice message

[1763] Specific processing flow for shopping assistance

[1764] Product Selection

[1765] Step 1:

[1766] The user voice-inputs "I want to buy tomatoes" into their smartphone.

[1767] Input: Voice command ("I want to buy tomatoes")

[1768] Output: Audio data

[1769] Step 2:

[1770] The device uses speech recognition technology (Google Cloud Speech-to-Text) to convert this speech into text.

[1771] Specifically, audio data is sent to the server, and the server returns text data.

[1772] Input: Audio data

[1773] Output: Text data ("I want to buy tomatoes")

[1774] Step 3:

[1775] The terminal sends the generated text data to the server. The server receives this data and processes it as product information.

[1776] Input: Text data ("I want to buy tomatoes")

[1777] Output: Product Information

[1778] Product recognition

[1779] Step 4:

[1780] The server prepares a product recognition model based on the product information.

[1781] Input: Product Information

[1782] Output: Recognition Model

[1783] Step 5:

[1784] The user points the camera at the tomato trellis.

[1785] Input: Image of a product shelf

[1786] Output: Camera data

[1787] Step 6:

[1788] The device captures camera footage and sends it to the server.

[1789] Input: Camera data

[1790] Output: Video data sent to the server

[1791] Step 7:

[1792] The server uses deep learning technology (TensorFlow) to analyze the received video and recognize the product.

[1793] Input: Video data

[1794] Output: Analysis results (product recognition)

[1795] Step 8:

[1796] The server generates the recognition results in text format and sends them back to the terminal.

[1797] Input: Analysis results

[1798] Output: Text data (product recognition)

[1799] Freshness evaluation

[1800] Step 9:

[1801] The server evaluates the freshness and quality of the recognized products. Specifically, it evaluates them based on criteria such as color, shape, and gloss.

[1802] Input: Product recognition result

[1803] Output: Freshness evaluation results

[1804] Step 10:

[1805] The server generates the evaluation results in text format and sends them back to the terminal.

[1806] Input: Freshness evaluation results

[1807] Output: Text data (evaluation results)

[1808] Step 11:

[1809] The device uses Google Cloud Text-to-Speech to generate the evaluation results as an audio message and notify the user.

[1810] For example, you might say, "These tomatoes are extremely fresh."

[1811] Input: Text data (evaluation results)

[1812] Output: Voice message

[1813] (Application Example 1)

[1814] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1815] Visually impaired individuals face numerous challenges when attempting to travel and shop independently. In particular, accurately perceiving their surroundings is difficult when seeking directions to a destination or purchasing goods. They struggle to recognize traffic signals, pedestrians, animals, and other moving objects. Furthermore, locating products and assessing their freshness and quality is extremely challenging. Effective support systems are needed to address these issues.

[1816] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1817] In this invention, the server includes: means for receiving voice input from the user; means for converting the received voice input into text; means for identifying a destination based on the converted text and calculating a route using map information; means for communicating the calculated route information to the user by voice; means for capturing and analyzing video of the user's surroundings; means for recognizing obstacles and traffic light conditions in the captured video; means for notifying the user of the recognition results by voice; means for the user to input a product name by voice, analyzing that information and converting it into text; means for recognizing a product from the camera image based on the converted text; and means for providing information about the recognized product. The system includes means for evaluating degree and quality; means for notifying the user of the evaluation results by voice; means for identifying the user's current location and updating the route in real time using map information; means for instructing the user on the direction of travel and the next action by voice based on the updated route information; means for recognizing moving objects such as pedestrians and animals in the surroundings other than obstacles and traffic lights, and notifying the user of that information by voice; means for the user to input searched products by voice, and for analyzing that information and converting it into text; means for calculating the evaluation results of product candidates based on the converted text using deep learning technology; and means for guiding the user to the desired product in real time based on the evaluation results. This enables visually impaired people to move and shop independently, safely and effectively.

[1818] A "system" refers to the entire apparatus, including a series of means, designed to support visually impaired individuals in independently navigating and shopping.

[1819] "Users" refers to visually impaired individuals who use this system to receive guidance and shopping assistance.

[1820] "Voice input" refers to the voice signals that a user speaks to the system.

[1821] "Means for receiving voice input" refers to devices or software that acquire the user's voice and allow the system to interpret it.

[1822] "Means of converting speech to text" refers to devices or software that analyze speech input and convert it into textual information.

[1823] "Means of identifying a destination based on text" refers to devices or software that determine a destination based on converted character information.

[1824] "Map information" refers to a database containing geographical information, providing destinations, current location, route information, and more.

[1825] "Means of calculating routes" refers to devices or software that use map information to calculate the optimal route to a destination.

[1826] "Means of conveying route information by voice" refers to devices or software that guide users through calculated routes by voice.

[1827] "Means of capturing video of the user's surroundings" refers to devices that acquire images or videos of the user's surroundings using cameras or similar devices.

[1828] "Means of analyzing video" refers to devices or software that analyze captured video to extract useful information.

[1829] "Means for recognizing obstacles and traffic signal status" refers to devices and software that detect the status of surrounding obstacles and traffic signals through video analysis.

[1830] "Means of notifying the user of recognition results by voice" refers to devices or software that inform the user of the analysis results by voice.

[1831] "Methods for inputting product names by voice" refer to devices or software that allow users to tell the system by voice the products they wish to purchase.

[1832] "Means of recognizing products from camera footage" refers to devices or software that use a camera to acquire images of products and then recognize those products.

[1833] "Means for evaluating freshness and quality" refers to devices or software used to evaluate the condition of recognized products.

[1834] "Means for determining current location" refers to devices or software used to determine the user's current geographical location.

[1835] "Means of updating routes in real time" refers to devices or software that continuously update route guidance in accordance with the user's movements.

[1836] "Means of providing voice instructions for direction of travel or next action" refers to devices or software that provide voice guidance for the next action based on updated route information.

[1837] "Means of recognizing moving objects such as pedestrians and animals" refers to devices and software used to detect moving objects in the surrounding environment.

[1838] "Methods for inputting searched products by voice" refers to devices or software that allow users to tell the system by voice what product they are looking for.

[1839] "Methods for calculating evaluation results of candidate products using deep learning technology" refers to devices or software that analyze and evaluate the characteristics of candidate products using deep learning.

[1840] "A means of guiding users to desired products in real time based on evaluation results" refers to devices or software that guide users to products they are looking for based on evaluation results.

[1841] System Configuration

[1842] To implement this invention, a server, terminals, and users are required. The terminals consist of smartphones or devices with built-in cameras. The server is responsible for data processing and analysis and operates in a cloud computing environment. The system integrates multiple functions, including voice input recognition, text conversion, route guidance using map information, real-time video analysis, product recognition, and quality evaluation.

[1843] Program and Processing Overview

[1844] Voice input acceptance and text conversion:

[1845] The user provides voice input. For example, they might say, "Take me to the nearest convenience store." The device receives this input and uses speech recognition technology to convert it into text. The Google Speech-to-Text API is used for this process.

[1846] Destination identification and route calculation:

[1847] The converted audio data is sent to a server, which identifies the destination based on the converted text. It then uses map information to calculate the optimal route. The Google Maps API is used here. The calculated route information is returned to the device.

[1848] Voice-guided route guidance:

[1849] The terminal analyzes route information calculated using speech synthesis technology and provides voice instructions to the user regarding the next direction and distance to travel. For example, it might say, "Go 50 meters and turn right."

[1850] Real-time video capture and analysis:

[1851] The device's camera captures video of the user's surroundings in real time and sends the data to a server. The server uses deep learning technology to analyze the video data and recognize traffic lights and obstacles. The results of this analysis are returned to the device and notified to the user via voice.

[1852] Product recognition and quality assessment:

[1853] The user voice-inputs "I want to buy tomatoes." This voice is converted back into text, and the server uses a model that recognizes products from camera footage based on the text information. When the user points the camera at a shelf of tomatoes, the image is sent to the server. Deep learning technology is used to evaluate the freshness and quality of the tomatoes, and the results are returned to the terminal. For example, the user might receive a voice notification saying, "These tomatoes are very fresh."

[1854] Current location and real-time route updates:

[1855] The server uses GPS to determine the user's current location and updates the route in real time based on map information. Based on the updated route information, the terminal provides voice instructions to the user regarding the direction to go and the next action.

[1856] Recognition and notification of moving objects:

[1857] The server uses video analysis to recognize moving objects in the surrounding area, such as pedestrians and animals, other than obstacles and traffic lights, and notifies the user of this information via voice.

[1858] Specific example

[1859] 1. Specific Examples of Navigation

[1860] User: "Take me to the nearest convenience store."

[1861] Terminal: Converts speech to text and sends it to the server.

[1862] Server: Route calculation using Google Maps API

[1863] Device: Voice-guided route instructions (e.g., "Go 50 meters and turn right")

[1864] Terminal: Captures camera footage and sends the traffic light status to the server.

[1865] Server: Recognizes the status of the traffic light and notifies, "The traffic light is red. Please stop."

[1866] Device: Notifies the user via voice.

[1867] 2. Specific examples of shopping

[1868] User: "I want to buy tomatoes."

[1869] Terminal: Converts speech to text and sends it to the server.

[1870] Server: Uses a model for product recognition.

[1871] User: Pointing the camera at the tomato trellis.

[1872] Terminal: Captures camera footage and sends it to the server.

[1873] Server: Evaluates the freshness and quality of the tomatoes and notifies, "These tomatoes are very fresh."

[1874] Device: Notifies the user via voice.

[1875] This system assists users in moving around and shopping safely and conveniently. Its purpose is to provide an environment where visually impaired individuals can live independently, utilizing technologies such as voice recognition, deep learning, and GPS.

[1876] Example of a prompt

[1877] 1. Navigation prompt text:

[1878] Please guide me to the nearest convenience store.

[1879] 2. Product purchase prompt message:

[1880] I bought tomatoes.

[1881] As described above, this invention provides comprehensive support for visually impaired individuals to move around and shop independently.

[1882] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1883] Flow of processing steps

[1884] Navigation function

[1885] Step 1:

[1886] Input: The user inputs "Take me to the nearest convenience store" by voice.

[1887] Operation: The device (smartphone) accepts voice input.

[1888] Output: Audio data.

[1889] Step 2:

[1890] Input: Audio data.

[1891] Operation: The device uses speech recognition technology to convert speech data into text. The Google Speech-to-Text API is used here.

[1892] Output: Text data.

[1893] Step 3:

[1894] Input: Text data.

[1895] Operation: The server identifies the destination based on text data and calculates the route using the Google Maps API.

[1896] Output: Route information.

[1897] Step 4:

[1898] Input: Route information.

[1899] Operation: The server sends route information back to the terminal, and the terminal parses this information.

[1900] Output: Route guidance information after analysis.

[1901] Step 5:

[1902] Input: Route guidance information after analysis.

[1903] Operation: The device uses speech synthesis technology to provide instructions to the user. For example, it might say, "Go 50 meters and turn right."

[1904] Output: Voice guidance.

[1905] Step 6:

[1906] Input: Video of the user's surroundings.

[1907] Operation: The device's camera captures video and sends it to the server.

[1908] Output: Video data.

[1909] Step 7:

[1910] Input: Video data.

[1911] Operation: The server uses deep learning technology to analyze video and recognize traffic lights and obstacles.

[1912] Output: Recognition result.

[1913] Step 8:

[1914] Input: Recognition result.

[1915] Operation: The server sends the recognition result back to the terminal, and the terminal notifies the user by voice, "The traffic light is red. Please stop."

[1916] Output: Voice notification.

[1917] Shopping support function

[1918] Step 1:

[1919] Input: The user enters "I want to buy tomatoes" by voice.

[1920] Operation: The device (smartphone) accepts voice input.

[1921] Output: Audio data.

[1922] Step 2:

[1923] Input: Audio data.

[1924] Operation: The device uses speech recognition technology to convert speech data into text. The Google Speech-to-Text API is used here.

[1925] Output: Text data.

[1926] Step 3:

[1927] Input: Text data.

[1928] Operation: The server prepares a product recognition model based on text data.

[1929] Output: Recognition model.

[1930] Step 4:

[1931] Input: Camera footage.

[1932] Operation: The user points their smartphone camera at a product shelf (e.g., a tomato shelf). This video data is sent from the device to the server.

[1933] Output: Video data.

[1934] Step 5:

[1935] Input: Video data.

[1936] Operation: The server uses deep learning technology to recognize products from images and evaluate the freshness and quality of tomatoes.

[1937] Output: Evaluation results.

[1938] Step 6:

[1939] Input: Evaluation results.

[1940] Operation: The server sends the evaluation result back to the terminal, and the terminal notifies the user by voice, "This tomato is very fresh."

[1941] Output: Voice notification.

[1942] Real-time route update function

[1943] Step 1:

[1944] Input: Current location.

[1945] Operation: The device's GPS function determines the user's current location and sends it to the server.

[1946] Output: Current location information.

[1947] Step 2:

[1948] Input: Current location information.

[1949] Operation: The server updates the route in real time using map information.

[1950] Output: Updated route information.

[1951] Step 3:

[1952] Input: Updated route information.

[1953] Operation: The device provides voice instructions to the user regarding the direction of travel and the next action based on updated route information.

[1954] Output: Voice command.

[1955] Recognition function for moving objects

[1956] Step 1:

[1957] Input: Video of the user's surroundings.

[1958] Operation: The device's camera captures the surrounding video and sends it to the server.

[1959] Output: Video data.

[1960] Step 2:

[1961] Input: Video data.

[1962] Operation: The server uses deep learning technology to analyze video and recognize moving objects such as pedestrians and animals.

[1963] Output: Recognition result.

[1964] Step 3:

[1965] Input: Recognition result.

[1966] Operation: The server sends the recognition result back to the terminal, and the terminal notifies the user by voice, "There is a pedestrian ahead. Please be careful."

[1967] Output: Voice notification.

[1968] The above outlines the specific processing steps of this system. This will enable visually impaired individuals to move around and shop independently.

[1969] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1970] System Overview

[1971] This invention is an assistance system for visually impaired individuals to independently navigate and perform daily shopping. By incorporating an emotion engine, it provides optimal support based on the user's current emotional state. The system combines voice input, speech synthesis, image recognition, GPS, map information, AI technology, and emotion recognition technology.

[1972] Navigation function

[1973] Enter destination

[1974] The user speaks into their smartphone and says, "Take me to the nearest convenience store." The device uses speech recognition technology to convert the speech into text. This text information is sent to a server and processed as destination information.

[1975] Root calculation

[1976] The server extracts destination information from the received text information and calculates the route from the current location to the destination using map information. It then sends the calculated route information back to the terminal.

[1977] Voice directions

[1978] The terminal analyzes the calculated route information and uses speech synthesis technology to instruct the user on the next direction and distance to go. For example, it provides step-by-step instructions such as, "Go 50 meters and turn right."

[1979] Real-time video analysis

[1980] The device uses a camera to capture video of the user's surroundings in real time. This video data is sent to a server, which analyzes the video using deep learning technology. Specifically, it recognizes the status of traffic lights and obstacles. The results of this analysis are sent back to the device, and the user is notified of the situation via voice, such as, "The traffic light is red. Please stop."

[1981] Emotion recognition and regulation

[1982] The device captures the user's voice and facial expressions and uses an emotion engine to recognize their current emotional state. For example, if the user is anxious, the guidance voice will slow down and provide detailed instructions to create a sense of security. Conversely, if the user is relaxed, the guidance will be given at a normal speed.

[1983] Shopping assistance

[1984] Product Selection

[1985] The user voice-inputs, "I want to buy tomatoes." The device uses speech recognition technology to convert the voice into text and sends it to the server.

[1986] Product recognition

[1987] The server uses a model to recognize products from the terminal's camera footage based on the received text information. When the user points the camera at a tomato shelf, the footage is sent to the server, and image analysis is performed using deep learning technology.

[1988] Freshness evaluation

[1989] The freshness and quality of tomatoes are evaluated based on image analysis. For example, the evaluation is based on color, shape, and gloss. The evaluation results are sent back to the device, and the user is notified by voice, "This tomato is very fresh."

[1990] Emotion recognition and product recommendations

[1991] The device captures the user's voice and facial expressions and uses an emotion engine to recognize their current emotional state. For example, if the user is undecided, it suggests other items on the same shelf to broaden their options. On the other hand, if the user is in a hurry, it quickly suggests the highest-rated product.

[1992] Specific example

[1993] Examples of directions

[1994] User: Inputs "Take me to the nearest convenience store" by voice.

[1995] Terminal: Converts speech to text and sends it to the server.

[1996] Server: Calculates the route from the current location to the destination and sends it back to the terminal.

[1997] Terminal: Provides route information to the user via voice (e.g., "Go 50 meters and turn right").

[1998] Terminal: Captures camera footage and sends it to the server to analyze the status of the traffic lights.

[1999] Server: Recognizes that the traffic light is red and notifies, "The traffic light is red. Please stop."

[2000] Terminal: Notifies the user via voice.

[2001] Terminal: If the user is anxious, slow down the voice guidance and add more detailed explanations.

[2002] Specific examples of shopping

[2003] User: Inputs "I want to buy tomatoes" by voice.

[2004] Terminal: Converts speech to text and sends it to the server.

[2005] Server: Uses a model for product recognition.

[2006] User: Point the camera at the tomato trellis.

[2007] Terminal: Captures camera footage and sends it to the server.

[2008] Server: Evaluates the freshness and quality of tomatoes and sends the results back to the terminal.

[2009] Device: Notifies the user of the evaluation results via voice (e.g., "This tomato is very fresh").

[2010] Terminal: If the user is undecided, suggest other products.

[2011] In this way, the system of the present invention can provide optimal support tailored to the emotional state of visually impaired individuals when they are independently moving around or shopping.

[2012] The following describes the processing flow.

[2013] Processing steps for the navigation function

[2014] Step 1:

[2015] The user uses voice input to tell the terminal, "Take me to the nearest convenience store."

[2016] Step 2:

[2017] The device captures voice input and uses speech recognition technology to convert the voice signal into text.

[2018] Step 3:

[2019] The terminal sends the converted text data to the server.

[2020] Step 4:

[2021] The server extracts destination information from the received text data and uses map information to calculate the route from the current location to the destination.

[2022] Step 5:

[2023] The server returns the calculated route information to the terminal.

[2024] Step 6:

[2025] The terminal analyzes route information and uses speech synthesis technology to instruct the user verbally on the next direction and distance to go. For example, it might say, "Go 50 meters and turn right."

[2026] Step 7:

[2027] The device uses its camera to capture images of its surroundings in real time.

[2028] Step 8:

[2029] The terminal sends the captured video to the server.

[2030] Step 9:

[2031] The server receives the video data and uses a deep learning model to analyze obstacles and traffic light conditions within the video.

[2032] Step 10:

[2033] The server sends the analysis results back to the terminal. For example, if the traffic light is red, it will send back an instruction such as, "The traffic light is red. Please stop."

[2034] Step 11:

[2035] The device notifies the user of the analysis results via voice. The notification will be in the form of, "The traffic light is red. Please stop."

[2036] Step 12:

[2037] The device captures the user's voice and facial expressions and uses an emotion engine to recognize their current emotional state.

[2038] Step 13:

[2039] The device adjusts the speed and content of voice guidance based on the recognized emotional state. For example, if the user is anxious, the voice guidance will slow down and more detailed instructions will be added.

[2040] Processing steps for the shopping support function

[2041] Step 1:

[2042] The user uses voice input to say "I want to buy tomatoes" to the device.

[2043] Step 2:

[2044] The device captures voice input and uses speech recognition technology to convert the voice signal into text.

[2045] Step 3:

[2046] The terminal sends the converted text data to the server.

[2047] Step 4:

[2048] The server extracts corresponding product information based on the received text data.

[2049] Step 5:

[2050] The user points the camera at the tomato trellis.

[2051] Step 6:

[2052] The terminal captures video of the product shelves in real time and sends that video to the server.

[2053] Step 7:

[2054] The server analyzes the received video data and identifies the recognized products using an image recognition model.

[2055] Step 8:

[2056] The server evaluates the freshness and quality of the products.

[2057] Step 9:

[2058] The server sends the evaluation results back to the terminal. For example, it might return an evaluation result such as, "This tomato is very fresh."

[2059] Step 10:

[2060] The device notifies the user of the evaluation results via voice. The notification may be in the form of, "This tomato is very fresh."

[2061] Step 11:

[2062] The device captures the user's voice and facial expressions and uses an emotion engine to recognize their current emotional state.

[2063] Step 12:

[2064] The device makes product suggestions based on the recognized emotional state. For example, if the user is undecided, it adds suggestions for other products. On the other hand, if the user is in a hurry, it quickly suggests highly-rated products.

[2065] These detailed processing steps enable the system of the present invention to allow visually impaired individuals to independently receive support in various aspects of their daily lives. Furthermore, it detects the user's emotional state and provides optimal support accordingly.

[2066] (Example 2)

[2067] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[2068] There is a need to ensure that visually impaired individuals receive appropriate support when they are able to move around independently and do their daily shopping. In particular, flexible responses that take into account the user's emotional state are necessary in areas such as navigation to destinations, product selection, and recognition of obstacles along the way. However, current systems have difficulty meeting all of these requirements simultaneously, which has been a problem as it makes it inconvenient for visually impaired individuals to move around and shop with peace of mind.

[2069] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[2070] In this invention, the server includes means for receiving voice input from a user, means for converting the received voice input into text, means for identifying a destination based on the converted text and calculating a route using map information, means for communicating the calculated route information to the user by voice, means for capturing and analyzing video of the user's surroundings, means for recognizing obstacles and traffic light conditions in the captured video, means for notifying the user of the recognition results by voice, and means for recognizing the user's emotional state and adjusting voice guidance based on that emotion. This makes it possible for visually impaired people to receive optimal support in real time, tailored to environmental information and their emotional state, when they are able to move around and shop independently with peace of mind.

[2071] "Means for receiving user voice input" refers to a mechanism for capturing the voice spoken by the user into the device.

[2072] "Means of converting received voice input into text" refers to the process of converting voice data into text data.

[2073] "A means of identifying a destination based on converted text and calculating a route using map information" refers to a process that analyzes text data to determine the destination and then calculates a route using map information.

[2074] "A means of communicating calculated route information to the user via voice" refers to a mechanism that informs the user of calculated route information via voice.

[2075] "Means for capturing and analyzing video footage of the user's surroundings" refers to a mechanism that uses a camera to capture video footage of the user's surroundings and then analyzes that footage.

[2076] "Means for recognizing obstacles and traffic light conditions in captured video" refers to technologies for identifying the condition of obstacles and traffic lights from captured video footage.

[2077] "Means for notifying the user of recognition results by voice" refers to a mechanism that communicates the analysis results to the user by voice.

[2078] "Means for recognizing the user's emotional state and adjusting voice guidance based on those emotions" refers to technology that analyzes the user's emotions and adjusts the content and speed of voice guidance accordingly.

[2079] "A method for inputting product names by voice, analyzing that information, and converting it into text" refers to a method in which a user inputs a product name by voice, and that voice is analyzed and converted into text data.

[2080] "Means for recognizing products from camera footage based on converted text" refers to a technology that uses converted text to identify specific products from camera footage.

[2081] "Means for evaluating the freshness and quality of recognized products" refers to the process of analyzing the freshness and quality of identified products.

[2082] "Means of notifying users of evaluation results by voice" refers to a mechanism that informs users of the product evaluation results by voice.

[2083] "A means of identifying the user's current location and updating the route in real time using map information" refers to a technology that obtains the user's current location and recalculates the route in real time based on that location information.

[2084] "A means of instructing users on their direction of travel and next actions via voice based on updated route information" refers to a voice guidance system that uses recalculated route information to provide instructions to users.

[2085] "A means of recognizing moving objects in the surroundings, such as pedestrians and animals other than obstacles and traffic lights, and notifying the user of that information by voice" refers to a technology that identifies moving objects in the surroundings and transmits that information to the user by voice.

[2086] This invention is an assistance system for visually impaired individuals to independently navigate and perform daily shopping, providing optimal support based on the user's emotional state. This system is realized by combining voice input, speech synthesis, image recognition, GPS, map information, AI technology, and emotion recognition technology.

[2087] System Configuration

[2088] The system mainly consists of the following components:

[2089] Speech recognition technology

[2090] Speech synthesis technology

[2091] Image recognition technology

[2092] GPS

[2093] Map information

[2094] emotion recognition technology

[2095] server

[2096] Devices (smartphones and mobile devices)

[2097] Voice input and conversion

[2098] The user provides voice input. For example, they might say, "Take me to the nearest convenience store." This voice data is captured by the device. The device uses speech recognition technology (e.g., Google Cloud Speech-to-Text) to convert the voice data into text data. The converted text data is sent to a server and processed as destination information.

[2099] Route calculation and communication

[2100] The server analyzes the received text data and extracts destination information. Then, using map information (such as the Google Maps API), the server calculates the optimal route from the user's current location to the destination and sends the calculated route information back to the device. The device then uses speech synthesis technology (for example, Amazon Polly) to communicate the calculated route information to the user verbally.

[2101] Surrounding video analysis

[2102] The device uses its camera to capture video of the user's surroundings in real time and sends the video data to a server. The server analyzes the video data using deep learning technology (such as TensorFlow). This analysis allows for the recognition of traffic lights and obstacles. The analysis results are sent back from the server to the device, and the user is notified of the situation via voice. For example, it might say, "The traffic light is red. Please stop."

[2103] Emotion recognition and regulation

[2104] The device captures the user's voice and facial expressions through its camera and microphone. The captured data is analyzed using emotion recognition technology (e.g., IBM Watson Tone Analyzer). Based on the analysis results, the user's emotional state is determined. For example, if the user is anxious, the speed of the voice guidance is slowed down, and more detailed instructions are added.

[2105] Shopping assistance

[2106] The user voice-inputs, "I want to buy tomatoes." The terminal uses speech recognition technology to convert the voice into text and sends it to the server. Based on the received text information, the server recognizes the products from the terminal's camera image. When the user points the camera at the tomato shelf, the image is sent to the server and analyzed using deep learning technology. The freshness and quality of the recognized tomatoes are evaluated, and the evaluation results are sent back to the terminal. The terminal then notifies the user by voice, "These tomatoes are very fresh."

[2107] Emotion recognition and product recommendations

[2108] The device then captures the user's voice and facial expressions and analyzes them using emotion recognition technology. For example, if the user is unsure, it suggests other products on the same shelf. On the other hand, if the user is in a hurry, it quickly suggests the highest-rated product.

[2109] Specific example

[2110] Examples of directions

[2111] User: "Take me to the nearest convenience store" (voice input).

[2112] Terminal: Converts audio data to text and sends it to the server.

[2113] Server: Extracts destination information, calculates the optimal route, and sends it back to the terminal.

[2114] Terminal: Provides calculated route information via voice (e.g., "Go 50 meters and turn right").

[2115] Terminal: Captures camera footage and sends it to the server.

[2116] Server: Analyzes video and recognizes the signal status.

[2117] Terminal: Announces analysis results by voice (e.g., "The traffic light is red. Please stop.").

[2118] Terminal: Adjusts the content of voice guidance according to the user's emotional state.

[2119] Specific examples of shopping

[2120] User: "I want to buy tomatoes" (voice input).

[2121] Terminal: Converts speech to text and sends it to the server.

[2122] Server: Uses a product recognition model, and the user points their camera at a tomato shelf.

[2123] Terminal: Captures video and sends it to the server.

[2124] Server: Analyzes video footage to evaluate the freshness of tomatoes.

[2125] Device: Notifies the evaluation result by voice (e.g., "This tomato is very fresh").

[2126] Terminal: If the user is undecided, suggest other products.

[2127] Example of a prompt

[2128] The following are examples of prompts to input into the generative AI model:

[2129] "Please describe the overview, functions, and specific examples of support systems for visually impaired individuals to shop independently."

[2130] "Please describe the specific processing steps and technologies used in a system that combines speech recognition and image recognition technologies."

[2131] The above describes specific embodiments for carrying out the present invention.

[2132] The flow of the specific processing in Example 2 will be explained using Figure 13.

[2133] Step 1: Obtaining user voice input

[2134] The user voice-inputs "Take me to the nearest convenience store" into their smartphone. The device receives this voice input via its microphone and stores it as audio data.

[2135] Input data: Audio data

[2136] Output data: Stored audio data

[2137] Step 2: Convert audio to text

[2138] The device analyzes the accumulated audio data using speech recognition technology (for example, Google Cloud Speech-to-Text) and converts it into text data.

[2139] Input data: Audio data

[2140] Data processing: Analysis using speech recognition technology

[2141] Output data: Text data

[2142] Step 3: Send destination information

[2143] The converted text data is sent from the terminal to the server. The server receives this text data and processes it as destination information.

[2144] Input data: Text data

[2145] Data processing: Data transmission

[2146] Output data: Text data sent to the server

[2147] Step 4: Calculating square roots

[2148] The server parses the received text data and extracts destination information. Next, it uses map information (for example, Google Maps API) and GPS information to calculate the route from the current location to the destination.

[2149] Input data: Text data, map information, GPS information

[2150] Data processing: Extraction of destination information and route calculation.

[2151] Output data: Route information

[2152] Step 5: Return route information

[2153] The calculated route information is sent back from the server to the terminal. The terminal receives this route information.

[2154] Input data: Route information

[2155] Data processing: Data transmission

[2156] Output data: Route information sent to the terminal

[2157] Step 6: Generate audio guide

[2158] The device analyzes the received route information and generates a voice guide using speech synthesis technology (for example, Amazon Polly).

[2159] Input data: Route information

[2160] Data processing: Voice guide generation using speech synthesis technology

[2161] Output data: Audio guide

[2162] Step 7: Providing audio guides

[2163] The device provides the user with generated audio guidance. Specifically, it gives instructions via voice, such as, "Go 50 meters and turn right."

[2164] Input data: Audio guide

[2165] Data processing: Audio guide playback

[2166] Output data: Voice guidance

[2167] Step 8: Video Capture

[2168] The device uses its camera to capture video footage of the user's surroundings in real time. This video data is then sent to a server.

[2169] Input data: Camera video

[2170] Data processing: Video capture and transmission

[2171] Output data: Video data sent to the server

[2172] Step 9: Video Analysis

[2173] The server analyzes the received video data using deep learning technology (for example, TensorFlow). This allows it to recognize traffic lights and obstacles.

[2174] Input data: Video data

[2175] Data processing: Video analysis using deep learning technology

[2176] Output data: Recognition result

[2177] Step 10: Return the recognition results

[2178] The server sends the analysis results back to the terminal. Specifically, it includes information such as, "The signal is red. Please stop." The terminal receives this information.

[2179] Input data: Recognition result

[2180] Data processing: Data transmission

[2181] Output data: Recognition results sent to the terminal

[2182] Step 11: Notification of Recognition Results

[2183] The device notifies the user of the received recognition results via voice.

[2184] Input data: Recognition result

[2185] Data processing: Voice notification

[2186] Output data: Voice notification

[2187] Step 12: Recognizing your emotional state

[2188] The device captures the user's voice and facial expressions and analyzes their emotional state using emotion recognition technology (for example, IBM Watson Tone Analyzer).

[2189] Input data: Voice and facial expression data

[2190] Data processing: Analysis using emotion recognition technology

[2191] Output data: Emotional state

[2192] Step 13: Adjusting the voice guidance

[2193] The device adjusts the content and speed of voice guidance based on the recognized emotional state. For example, if the user is anxious, the guidance voice will slow down and more detailed explanations will be added.

[2194] Input data: Emotional state

[2195] Data processing: Adjustment of the content and speed of voice guidance.

[2196] Output data: Adjusted voice guidance

[2197] (Application Example 2)

[2198] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[2199] Conventional support systems for the visually impaired possessed basic functions such as voice guidance and image recognition, but lacked adaptive support that took into account the user's emotional state. As a result, they were unable to provide appropriate guidance in situations where the user felt anxious or unsure, or adequate support when the user was unsure about product selection. Furthermore, they were insufficient in supporting independent shopping in physical stores and in real-time adjustments during route guidance. Therefore, this invention proposes a system that solves these problems and provides optimal support according to the user's emotional state.

[2200] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for capturing the user's voice and facial expressions and recognizing their emotional state; means for adjusting the speed and detail of voice guidance based on the recognized emotional state; means for recognizing the user's emotional state and suggesting the most suitable product according to the emotion; and means for dynamically adjusting the guidance content based on the emotional state. As a result, the user can receive guidance and support optimized for their own emotional state.

[2201] "Users" refers to the users of the system, and specifically to individuals with visual impairments.

[2202] "Voice input" refers to the act of capturing the user's voice into the system via a microphone.

[2203] "Text conversion" is the process of converting voice input into text information.

[2204] "Destination" refers to the specific place the user wishes to travel to.

[2205] "Map information" refers to data that includes the route from the current location to the destination and surrounding geographical information.

[2206] "Route calculation" is the process of using map information to determine the optimal route from the current location to the destination.

[2207] "Voice guidance" refers to the act of conveying calculated route information and other instructions to users via voice.

[2208] "Video capture" is the act of capturing images of the surroundings using devices such as cameras.

[2209] "Video analysis" is the process of processing captured video data to identify specific information or objects.

[2210] An "obstacle" refers to a physical object that hinders the user's progress.

[2211] A "traffic light" is a light signaling device used to control road traffic.

[2212] "Recognition results" refer to information obtained through methods such as video analysis and audio analysis.

[2213] "Emotional state" refers to the user's current psychological and emotional condition.

[2214] "Emotion recognition means" refers to tools and technologies used to identify a user's emotional state by analyzing their voice and facial expressions.

[2215] "Product" refers to a specific item that a customer wishes to purchase within a physical store.

[2216] "Product recognition" is the process of identifying a specific product from camera footage or other images.

[2217] "Freshness" refers to an indicator related to the freshness and quality of a product.

[2218] "Quality evaluation" is the process of analyzing the freshness and condition of a product to determine its value.

[2219] "Voice notification" refers to the act of a system communicating analysis results, instructions, and other information to the user via voice.

[2220] "Guidance adjustment" is the process of appropriately changing the content and speed of guidance according to the user's emotional state and circumstances.

[2221] This invention is a support system for visually impaired individuals to independently navigate and perform daily shopping. By incorporating an emotion engine, the system provides optimal support based on the user's current emotional state. This system is realized by combining voice input, speech synthesis, image recognition, GPS, map information, AI technology, and emotion recognition technology.

[2222] System Configuration

[2223] Voice Input and Speech Synthesis: This system uses a terminal equipped with a microphone to accept voice input. Speech synthesis is performed using a speech synthesis library such as pyttsx3.

[2224] Image Recognition: A device equipped with a camera captures video of the surroundings and sends the data to a server. The server uses deep learning technology and image analysis tools such as OpenCV to analyze the video and recognize obstacles and the status of traffic lights.

[2225] GPS and Map Information: Use GPS modules or APIs (e.g., Google Maps API) to obtain current location information and calculate the optimal route to the destination.

[2226] Emotion Recognition: The `pipeline` function from the `transformers` library is used to recognize the user's emotional state from their voice and facial expressions.

[2227] Main Features

[2228] 1. In-store navigation

[2229] Users use their smartphone's voice input function to give commands such as, "Take me to the nearest milk stand."

[2230] The voice input is converted to text using the speech_recognition library and sent to the server.

[2231] The server uses map information based on the text to calculate the route. It then sends the calculation result back to the terminal and provides the user with voice guidance such as, "Go 20 meters and turn right."

[2232] 2. Product Recognition and Evaluation

[2233] When a user is looking for a specific product, they can use voice input, for example, by saying, "Tell me how fresh this milk is."

[2234] A camera is used to photograph the target product, and the video is sent to a server.

[2235] The server uses a deep learning model to recognize products and evaluate their freshness and quality. The evaluation results are sent back to the terminal, and the user is notified via voice message, "This milk is very fresh."

[2236] 3. Support tailored to emotional state

[2237] The system captures the user's voice and facial expressions, and recognizes their emotional state using the transformers library's pipeline.

[2238] Based on the recognized emotional state, the system automatically adjusts the speed and level of detail of voice guidance. For example, if the user is anxious, the guidance will be slow and detailed; if they are relaxed, the guidance will be at a normal speed.

[2239] Examples of specific cases and prompt statements

[2240] Specific example

[2241] User: "Take me to the nearest milk stand."

[2242] Terminal message: "Go 20 meters and turn right."

[2243] User: While shopping, asked, "Please tell me how fresh this milk is."

[2244] Device: Captures camera footage, analyzes it on the server, and then sends a notification saying, "This milk is very fresh."

[2245] Example of a prompt

[2246] "Take me to the nearest milk stand."

[2247] "Please tell me how fresh this milk is."

[2248] "Please turn right next."

[2249] "There is an obstacle in the direction you are going."

[2250] "This product is extremely fresh."

[2251] This invention enables users to navigate and shop efficiently and safely in physical stores by receiving optimized guidance and support tailored to their emotional state. Furthermore, emotion recognition technology can reduce the psychological burden on users.

[2252] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[2253] Step 1:

[2254] Users use their smartphone's voice input function to enter commands regarding destinations or products. This voice input is captured by the smartphone's built-in microphone. The input data consists of the user's voice, specifically commands such as "Take me to the nearby milk section" or "Tell me the freshness of this milk."

[2255] Step 2:

[2256] The voice input data is converted into text data using the terminal's speech_recognition library. At this stage, the input is the user's voice data, and the output is text data obtained by analyzing the voice. The text data will be in the format of "Take me to the nearest milk stand" or "Tell me how fresh this milk is."

[2257] Step 3:

[2258] The data, converted to text, is sent to the server, which analyzes the data to identify destination and product information. The input here is a text-converted instruction, which the server then analyzes. The output includes destination information (milk section) and product information (milk).

[2259] Step 4:

[2260] Once a destination is identified, the server uses map information to calculate the optimal route from the current location to the destination. The input data consists of the current location and destination information, and the server executes a route calculation algorithm based on this data. The output data generates detailed route information, such as "Go 20 meters and turn right."

[2261] Step 5:

[2262] The calculated route information is sent back to the terminal, which then uses a speech synthesis library (such as pyttsx3) to provide voice guidance. The input data is route information obtained from the server, which the terminal converts into voice data. The output data is the voice guidance that the user can hear. Specifically, it generates voice guidance such as, "Go 20 meters and turn right."

[2263] Step 6:

[2264] When a user selects a product, the camera is used to photograph the product, and the video data is sent to the server. The input data for this step is the video data captured by the camera. The video data is an image containing the product, specifically including images of milk, for example.

[2265] Step 7:

[2266] The server analyzes the received video data using deep learning techniques and image recognition tools (such as OpenCV) to identify products. The input data is video data transmitted from the camera, and the output data is information about the analyzed products. Specifically, it recognizes milk in the video and evaluates its quality and freshness.

[2267] Step 8:

[2268] Product information analyzed by the server undergoes a quality evaluation, and the evaluation results are sent back to the terminal. The input data is the product recognition result, and the server performs the evaluation using a quality evaluation algorithm. The output data is the freshness evaluation result, and information such as "This milk is very fresh" can be obtained.

[2269] Step 9:

[2270] The terminal uses speech synthesis technology (such as pyttsx3) to notify the user of the received evaluation results via voice. The input data is the freshness evaluation results returned from the server, which the terminal converts into voice data. The output data is a voice notification of the quality evaluation results that the user can hear.

[2271] Step 10:

[2272] The user's voice and facial expressions are captured and acquired by the device as data to recognize their emotional state. The input data consists of the user's voice and facial expressions, which the device captures.

[2273] Step 11:

[2274] The acquired emotion data is analyzed using emotion recognition technology (the pipeline function in the transformers library). The input data consists of voice and facial capture data, and the output data is the recognized emotional state. Specifically, emotional states such as "anxious" or "relaxed" can be obtained.

[2275] Step 12:

[2276] Based on the recognized emotional state, the speed and detail of the voice guidance are adjusted. The input data is the result of emotion recognition, and the terminal changes the voice guidance settings based on this. The output data is the adjusted voice guidance; for example, if the user is "anxious," slow and detailed guidance will be provided.

[2277] Through this series of steps, users can independently navigate and shop in physical stores, and receive optimal support tailored to their emotional state.

[2278] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[2279] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2280] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[2281] [Fourth Embodiment]

[2282] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[2283] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[2284] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[2285] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[2286] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[2287] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[2288] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[2289] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[2290] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[2291] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[2292] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[2293] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[2294] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[2295] System Overview

[2296] This invention is an assistance system for visually impaired individuals to enable them to move around and perform daily shopping independently. This system combines voice input, speech synthesis, image recognition, GPS, map information, and AI technology. The embodiments are described below.

[2297] Navigation function

[2298] Enter destination

[2299] The user voice-inputs "Take me to the nearest convenience store" into their smartphone. The device uses speech recognition technology to convert the voice into text. This text information is sent to a server and processed as destination information.

[2300] Root calculation

[2301] The server calculates a route using map information based on the received text information (destination). Specifically, it uses an API to calculate the optimal walking route from the current location to the destination. The calculated route information is then returned to the device.

[2302] Voice directions

[2303] The terminal analyzes the calculated route information and provides voice instructions to the user regarding the next direction and distance to travel. For example, it might say, "Go 50 meters and turn right," providing step-by-step guidance.

[2304] Real-time video analysis

[2305] The device uses a camera to capture video of the user's surroundings in real time. This video data is sent to a server, which analyzes the video using deep learning technology. Specifically, it recognizes the status of traffic lights and obstacles. The results of this analysis are returned to the device, and the user is notified of the situation via voice, such as, "The traffic light is red. Please stop."

[2306] Shopping assistance

[2307] Product Selection

[2308] The user voice-inputs, "I want to buy tomatoes." The device then uses speech recognition technology to convert this voice into text and sends it to the server.

[2309] Product recognition

[2310] The server uses a model to recognize products from the terminal's camera footage based on the received text information (product information). When the user points the camera at a tomato shelf, the footage is sent to the server, and image analysis is performed using deep learning technology.

[2311] Freshness evaluation

[2312] The freshness and quality of tomatoes are evaluated based on image analysis. For example, the evaluation is based on color, shape, and gloss. The evaluation results are returned to the device, and the user is notified by voice, "This tomato is very fresh."

[2313] Specific example

[2314] Examples of directions

[2315] User: Inputs "Take me to the nearest convenience store" by voice.

[2316] Terminal: Converts speech to text and sends it to the server.

[2317] Server: Uses the Google Maps API to calculate the route from the current location to the destination and sends it back to the device.

[2318] Terminal: Provides voice guidance to the user regarding the calculated route information (e.g., "Go 50 meters and turn right").

[2319] Terminal: Captures camera footage and sends it to a server to recognize the status of traffic lights.

[2320] Server: Recognizes that the traffic light is red and notifies the terminal with the message, "The traffic light is red. Please stop."

[2321] Terminal: Notifies the user via voice.

[2322] Specific examples of shopping

[2323] User: Inputs "I want to buy tomatoes" by voice.

[2324] Terminal: Converts speech to text and sends it to the server.

[2325] Server: Uses a model for product recognition.

[2326] User: Point the camera at the tomato trellis.

[2327] Terminal: Captures camera footage and sends it to the server.

[2328] Server: Evaluates the freshness and quality of tomatoes and sends the results back to the terminal.

[2329] Device: Notifies the user of the evaluation results via voice (e.g., "This tomato is very fresh").

[2330] In this way, the system of the present invention can effectively provide support to visually impaired persons when they are able to move around and shop independently.

[2331] The following describes the processing flow.

[2332] Processing steps for the navigation function

[2333] Step 1:

[2334] The user uses voice input to tell the terminal, "Take me to the nearest convenience store."

[2335] Step 2:

[2336] The device captures voice input and uses speech recognition technology to convert the voice signal into text.

[2337] Step 3:

[2338] The terminal sends the converted text data to the server.

[2339] Step 4:

[2340] The server extracts destination information from the received text data and uses map information to calculate the route from the current location to the destination.

[2341] Step 5:

[2342] The server returns the calculated route information to the terminal.

[2343] Step 6:

[2344] The terminal analyzes route information and uses speech synthesis technology to instruct the user verbally on the next direction and distance to go. For example, it might say, "Go 50 meters and turn right."

[2345] Step 7:

[2346] The device uses its camera to capture images of its surroundings in real time.

[2347] Step 8:

[2348] The terminal sends the captured video to the server.

[2349] Step 9:

[2350] The server receives the video data and uses a deep learning model to analyze obstacles and traffic light conditions within the video.

[2351] Step 10:

[2352] The server sends the analysis results back to the terminal. For example, if the traffic light is red, it will send back an instruction such as, "The traffic light is red. Please stop."

[2353] Step 11:

[2354] The device notifies the user of the analysis results via voice. The notification will be in the form of, "The traffic light is red. Please stop."

[2355] Processing steps for the shopping support function

[2356] Step 1:

[2357] The user uses voice input to say "I want to buy tomatoes" to the device.

[2358] Step 2:

[2359] The device captures voice input and uses speech recognition technology to convert the voice signal into text.

[2360] Step 3:

[2361] The terminal sends the converted text data to the server.

[2362] Step 4:

[2363] The server extracts product information from the received text data and selects an appropriate image recognition model.

[2364] Step 5:

[2365] The user points the camera at the tomato trellis.

[2366] Step 6:

[2367] The device uses a camera to capture real-time video of the product shelves.

[2368] Step 7:

[2369] The terminal sends the captured video to the server.

[2370] Step 8:

[2371] The server receives the video data and uses an image recognition model to identify tomatoes in the video.

[2372] Step 9:

[2373] The server evaluates the freshness and quality of the identified tomatoes.

[2374] Step 10:

[2375] The server sends the evaluation results back to the terminal. For example, it might return an evaluation result such as, "This tomato is very fresh."

[2376] Step 11:

[2377] The device notifies the user of the evaluation results via voice. The notification may be in the form of, "This tomato is very fresh."

[2378] These detailed processing steps enable the system of the present invention to allow visually impaired individuals to independently receive support in various aspects of their daily lives.

[2379] (Example 1)

[2380] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[2381] There is a lack of technical support for visually impaired individuals to move independently, safely, and comfortably, and to perform daily shopping. To address this challenge, a system is needed that recognizes destinations and products through voice input and provides appropriate instructions through voice guidance and video analysis. Furthermore, real-time updated navigation information and product freshness assessments based on this system are also required.

[2382] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[2383] In this invention, the server includes means for receiving voice input from a user, means for converting the received voice input into text, means for identifying a destination based on the converted text and calculating a route using map information, means for communicating the calculated route information to the user by voice, means for capturing and analyzing video of the user's surroundings, means for recognizing obstacles and traffic light conditions in the captured video, and means for notifying the user of the recognition results by voice. This enables visually impaired individuals to recognize destinations and products through voice input and to move and shop safely while receiving appropriate guidance and instructions in real time.

[2384] "Means for receiving user voice input" refers to a device or program that allows visually impaired individuals to input voice commands using a smartphone or other device.

[2385] "Means for converting received voice input into text" refers to a device or program that uses speech recognition technology to convert voice data into text data.

[2386] "Means for identifying a destination based on converted text and calculating a route using map information" refers to a device or program that analyzes text data, refers to a map database based on a specified destination, and calculates the optimal route.

[2387] "Means of conveying calculated route information to users via voice" refers to a device or program that provides calculated route information as a voice message using speech synthesis technology.

[2388] "Means for capturing and analyzing video footage of the user's surroundings" refers to a device or program that uses a camera to capture images of the user's surroundings in real time and analyzes that image data.

[2389] "Means for recognizing obstacles and traffic light status in captured video" refers to a device or program that uses deep learning technology to recognize information about obstacles and traffic lights from acquired video data.

[2390] "Means for notifying the user of the recognition results by voice" refers to a device or program that provides the recognized information as a voice message using speech synthesis technology.

[2391] "A means of inputting a product name by voice, analyzing that information, and converting it into text" refers to a device or program that allows a user to input a product name by voice and convert that information into text data.

[2392] "Means for recognizing products from camera footage" refers to a device or program that uses deep learning technology to identify specific products from video data acquired using a camera.

[2393] "Means for evaluating the freshness and quality of recognized goods" refers to a device or program that evaluates the freshness and quality of recognized goods based on criteria such as color, shape, and gloss.

[2394] "Means for notifying users of evaluation results by voice" refers to a device or program that provides evaluation results of freshness and quality as a voice message using speech synthesis technology.

[2395] "Means for updating routes in real time" refers to a device or program that dynamically updates the navigation route for each action based on the user's current location information.

[2396] "Means of providing voice instructions for the direction of travel or the next action" refers to a device or program that uses speech synthesis technology to provide the user with voice messages indicating the direction and action they should take next.

[2397] "Means for recognizing moving objects such as pedestrians and animals in the surrounding area and notifying the user of that information via voice" refers to a device or program that recognizes moving objects from camera footage and provides information for avoiding danger as a voice message.

[2398] System Overview

[2399] This invention is an assistance system for visually impaired individuals to enable them to move around and perform daily shopping independently. This system is comprised of a combination of voice input, speech synthesis, image recognition, GPS, map information, and AI technology.

[2400] Navigation function

[2401] Enter destination

[2402] The user inputs a voice command into their smartphone saying, "Take me to the nearest convenience store."

[2403] The device uses speech recognition technology to convert this speech into text. Specifically, it uses Google Cloud Speech-to-Text.

[2404] The device sends the generated text data ("Please guide me to the nearest convenience store") to the server.

[2405] The server receives the text data and uses it to determine the destination.

[2406] Root calculation

[2407] The server determines the user's current location based on the destination information received.

[2408] The server uses map information and the Google Maps API to calculate the optimal walking route from the current location to the destination.

[2409] The server generates the calculated route information in JSON format and sends it back to the terminal.

[2410] Voice directions

[2411] The terminal parses the received route information in JSON format.

[2412] To guide the user verbally in the next direction and distance, a voice message is generated using Google Cloud Text-to-Speech.

[2413] For example, based on route information, it might instruct you to "Go 50 meters and turn right."

[2414] Real-time video analysis

[2415] The device activates its camera and captures video of the user's surroundings in real time.

[2416] The terminal sends the captured video data to the server.

[2417] The server analyzes the received video using deep learning technology (TensorFlow). It recognizes traffic lights and obstacles, generates that information in text format, and sends it back to the terminal.

[2418] The device generates a voice message using speech synthesis technology (Google Cloud Text-to-Speech) based on the received recognition results and notifies the user.

[2419] For example, it might announce, "The traffic light is red. Please stop."

[2420] Shopping assistance

[2421] Product Selection

[2422] The user voice-inputs "I want to buy tomatoes" into their smartphone.

[2423] The device uses speech recognition technology (Google Cloud Speech-to-Text) to convert this speech into text.

[2424] The terminal sends the generated text data ("I want to buy tomatoes") to the server.

[2425] The server receives the text data and processes it as product information.

[2426] Product recognition

[2427] The server uses a deep learning model to recognize tomatoes based on product information.

[2428] The user points the camera at the tomato trellis.

[2429] The device captures camera footage and sends it to the server.

[2430] The server analyzes the received video using deep learning technology (TensorFlow) to recognize tomatoes.

[2431] Freshness evaluation

[2432] The server evaluates the freshness and quality of the recognized tomatoes. Specifically, it evaluates them based on criteria such as color, shape, and gloss.

[2433] The server generates the evaluation results in text format and sends them back to the terminal.

[2434] The device generates a voice message saying, "These tomatoes are very fresh," and notifies the user.

[2435] Specific example

[2436] Examples of directions

[2437] User: Inputs "Take me to the nearest convenience store" by voice.

[2438] Terminal: Use Google Cloud Speech-to-Text to convert speech to text and send it to the server.

[2439] Server: Uses the Google Maps API to calculate the route from the current location to the destination and returns it to the device in JSON format.

[2440] Terminal: Analyzes the received route information and uses Google Cloud Text-to-Speech to provide voice guidance such as, "Go 50 meters and turn right."

[2441] Terminal: Captures camera footage and sends it to a server to recognize the status of traffic lights.

[2442] Server: Uses deep learning technology (TensorFlow) to recognize that the signal is red and sends a message back to the terminal.

[2443] Terminal: Based on the received information, it notifies the user by voice, "The traffic light is red. Please stop."

[2444] Specific examples of shopping

[2445] User: Inputs "I want to buy tomatoes" by voice.

[2446] Terminal: Use Google Cloud Speech-to-Text to convert speech to text and send it to the server.

[2447] Server: Prepares to recognize tomatoes using deep learning technology.

[2448] User: Point the camera at the tomato trellis.

[2449] Terminal: Captures camera footage and sends it to the server.

[2450] Server: Uses deep learning technology (TensorFlow) to analyze received video and evaluate the freshness and quality of tomatoes.

[2451] Server: Generates evaluation results in text format and sends them back to the terminal.

[2452] Device: Use Google Cloud Text-to-Speech to announce, "These tomatoes are very fresh."

[2453] In this way, the system of the present invention can effectively provide support to visually impaired persons when they are able to move around and shop independently.

[2454] The flow of the specific processing in Example 1 will be explained using Figure 11.

[2455] Specific processing flow of the navigation function

[2456] Enter destination

[2457] Step 1:

[2458] The user uses voice input on their smartphone, saying, "Take me to the nearest convenience store."

[2459] Input: Voice command ("Take me to the nearest convenience store.")

[2460] Output: Audio data

[2461] Step 2:

[2462] The device uses speech recognition technology (Google Cloud Speech-to-Text) to convert the input voice data into text data.

[2463] Specifically, audio data is sent to the server, and the server returns text data.

[2464] Input: Audio data

[2465] Output: Text data ("Please guide me to the nearest convenience store")

[2466] Step 3:

[2467] The terminal sends the generated text data to the server. The server receives this data and processes it as destination information.

[2468] Input: Text data ("Please guide me to the nearest convenience store")

[2469] Output: Destination information

[2470] Root calculation

[2471] Step 4:

[2472] The server determines the user's current location based on the destination information received.

[2473] Input: Destination information

[2474] Output: Current location information

[2475] Step 5:

[2476] The server uses the Google Maps API to calculate the optimal walking route from the current location to the destination.

[2477] Input: Current location information, destination information

[2478] Output: Route information (JSON format)

[2479] Step 6:

[2480] The server generates the calculated route information in JSON format and sends it back to the terminal.

[2481] Input: Route information (in generated JSON format)

[2482] Output: Route information (JSON format)

[2483] Voice directions

[2484] Step 7:

[2485] The terminal analyzes the received route information and calculates the next direction and distance to proceed.

[2486] Input: Route information (JSON format)

[2487] Output: Guidance / Instructions

[2488] Step 8:

[2489] The device uses Google Cloud Text-to-Speech to generate guidance instructions as voice messages and notify the user.

[2490] For example, you might give directions like, "Go 50 meters and turn right."

[2491] Input: Guidance / Instructions

[2492] Output: Voice message

[2493] Real-time video analysis

[2494] Step 9:

[2495] The device activates its camera and captures video of the user's surroundings in real time.

[2496] Input: Camera video

[2497] Output: Video data

[2498] Step 10:

[2499] The terminal sends the captured video data to the server.

[2500] Input: Video data

[2501] Output: Video data sent to the server

[2502] Step 11:

[2503] The server uses deep learning technology (TensorFlow) to analyze the received video and recognize the status of traffic lights and obstacles.

[2504] Input: Video data

[2505] Output: Analysis results (status of traffic lights and obstacles)

[2506] Step 12:

[2507] The server generates the analysis results in text format and sends them back to the terminal.

[2508] Input: Analysis results

[2509] Output: Text data (analysis results)

[2510] Step 13:

[2511] The device generates a voice message using Google Cloud Text-to-Speech based on the received recognition results and notifies the user.

[2512] For example, you might announce, "The traffic light is red. Please stop."

[2513] Input: Text data (analysis results)

[2514] Output: Voice message

[2515] Specific processing flow for shopping assistance

[2516] Product Selection

[2517] Step 1:

[2518] The user voice-inputs "I want to buy tomatoes" into their smartphone.

[2519] Input: Voice command ("I want to buy tomatoes")

[2520] Output: Audio data

[2521] Step 2:

[2522] The device uses speech recognition technology (Google Cloud Speech-to-Text) to convert this speech into text.

[2523] Specifically, audio data is sent to the server, and the server returns text data.

[2524] Input: Audio data

[2525] Output: Text data ("I want to buy tomatoes")

[2526] Step 3:

[2527] The terminal sends the generated text data to the server. The server receives this data and processes it as product information.

[2528] Input: Text data ("I want to buy tomatoes")

[2529] Output: Product Information

[2530] Product recognition

[2531] Step 4:

[2532] The server prepares a product recognition model based on the product information.

[2533] Input: Product Information

[2534] Output: Recognition Model

[2535] Step 5:

[2536] The user points the camera at the tomato trellis.

[2537] Input: Image of a product shelf

[2538] Output: Camera data

[2539] Step 6:

[2540] The device captures camera footage and sends it to the server.

[2541] Input: Camera data

[2542] Output: Video data sent to the server

[2543] Step 7:

[2544] The server uses deep learning technology (TensorFlow) to analyze the received video and recognize the product.

[2545] Input: Video data

[2546] Output: Analysis results (product recognition)

[2547] Step 8:

[2548] The server generates the recognition results in text format and sends them back to the terminal.

[2549] Input: Analysis results

[2550] Output: Text data (product recognition)

[2551] Freshness evaluation

[2552] Step 9:

[2553] The server evaluates the freshness and quality of the recognized products. Specifically, it evaluates them based on criteria such as color, shape, and gloss.

[2554] Input: Product recognition result

[2555] Output: Freshness evaluation results

[2556] Step 10:

[2557] The server generates the evaluation results in text format and sends them back to the terminal.

[2558] Input: Freshness evaluation results

[2559] Output: Text data (evaluation results)

[2560] Step 11:

[2561] The device uses Google Cloud Text-to-Speech to generate the evaluation results as an audio message and notify the user.

[2562] For example, you might say, "These tomatoes are extremely fresh."

[2563] Input: Text data (evaluation results)

[2564] Output: Voice message

[2565] (Application Example 1)

[2566] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[2567] Visually impaired individuals face numerous challenges when attempting to travel and shop independently. In particular, accurately perceiving their surroundings is difficult when seeking directions to a destination or purchasing goods. They struggle to recognize traffic signals, pedestrians, animals, and other moving objects. Furthermore, locating products and assessing their freshness and quality is extremely challenging. Effective support systems are needed to address these issues.

[2568] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[2569] In this invention, the server includes: means for receiving voice input from the user; means for converting the received voice input into text; means for identifying a destination based on the converted text and calculating a route using map information; means for communicating the calculated route information to the user by voice; means for capturing and analyzing video of the user's surroundings; means for recognizing obstacles and traffic light conditions in the captured video; means for notifying the user of the recognition results by voice; means for the user to input a product name by voice, analyzing that information and converting it into text; means for recognizing a product from the camera image based on the converted text; and means for providing information about the recognized product. The system includes means for evaluating degree and quality; means for notifying the user of the evaluation results by voice; means for identifying the user's current location and updating the route in real time using map information; means for instructing the user on the direction of travel and the next action by voice based on the updated route information; means for recognizing moving objects such as pedestrians and animals in the surroundings other than obstacles and traffic lights, and notifying the user of that information by voice; means for the user to input searched products by voice, and for analyzing that information and converting it into text; means for calculating the evaluation results of product candidates based on the converted text using deep learning technology; and means for guiding the user to the desired product in real time based on the evaluation results. This enables visually impaired people to move and shop independently, safely and effectively.

[2570] A "system" refers to the entire apparatus, including a series of means, designed to support visually impaired individuals in independently navigating and shopping.

[2571] "Users" refers to visually impaired individuals who use this system to receive guidance and shopping assistance.

[2572] "Voice input" refers to the voice signals that a user speaks to the system.

[2573] "Means for receiving voice input" refers to devices or software that acquire the user's voice and allow the system to interpret it.

[2574] "Means of converting speech to text" refers to devices or software that analyze speech input and convert it into textual information.

[2575] "Means of identifying a destination based on text" refers to devices or software that determine a destination based on converted character information.

[2576] "Map information" refers to a database containing geographical information, providing destinations, current location, route information, and more.

[2577] "Means of calculating routes" refers to devices or software that use map information to calculate the optimal route to a destination.

[2578] "Means of conveying route information by voice" refer...

Claims

1. A means of receiving voice input from users, A means of converting received voice input into text, A means of identifying a destination based on converted text and calculating a route using map information, A means of conveying calculated route information to the user via voice, A means of capturing and analyzing video footage of the user's surroundings, A means of recognizing obstacles and traffic light status in captured video, A means of notifying the user of the recognition result by voice, A system that includes this.

2. A method for users to input product names by voice, and for that information to be analyzed and converted into text, A means of recognizing products from camera footage based on the converted text, A means of evaluating the freshness and quality of recognized products, A means of notifying the user of the evaluation results by voice, The system according to claim 1, including the following:

3. A means of identifying the user's current location and updating the route in real time using map information, A means of providing voice instructions to the user regarding the direction of travel and the next action based on updated route information, A means of recognizing moving objects in the surroundings, such as pedestrians and animals other than obstacles and traffic lights, and notifying the user of that information by voice, The system according to claim 1, including the following:

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A