system
A navigation system using AI to analyze braille and QR codes from external cameras provides audio and visual guidance, addressing the challenges of visually impaired and foreign travelers in unfamiliar environments, ensuring safe and efficient travel.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
Visually impaired individuals and foreign travelers face challenges in navigating unfamiliar environments due to the reliance on visual information and lack of support in multiple languages, leading to difficulties in moving safely and efficiently.
A system utilizing an external camera to capture images of braille guidance paths and QR codes, processed by a server with an AI model to provide audio and display guidance in selectable languages, enabling safe and efficient navigation without relying on visual cues.
Enables visually impaired individuals and foreign travelers to navigate safely and efficiently by providing audio and visual guidance based on machine learning analysis of braille paths and QR codes, overcoming location and language barriers.
Smart Images

Figure 2026073449000001_ABST
Abstract
Description
Technical Field
[0004] , , ,
[0005] , , , ,
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Avoiding confusion and danger when visually impaired people or foreign travelers move in unfamiliar places or areas with many obstacles is a major issue. In the current navigation method, since it depends on visual information, it is difficult for visually impaired people to obtain sufficient support. In addition, due to the lack of support in multiple languages that can be easily used, it is difficult for foreign travelers to move efficiently. In such a situation, there is a demand for providing a comprehensive support system that enables people to move safely.
Means for Solving the Problems
[0005] This invention utilizes an external camera to acquire image data, which is then received by a server. The system includes an artificial intelligence model that analyzes braille guidance paths and QR codes (registered trademark) using image recognition technology. Based on the analysis results, it provides audio guidance to visually impaired individuals and display guidance in selectable languages to travelers who require multiple languages. The system of this invention reduces the uncertainty of location information and language barriers faced by visually impaired individuals and foreign travelers, enabling them to reach their destinations safely and efficiently.
[0006] "External devices" refer to devices used to acquire images or codes within a system, and primarily include personal digital assistants (PADs).
[0007] "Image capture device" refers to a camera function mounted on an external device for acquiring images or codes.
[0008] "Image data" refers to digital image information, including braille guidance paths and QR codes, acquired by a camera.
[0009] "Means of receiving" refers to the method of receiving image data sent from an external device and importing it into the server.
[0010] The term "analytical artificial intelligence model" refers to a machine learning algorithm used to recognize and interpret braille guidance paths and QR codes based on received image data.
[0011] "Route guidance information" refers to information about directions provided to the user based on analyzed image data.
[0012] "Audio or display information" refers to guidance information provided to the user as output from an external device, and includes expressions such as voice messages and screen displays.
[0013] "Braille guidance paths" refer to tactile pavement installed to help visually impaired people understand their direction and surroundings while walking.
[0014] A "QR code" is a two-dimensional code used to embed information, and can include location information and related data. [Brief explanation of the drawing]
[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.
Embodiments for Carrying Out the Invention
[0016] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0019] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0020] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0023] [First Embodiment]
[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0036] This invention provides a navigation system for visually impaired people and foreign tourists to travel safely and efficiently, and mainly consists of an external device equipped with a camera, a server, and an artificial intelligence model.
[0037] First, the user uses an external device to acquire images of braille guidance paths and QR codes installed in the environment. This allows the actual movement environment information to be transmitted to the server as digital data.
[0038] The server utilizes an artificial intelligence model to analyze the received image data. This model is built on machine learning techniques and has the ability to accurately recognize braille guidance paths and QR codes. Based on the analysis results, the server calculates the optimal route from the user's current location to their destination and generates route guidance information.
[0039] The terminal receives route guidance information sent back from the server and provides it to the user as audio or screen display information. Audio guidance is primarily used for visually impaired users, while screen displays and audio guidance are provided to foreign travelers using a multi-language selection function. This system allows users to travel safely without relying on surrounding visual information.
[0040] As a concrete example, if a visually impaired person arrives at a train station in a city they are visiting for the first time, they can use their smartphone camera to photograph the tactile guidance path at their feet. This image is sent to a server, which analyzes it and calculates a safe route from the platform to the ticket gate. Then, the user receives instructions via voice message from their device, such as "Please go left. Turn right after 10 meters." This allows the visually impaired person to navigate independently.
[0041] The following describes the processing flow.
[0042] Step 1:
[0043] The user activates the camera on an external device and takes pictures of the tactile paving paths and QR codes on the ground. This image data is saved to the device.
[0044] Step 2:
[0045] The device immediately sends the acquired image data to the server. If a QR code is included, that data is also sent at the same time.
[0046] Step 3:
[0047] The server receives the transmitted image data and passes it to an artificial intelligence model. The model analyzes the braille guidance paths and QR codes to determine the current user's location.
[0048] Step 4:
[0049] The server calculates the optimal route to the destination based on the analysis results and generates route guidance information. This information includes the direction to travel and details about obstacles to be aware of.
[0050] Step 5:
[0051] The server sends the generated routing information back to the terminal.
[0052] Step 6:
[0053] The terminal provides the user with received route guidance information as voice guidance or on-screen display in the user's selected language. The terminal updates the guidance content according to the direction the user is heading.
[0054] Step 7:
[0055] The user travels to their destination according to the provided directions. They can also take additional photos as needed to ensure the accuracy of the directions.
[0056] (Example 1)
[0057] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0058] Navigating safely and efficiently in new environments is often challenging for visually impaired individuals and foreign visitors, due to their inability to rely on visual information. In particular, the lack of methods for reaching destinations without relying on visual guidance makes independent movement difficult. This project aims to address these issues and enable more people to navigate their surroundings with confidence.
[0059] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0060] In this invention, the server includes means for receiving acquired visual data, means for using a machine learning model to generate movement information based on the visual data, and means for transmitting the movement information to an external device as an acoustic or visual output. This enables visually impaired persons and foreign visitors to safely navigate the best route to their destination without relying on visual information.
[0061] "Acquired visual data" refers to digital data containing visual information acquired using external devices.
[0062] A "machine learning model for generating movement information based on visual data" is an algorithm that receives visual data as input, extracts the information necessary for movement based on the model's learning, and generates appropriate movement instructions.
[0063] "Means for transmitting movement information to external devices as acoustic or visual output" refers to technologies for providing the generated movement-related information to the user either as audio output or displayed on a screen such as a display.
[0064] The "best route" refers to the most efficient, safe, and accessible path for a user attempting to travel.
[0065] "External devices" refers to all devices that function as part of a system and are used for acquiring visual data, displaying movement information, and outputting audio.
[0066] This invention is a navigation system aimed at enabling visually impaired individuals and foreign visitors to move safely and efficiently. The system consists of external devices, a server, and a generated AI model.
[0067] Users use their smartphones or dedicated camera equipment to photograph the surrounding braille guidance paths and QR codes. The acquired visual data is transmitted to a server via Wi-Fi or a mobile network.
[0068] The server processes the received visual data using a generative AI model to analyze it. This AI model takes visual data as input and uses machine learning techniques to identify braille pathways and QR codes. Specific instructions for the model are given using prompts. For example, a prompt such as "Identify the pathways and QR codes contained in this image" might be input.
[0069] After analysis, the server calculates the optimal route from the user's current location to their destination and generates route guidance information. This guidance information includes a safe and efficient path for travel.
[0070] The terminal receives routing information from the server and provides guidance to the user via voice or screen display. For visually impaired users, speech synthesis technology is used to deliver the guidance messages. For foreign visitors, screen displays can also be provided in multiple languages.
[0071] As a concrete example, when a visually impaired person uses public transportation in a city they are visiting for the first time, they use their smartphone to photograph the tactile guidance paths within the station. This image data is sent to a server, where an AI model analyzes it and calculates a safe route. Then, the device provides specific instructions via voice, such as "Go straight. Turn right at the next corner."
[0072] Thus, based on the system of the present invention, it becomes possible to move autonomously and safely.
[0073] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0074] Step 1:
[0075] The user uses a smartphone as an external device to photograph surrounding braille guidance paths and QR codes. Visual data acquired from the real-world environment is obtained as input. In this operation, the user launches a camera app, frames the necessary objects, and takes a photograph. Visual data (image file) is generated as output.
[0076] Step 2:
[0077] The device transmits captured visual data to the server. The input is an image file stored on the device. The device then packets this data over the network and transmits it. The output is the digital image data sent to the server.
[0078] Step 3:
[0079] The server activates a generative AI model to analyze the received visual data. The input is the visual data received by the server. The server prompts the AI model with the command, "Identify the guidance paths and QR codes contained in this image," and the model analyzes the image. This analysis outputs information identifying the locations of the braille guidance paths and QR codes.
[0080] Step 4:
[0081] The server calculates the optimal route from the user's current location to their destination based on information obtained from the AI model. The input consists of environmental information obtained through analysis and the user's destination information. Using this data, the server applies a route calculation algorithm and outputs a safe and efficient travel route.
[0082] Step 5:
[0083] The server sends the generated routing information to the terminal. The input is the calculated routing information. The server encrypts this information and transmits it to the terminal over the network. The output is the routing information that reaches the terminal.
[0084] Step 6:
[0085] The terminal analyzes the received route guidance information and provides directions to the user. The input is route guidance information received from the server. The terminal uses speech synthesis technology to output specific instructions in voice, such as "Turn right. There is an elevator 50 meters ahead." It then initiates the guidance process, allowing the user to begin their journey to their destination.
[0086] (Application Example 1)
[0087] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0088] There is a need to alleviate the difficulties that visually impaired people and foreign tourists face in safely and smoothly reaching their destinations in commercial facilities and complex environments, and to provide more accurate and multilingual navigation. In particular, there is a need for technology that can efficiently provide route guidance without relying on visual information.
[0089] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0090] In this invention, the server includes means for receiving visual data acquired by the imaging function of an external device, means for using a machine learning model to analyze the visual data in order to generate route guidance information, means for transmitting the route guidance information to the external device as audio or visual information, and means for providing multilingual guidance to the user on a display terminal device. As a result, the user can move to their destination safely and efficiently without relying on visual guidance.
[0091] An "external device" is a device that is portable to the user and has the function of acquiring and displaying visual data.
[0092] "Shooting function" refers to a function that uses a camera or similar sensor to acquire images or videos.
[0093] "Visual data" refers to image and video information acquired by cameras and sensors.
[0094] "Means of receiving" refers to the function that allows a server to receive data transmitted from an external device.
[0095] A "machine learning model" is an algorithm trained to analyze visual data and extract meaningful information.
[0096] "Means of analysis" refers to the function of analyzing acquired visual data and extracting necessary information.
[0097] "Route guidance information" refers to instructions and navigation information that helps users reach their destination.
[0098] "Means of transmission" refers to the function of sending route guidance information as audio or visual information to an external device and presenting it to the user.
[0099] A "display terminal device" is a device used to provide audio or visual information to a user.
[0100] "Multilingual guidance" refers to a function that provides route guidance information in the language selected by the user.
[0101] The system for carrying out the invention comprises an external device, a server, and a display terminal device. The program for this system begins by acquiring visual data using the camera function of an external device carried by the user. The visual data is image data including identification codes and braille guidance paths.
[0102] The server uses a machine learning model to analyze the received visual data. This model extracts the user's location information from the visual data and generates optimal route guidance information to the destination. Specifically, a machine learning framework such as TENSORFLOW® is used. The server also performs language processing to provide this route guidance information in multiple languages, translating it into the user's preferred language.
[0103] The terminal receives route guidance information transmitted from the server and presents it to the user as audio or visual information. Display terminal devices include smart glasses and smartphones, which display multilingual guidance. Libraries such as Zebra Crossing are used for QR code recognition.
[0104] To give a specific example, a visually impaired person can use smart glasses in a shopping mall to scan a QR code at the entrance. At this time, the server analyzes the route to the "food section" in real time and provides voice guidance such as "Go 10 meters and turn right."
[0105] An example of a prompt sentence to input into the generating AI model is, "Please write about a route guidance system that provides audio guidance to help visually impaired people move safely within a shopping mall."
[0106] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0107] Step 1:
[0108] The user captures the surrounding environment using the camera function of an external device. The input is image data acquired from the external device, and the output is the captured image being saved as data on the external device. In this process, the camera mounted on the external device captures the image and secures the necessary visual data.
[0109] Step 2:
[0110] The terminal sends the captured image data to the server. The input is the image data stored on the external device, and the output is the image data transferred to the server. In this transmission operation, data communication functions are used to deliver the data to the server quickly and securely.
[0111] Step 3:
[0112] The server analyzes the received image data using a machine learning model. The input is the image data sent to the server, and the output is the analyzed location and identification information. The server uses the TensorFlow library to recognize identification codes and braille guidance paths within the image and determine the user's current location.
[0113] Step 4:
[0114] The server generates optimal route guidance information to the destination based on the analysis results. The input is the analyzed location information, and the output is route guidance information. The server uses a generation AI model to generate detailed directions, such as "Go straight to your destination and turn right after 10 meters."
[0115] Step 5:
[0116] The server converts the generated route guidance information into multiple languages and sends it to the terminal. The input is route guidance information, and the output is multilingual guidance information. In this process, natural language processing technology is used to convert the guidance into a language that is easy for the user to understand.
[0117] Step 6:
[0118] The terminal provides the user with received route guidance information as audio or visual information. The input is multilingual guidance information sent from the server, and the output is user-recognizable audio or display. The terminal uses speech synthesis technology or display functionality to convey instructions to the user.
[0119] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0120] The present invention provides a navigation system that takes into account the emotional state of the user when visually impaired people or foreign tourists are traveling, and comprises an external device, a server, an artificial intelligence model, and an emotion engine.
[0121] The user uses an external device to photograph braille guidance paths or QR codes. This image data is immediately transmitted to the server. The server analyzes the received image data using an artificial intelligence model to identify the user's current location and surrounding geographical information. This analysis result serves as the basic data for generating route guidance information.
[0122] The emotion engine analyzes the user's emotional state in real time from their voice and facial expressions using an external device. The analyzed emotional information is used to customize navigation. For example, if the system determines that the user is stressed, the route is adjusted to select a quieter and safer path.
[0123] The server generates optimal route guidance information that also takes into account the user's emotional state and transmits it to an external device. The external device then provides this information to the user as voice guidance or screen display. In this process, the guidance is adjusted according to the user's emotional state to provide more reassuring guidance.
[0124] As a concrete example, when a user gets lost in an area with many signs and a lot of noise, the system senses the user's stress. The emotion engine selects reassuring voices and relaxing routes for the user, and the server calculates a travel route based on this. The terminal then provides a tailored message such as, "Let's try taking a quieter route." In this way, the system provides support to help users reach their destination with peace of mind.
[0125] The following describes the processing flow.
[0126] Step 1:
[0127] The user uses an external device's camera to photograph surrounding braille pathways and QR codes. This data is temporarily stored on the device.
[0128] Step 2:
[0129] The device sends the captured image data to the server. Because this transmission requires real-time processing, the communication is optimized.
[0130] Step 3:
[0131] The server passes the received image data to an artificial intelligence model, which analyzes the braille guidance paths and QR codes. The model then uses this information to determine the user's current location and surrounding environment.
[0132] Step 4:
[0133] The device transmits data obtained from the user's voice input and facial recognition via the camera to the emotion engine. The emotion engine analyzes this data and estimates the user's emotional state.
[0134] Step 5:
[0135] The server generates optimal route guidance information by considering the analyzed location information and the user's emotional state obtained from the emotion engine. If the emotional state indicates stress, it prioritizes selecting a safer route.
[0136] Step 6:
[0137] The server sends the generated routing information back to the terminal.
[0138] Step 7:
[0139] The terminal provides the user with route guidance information received from the server as voice guidance or screen display. The guidance reflects the user's emotional state, employing relaxing voices and designs if necessary.
[0140] Step 8:
[0141] The user follows the provided instructions. If further shooting or emotion recognition is required, they return to step 1 as needed and repeat the process.
[0142] (Example 2)
[0143] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0144] When visually impaired individuals and foreign tourists navigate unfamiliar environments, there is a need for navigation that not only provides location information and routes, but also takes their emotional state into consideration to offer a more reassuring experience. However, current technology lacks navigation systems that consider the user's emotional state, which can increase the user's mental burden. Furthermore, real-time emotion analysis and the mechanism for reflecting it in navigation are complex and difficult to implement.
[0145] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0146] In this invention, the server includes means for receiving image data acquired by an external device's imaging device, means for using a machine learning model to analyze the image data in order to generate route guidance information, means for transmitting the route guidance information to the external device as voice or display data, means for using an emotion analysis engine in the external device to analyze the user's emotional state from their voice and facial expressions, and means for adjusting the route guidance information based on the emotional state. This makes it possible to provide safe and secure routes that are tailored to the user's emotional state.
[0147] An "external device" is a device that a user can carry with them and is used to acquire image data and analyze emotional states.
[0148] A "photography device" is a component mounted on an external device that acquires image data.
[0149] "Image data" refers to data containing visual information collected using a camera or imaging device.
[0150] "Means of receiving" refers to the technology used by a server to receive image data transmitted from an external device.
[0151] A "machine learning model for analysis" is an algorithm used to extract information from image data and generate path guidance information.
[0152] "Route guidance information" refers to information that includes recommended routes and methods for users to reach their destination.
[0153] "Voice or display data" refers to a format in which route guidance information is provided to the user, including voice guidance and visual displays.
[0154] "Means of transmission" refers to the technology for sending analyzed routing information to an external device and providing it to the user.
[0155] An "emotion analysis engine" is a system that identifies a user's emotional state in real time from their voice and facial expressions.
[0156] "Emotional state" refers to the user's mental and emotional state, including stress levels and feelings of security.
[0157] "Means of adjustment" refer to processes and techniques for optimizing path guidance information based on analyzed emotional states.
[0158] A "matrix barcode" is a two-dimensional code used to visually encode information, including formats such as QR codes.
[0159] This invention provides a system that offers navigation that takes into account the emotional state of visually impaired individuals and foreign tourists while they are traveling. The system comprises an external device, a server, a machine learning model, and an emotion analysis engine.
[0160] The user uses a portable external device to photograph braille guidance paths or matrix barcodes. The captured image data is immediately transmitted to a server via the network. The server receives this image data and analyzes it using a machine learning model. This analysis identifies the user's current location and surrounding geographical information, which then serves as the basis for generating route guidance information.
[0161] Furthermore, an emotion analysis engine installed in an external device analyzes the user's voice and facial expressions in real time to understand their emotional state. This information is used to customize navigation. Specifically, if the user is feeling stressed, the system adjusts the route guidance information to select a quieter and safer route.
[0162] The server generates optimal route guidance information that takes emotional states into account and transmits it to an external device as voice or display data. The external device uses speech synthesis technology or a display to provide the user with tailored navigation information. For example, if the user gets lost in a busy environment, it might display a message such as, "Let's try taking a quieter route," to provide reassuring guidance.
[0163] An example of a specific prompt for a generative AI model would be something like, "Analyze the user's current location and suggest the optimal route based on their emotional state." This prompt functions as an instruction for the AI model to generate an appropriate path.
[0164] In this way, the present invention realizes a navigation system that helps users reach their destination safely and comfortably.
[0165] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0166] Step 1:
[0167] The user uses an external camera to photograph braille guidance paths and matrix barcodes. The resulting input is visual data. This data is transmitted to a server via the network. The shooting process involves using a smartphone camera app to acquire visual data with clear focus.
[0168] Step 2:
[0169] The server utilizes a machine learning model to analyze the received visual data. The input is visual data, and the output is the user's current location and environmental information. In this process, the AI model reads matrix barcodes in the visual data and retrieves the corresponding location information from the database.
[0170] Step 3:
[0171] The device analyzes the user's voice and facial expressions in real time using an emotion analysis engine. Input is voice and video data, and output is the user's emotional state. Specifically, the device uses a microphone and camera to collect data and transmits it to the emotion analysis engine.
[0172] Step 4:
[0173] The server generates optimal route guidance information based on emotional state and geographical information. The inputs are the current location, environmental information, and the user's emotional state, and the output is the adjusted route guidance information. In this step, the server calculates and selects a quiet and safe route based on the emotional state.
[0174] Step 5:
[0175] The terminal provides the user with route guidance information received from the server as audio or display data. The input is the adjusted route guidance information, and the output is audio guidance or screen display to aid the user's understanding. As a concrete example, the terminal might output a voice message through its speaker saying, "Let's choose a slightly quieter route."
[0176] (Application Example 2)
[0177] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0178] There is a lack of safer and more comfortable route guidance for visually impaired people and foreign tourists navigating shopping malls and large stores. Furthermore, navigation that takes into account the user's emotional state is not provided, failing to reduce user stress and improve safety.
[0179] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0180] In this invention, the server includes means for receiving visual information acquired by an external imaging device, means for using an artificial intelligence algorithm to analyze the visual information in order to generate route guidance information, and means for analyzing the user's emotional state from their voice or facial expressions. This enables customized route guidance that takes into account the user's emotional state.
[0181] An "external device" is a terminal used to acquire visual information and transmit captured images and encoded data to a server.
[0182] "Visual information" refers to image data and encoded data acquired by external devices, and forms the basis for location information and path analysis.
[0183] An "artificial intelligence algorithm" is a computational method used to analyze acquired visual information and generate appropriate path guidance information.
[0184] "Emotional state" refers to the psychological state of a user as determined from their voice or facial expressions, and is an important element in customizing navigation.
[0185] "Route guidance information" refers to information that includes directions and routes for movement, generated based on analyzed visual information.
[0186] This invention provides a navigation system that takes emotional states into account, enabling users to move around shopping malls and large stores with peace of mind. The system comprises external devices, a server, an artificial intelligence algorithm, and an emotion analysis engine.
[0187] The server receives visual information acquired by an external device's vision sensor. Then, using an artificial intelligence algorithm, it analyzes this visual information to determine the user's location. Next, an emotion analysis engine analyzes the user's emotional state from their real-time voice and facial expressions. Based on this analysis, optimal route guidance information is generated.
[0188] External devices can be smart glasses or mobile terminals that acquire visual information and instantly transmit it to the server. Furthermore, the route guidance received by the user is perceived through voice and screen display. This allows the user to enjoy customized navigation tailored to their emotional state.
[0189] For example, if a user attempts to pass through a crowded area in a large facility, and the emotion analysis engine detects tension, the server calculates a quieter route or one that avoids the crowd and guides the user through an external device.
[0190] An example of a prompt message would be, "We want to develop a navigation system for smart glasses that analyzes the user's emotions and guides them along a safe and comfortable route."
[0191] This invention can be said to be a very useful system for visually impaired people and foreign travelers by providing flexible navigation that takes into account the user's emotional state.
[0192] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0193] Step 1:
[0194] The device uses a visual sensor to acquire visual information about the user's surroundings. Inputs include image data and encoded data (e.g., QR codes). This data is processed on the device and immediately sent to the server.
[0195] Step 2:
[0196] The server inputs the received visual information into an artificial intelligence algorithm. Based on this data, location information and surrounding environment are identified. The AI algorithm performs image analysis and generates output such as the user's current location and surrounding geographical information.
[0197] Step 3:
[0198] The device analyzes the user's voice and facial expressions using an emotion analysis engine. Biometric data (voice and facial expressions) serves as input. This analysis identifies the user's real-time emotional state and generates psychological state data as output.
[0199] Step 4:
[0200] The server combines acquired location information and emotional state to generate optimal route guidance information. This involves data calculations using a generative AI model, resulting in a customized travel route that takes the user's emotional state into consideration.
[0201] Step 5:
[0202] The terminal receives route guidance information transmitted from the server and provides guidance to the user through voice and screen displays. By receiving visually and audibly adjusted information, the user can travel to their destination with peace of mind.
[0203] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0204] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0205] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0206] [Second Embodiment]
[0207] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0208] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0209] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0210] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0211] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0212] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0213] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0214] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0215] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0216] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0217] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0218] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0219] This invention provides a navigation system for visually impaired people and foreign tourists to travel safely and efficiently, and mainly consists of an external device equipped with a camera, a server, and an artificial intelligence model.
[0220] First, the user uses an external device to acquire images of braille guidance paths and QR codes installed in the environment. This allows the actual movement environment information to be transmitted to the server as digital data.
[0221] The server utilizes an artificial intelligence model to analyze the received image data. This model is built on machine learning techniques and has the ability to accurately recognize braille guidance paths and QR codes. Based on the analysis results, the server calculates the optimal route from the user's current location to their destination and generates route guidance information.
[0222] The terminal receives route guidance information sent back from the server and provides it to the user as audio or screen display information. Audio guidance is primarily used for visually impaired users, while screen displays and audio guidance are provided to foreign travelers using a multi-language selection function. This system allows users to travel safely without relying on surrounding visual information.
[0223] As a concrete example, if a visually impaired person arrives at a train station in a city they are visiting for the first time, they can use their smartphone camera to photograph the tactile guidance path at their feet. This image is sent to a server, which analyzes it and calculates a safe route from the platform to the ticket gate. Then, the user receives instructions via voice message from their device, such as "Please go left. Turn right after 10 meters." This allows the visually impaired person to navigate independently.
[0224] The following describes the processing flow.
[0225] Step 1:
[0226] The user activates the camera on an external device and takes pictures of the tactile guidance paths on the ground and the installed QR codes. This image data is saved on the device.
[0227] Step 2:
[0228] The device immediately sends the acquired image data to the server. If a QR code is included, that data is also sent at the same time.
[0229] Step 3:
[0230] The server receives the transmitted image data and passes it to an artificial intelligence model. The model analyzes the braille guidance paths and QR codes to determine the current user's location.
[0231] Step 4:
[0232] The server calculates the optimal route to the destination based on the analysis results and generates route guidance information. This information includes the direction to travel and details about obstacles to be aware of.
[0233] Step 5:
[0234] The server sends the generated routing information back to the terminal.
[0235] Step 6:
[0236] The terminal provides the user with received route guidance information as voice guidance or on-screen display in the user's selected language. The terminal updates the guidance content according to the direction the user is heading.
[0237] Step 7:
[0238] The user travels to their destination according to the provided directions. They can also take additional photos as needed to ensure the accuracy of the directions.
[0239] (Example 1)
[0240] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0241] Navigating safely and efficiently in new environments is often challenging for visually impaired individuals and foreign visitors, due to their inability to rely on visual information. In particular, the lack of methods for reaching destinations without relying on visual guidance makes independent movement difficult. This project aims to address these issues and enable more people to navigate their surroundings with confidence.
[0242] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0243] In this invention, the server includes means for receiving acquired visual data, means for using a machine learning model to generate movement information based on the visual data, and means for transmitting the movement information to an external device as an acoustic or visual output. This enables visually impaired persons and foreign visitors to safely navigate the best route to their destination without relying on visual information.
[0244] "Acquired visual data" refers to digital data containing visual information acquired using external devices.
[0245] A "machine learning model for generating movement information based on visual data" is an algorithm that receives visual data as input, extracts the information necessary for movement based on the model's learning, and generates appropriate movement instructions.
[0246] "Means for transmitting movement information to external devices as acoustic or visual output" refers to technologies for providing the generated movement-related information to the user either as audio output or displayed on a screen such as a display.
[0247] The "best route" refers to the most efficient, safe, and accessible path for a user attempting to travel.
[0248] "External devices" refer to all devices that function as part of a system and are used for acquiring visual data, displaying movement information, and outputting audio.
[0249] This invention is a navigation system aimed at enabling visually impaired individuals and foreign visitors to move safely and efficiently. The system consists of external devices, a server, and a generated AI model.
[0250] Users use their smartphones or dedicated camera equipment to photograph the surrounding braille guidance paths and QR codes. The acquired visual data is transmitted to a server via Wi-Fi or a mobile network.
[0251] The server processes the received visual data using a generative AI model to analyze it. This AI model takes visual data as input and uses machine learning techniques to identify braille pathways and QR codes. Specific instructions for the model are given using prompts. For example, a prompt such as "Identify the pathways and QR codes contained in this image" might be input.
[0252] After analysis, the server calculates the optimal route from the user's current location to their destination and generates route guidance information. This guidance information includes a safe and efficient path for travel.
[0253] The terminal receives routing information from the server and provides guidance to the user via voice or screen display. For visually impaired users, speech synthesis technology is used to deliver the guidance messages. For foreign visitors, screen displays can also be provided in multiple languages.
[0254] As a concrete example, when a visually impaired person uses public transportation in a city they are visiting for the first time, they use their smartphone to photograph the tactile guidance paths within the station. This image data is sent to a server, where an AI model analyzes it and calculates a safe route. Then, the device provides specific instructions via voice, such as "Go straight. Turn right at the next corner."
[0255] Thus, based on the system of the present invention, it becomes possible to move autonomously and safely.
[0256] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0257] Step 1:
[0258] The user uses a smartphone as an external device to photograph surrounding braille guidance paths and QR codes. Visual data acquired from the real-world environment is obtained as input. In this operation, the user launches a camera app, frames the necessary objects, and takes a photograph. Visual data (image file) is generated as output.
[0259] Step 2:
[0260] The device transmits captured visual data to the server. The input is an image file stored on the device. The device then packets this data over the network and transmits it. The output is the digital image data sent to the server.
[0261] Step 3:
[0262] The server activates a generative AI model to analyze the received visual data. The input is the visual data received by the server. The server prompts the AI model with the command, "Identify the guidance paths and QR codes contained in this image," and the model analyzes the image. This analysis outputs information identifying the locations of the braille guidance paths and QR codes.
[0263] Step 4:
[0264] The server calculates the optimal route from the user's current location to their destination based on information obtained from the AI model. The input consists of environmental information obtained through analysis and the user's destination information. Using this data, the server applies a route calculation algorithm and outputs a safe and efficient travel route.
[0265] Step 5:
[0266] The server sends the generated routing information to the terminal. The input is the calculated routing information. The server encrypts this information and transmits it to the terminal over the network. The output is the routing information that reaches the terminal.
[0267] Step 6:
[0268] The terminal analyzes the received route guidance information and provides directions to the user. The input is route guidance information received from the server. The terminal uses speech synthesis technology to output specific instructions in voice, such as "Turn right. There is an elevator 50 meters ahead." It then initiates the guidance process, allowing the user to begin their journey to their destination.
[0269] (Application Example 1)
[0270] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0271] There is a need to alleviate the difficulties that visually impaired people and foreign tourists face in safely and smoothly reaching their destinations in commercial facilities and complex environments, and to provide more accurate and multilingual navigation. In particular, there is a need for technology that can efficiently provide route guidance without relying on visual information.
[0272] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0273] In this invention, the server includes means for receiving visual data acquired by the imaging function of an external device, means for using a machine learning model to analyze the visual data in order to generate route guidance information, means for transmitting the route guidance information to the external device as audio or visual information, and means for providing multilingual guidance to the user on a display terminal device. As a result, the user can move to their destination safely and efficiently without relying on visual guidance.
[0274] An "external device" is a device that is portable to the user and has the function of acquiring and displaying visual data.
[0275] The "imaging function" is a function that utilizes a camera or a similar sensor for acquiring images and videos.
[0276] "Visual data" refers to image and video information acquired by cameras or sensors.
[0277] The "means for receiving" is a function for the server to receive data transmitted from an external device.
[0278] A "machine learning model" is an algorithm trained to analyze visual data and extract meaningful information.
[0279] The "means for analyzing" is a function for analyzing the acquired visual data and extracting necessary information.
[0280] "Route guidance information" refers to instructions and navigation information for a user to reach a destination.
[0281] The "means for transmitting" is a function for sending route guidance information to an external device as audio or visual information and presenting it to the user.
[0282] A "display terminal device" is a device for providing audio or visual information to a user.
[0283] "Multilingual guidance" is a function for providing route guidance information in the language selected by the user.
[0284] The system for implementing the invention has a configuration including an external device, a server, and a display terminal device. The program of this system starts from acquiring visual data by utilizing the imaging function of the external device carried by the user. The visual data is image data including identification codes and Braille guide paths.
[0285] The server analyzes the data using a machine learning model to process the received visual data. This model extracts the user's location information from the visual data and generates optimal route guidance information to the destination. As specific software, a machine learning framework such as TensorFlow is used. The server also performs language processing to provide this route guidance information in multiple languages and converts it into the language corresponding to the user.
[0286] The terminal receives the route guidance information sent from the server and presents it to the user as audio or visual information. Display terminal devices include smart glasses and smartphones, etc., which display multi-language guidance. Also, libraries such as Zebra Crossing are used for QR code recognition.
[0287] Taking a specific example, a visually impaired person uses smart glasses in a shopping mall to scan the QR code at the entrance. At this time, the server analyzes the route to the "food sales area" in real time and provides guidance such as "Go straight for 10 meters and turn right" in voice.
[0288] An example of a prompt sentence input to the generation AI model is "Please write about a route guidance system that provides voice guidance for a visually impaired person to move safely inside a shopping mall."
[0289] The flow of specific processing in Application Example 1 will be described using FIG. 12.
[0290] Step 1:
[0291] The user uses the camera function of an external device to photograph the surrounding environment. The input is the image data obtained from the external device, and the output is that the photographed image is saved as data in the external device. In this process, the camera installed in the external device captures the image and ensures the necessary visual data.
[0292] Step 2:
[0293] The terminal sends the captured image data to the server. The input is the image data stored on the external device, and the output is the image data transferred to the server. In this transmission operation, data communication functions are used to deliver the data to the server quickly and securely.
[0294] Step 3:
[0295] The server analyzes the received image data using a machine learning model. The input is the image data sent to the server, and the output is the analyzed location and identification information. The server uses the TensorFlow library to recognize identification codes and braille guidance paths within the image and determine the user's current location.
[0296] Step 4:
[0297] The server generates optimal route guidance information to the destination based on the analysis results. The input is the analyzed location information, and the output is route guidance information. The server uses a generation AI model to generate detailed directions, such as "Go straight to your destination and turn right after 10 meters."
[0298] Step 5:
[0299] The server converts the generated route guidance information into multiple languages and sends it to the terminal. The input is route guidance information, and the output is multilingual guidance information. In this process, natural language processing technology is used to convert the guidance into a language that is easy for the user to understand.
[0300] Step 6:
[0301] The terminal provides the user with received route guidance information as audio or visual information. The input is multilingual guidance information sent from the server, and the output is user-recognizable audio or display. The terminal uses speech synthesis technology or display functionality to convey instructions to the user.
[0302] Furthermore, an emotion engine for estimating the user's emotions may be combined. That is, the specific processing unit 290 may estimate the user's emotions using the emotion identification model 59 and perform specific processing using the user's emotions.
[0303] The present invention is a system that provides navigation considering the user's emotional state when visually impaired people or foreign travelers move, and includes an external device, a server, an artificial intelligence model, and an emotion engine.
[0304] The user uses an external device to photograph a braille guide path or a QR code. This image data is immediately transmitted to the server. The server analyzes the received image data using an artificial intelligence model to identify the user's current location and surrounding geographical information. This analysis result serves as basic data for generating route guidance information.
[0305] The emotion engine analyzes the user's emotional state in real time from the user's voice and expression on the external device. The analyzed emotion information is used for customizing the navigation. For example, when it is determined that the user is in a stressed state, the route is adjusted to select a quieter and safer route.
[0306] The server generates optimal route guidance information considering the user's emotional state and transmits it to the external device. The external device provides this information to the user as voice guidance or on-screen display. At this time, the guidance is adjusted according to the user's emotional state, and more reassuring guidance content is provided.
[0307] As a specific example, when the user gets lost on the road in an area with many signs and noises, the system senses the user's stress. The emotion engine selects a voice that gives the user a sense of security or a relaxing route, and the server calculates a moving route based on this and provides an adjusted message such as "Let's choose a slightly quieter road" from the terminal. In this way, support is realized for the user to reach the destination with peace of mind.
[0308] The following describes the processing flow.
[0309] Step 1:
[0310] The user uses an external device's camera to photograph surrounding braille pathways and QR codes. This data is temporarily stored on the device.
[0311] Step 2:
[0312] The device sends the captured image data to the server. Because this transmission requires real-time processing, the communication is optimized.
[0313] Step 3:
[0314] The server passes the received image data to an artificial intelligence model, which analyzes the braille guidance paths and QR codes. The model then uses this information to determine the user's current location and surrounding environment.
[0315] Step 4:
[0316] The device transmits data obtained from the user's voice input and facial recognition via the camera to the emotion engine. The emotion engine analyzes this data and estimates the user's emotional state.
[0317] Step 5:
[0318] The server generates optimal route guidance information by considering the analyzed location information and the user's emotional state obtained from the emotion engine. If the emotional state indicates stress, it prioritizes selecting a safer route.
[0319] Step 6:
[0320] The server sends the generated routing information back to the terminal.
[0321] Step 7:
[0322] The terminal provides the user with route guidance information received from the server as voice guidance or screen display. The guidance reflects the user's emotional state, employing relaxing voices and designs if necessary.
[0323] Step 8:
[0324] The user follows the provided instructions. If further shooting or emotion recognition is required, they return to step 1 as needed and repeat the process.
[0325] (Example 2)
[0326] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0327] When visually impaired individuals and foreign tourists navigate unfamiliar environments, there is a need for navigation that not only provides location information and routes, but also takes their emotional state into consideration to offer a more reassuring experience. However, current technology lacks navigation systems that consider the user's emotional state, which can increase the user's mental burden. Furthermore, real-time emotion analysis and the mechanism for reflecting it in navigation are complex and difficult to implement.
[0328] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0329] In this invention, the server includes means for receiving image data acquired by an external device's imaging device, means for using a machine learning model to analyze the image data in order to generate route guidance information, means for transmitting the route guidance information to the external device as voice or display data, means for using an emotion analysis engine in the external device to analyze the user's emotional state from their voice and facial expressions, and means for adjusting the route guidance information based on the emotional state. This makes it possible to provide safe and secure routes that are tailored to the user's emotional state.
[0330] An "external device" is a device that a user can carry with them and is used to acquire image data and analyze emotional states.
[0331] A "photography device" is a component mounted on an external device that acquires image data.
[0332] "Image data" refers to data containing visual information collected using a camera or imaging device.
[0333] "Means of receiving" refers to the technology used by a server to receive image data transmitted from an external device.
[0334] A "machine learning model for analysis" is an algorithm used to extract information from image data and generate path guidance information.
[0335] "Route guidance information" refers to information that includes recommended routes and methods for users to reach their destination.
[0336] "Voice or display data" refers to a format in which route guidance information is provided to the user, including voice guidance and visual displays.
[0337] "Means of transmission" refers to the technology for sending analyzed routing information to an external device and providing it to the user.
[0338] An "emotion analysis engine" is a system that identifies a user's emotional state in real time from their voice and facial expressions.
[0339] "Emotional state" refers to the user's mental and emotional state, including stress levels and feelings of security.
[0340] "Means of adjustment" refer to processes and techniques for optimizing path guidance information based on analyzed emotional states.
[0341] A "matrix barcode" is a two-dimensional code used to visually encode information, including formats such as QR codes.
[0342] This invention provides a system that offers navigation that takes into account the emotional state of visually impaired individuals and foreign tourists while they are traveling. The system comprises an external device, a server, a machine learning model, and an emotion analysis engine.
[0343] The user uses a portable external device to photograph braille guidance paths or matrix barcodes. The captured image data is immediately transmitted to a server via the network. The server receives this image data and analyzes it using a machine learning model. This analysis identifies the user's current location and surrounding geographical information, which then serves as the basis for generating route guidance information.
[0344] Furthermore, an emotion analysis engine installed in an external device analyzes the user's voice and facial expressions in real time to understand their emotional state. This information is used to customize navigation. Specifically, if the user is feeling stressed, the system adjusts the route guidance information to select a quieter and safer route.
[0345] The server generates optimal route guidance information that takes emotional states into account and transmits it to an external device as voice or display data. The external device uses speech synthesis technology or a display to provide the user with tailored navigation information. For example, if the user gets lost in a busy environment, it might display a message such as, "Let's try taking a quieter route," to provide reassuring guidance.
[0346] An example of a specific prompt for a generative AI model would be something like, "Analyze the user's current location and suggest the optimal route based on their emotional state." This prompt functions as an instruction for the AI model to generate an appropriate path.
[0347] In this way, the present invention realizes a navigation system that helps users reach their destination safely and comfortably.
[0348] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0349] Step 1:
[0350] The user uses an external camera to photograph braille guidance paths and matrix barcodes. The resulting input is visual data. This data is transmitted to a server via the network. The shooting process involves using a smartphone camera app to acquire visual data with clear focus.
[0351] Step 2:
[0352] The server utilizes a machine learning model to analyze the received visual data. The input is visual data, and the output is the user's current location and environmental information. In this process, the AI model reads matrix barcodes in the visual data and retrieves the corresponding location information from the database.
[0353] Step 3:
[0354] The device analyzes the user's voice and facial expressions in real time using an emotion analysis engine. Input is voice and video data, and output is the user's emotional state. Specifically, the device uses a microphone and camera to collect data and transmits it to the emotion analysis engine.
[0355] Step 4:
[0356] The server generates optimal route guidance information based on emotional state and geographical information. The inputs are the current location, environmental information, and the user's emotional state, and the output is the adjusted route guidance information. In this step, the server calculates and selects a quiet and safe route based on the emotional state.
[0357] Step 5:
[0358] The terminal provides the user with route guidance information received from the server as audio or display data. The input is the adjusted route guidance information, and the output is audio guidance or screen display to aid the user's understanding. As a concrete example, the terminal might output a voice message through its speaker saying, "Let's choose a slightly quieter route."
[0359] (Application Example 2)
[0360] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0361] There is a lack of safer and more comfortable route guidance for visually impaired people and foreign tourists navigating shopping malls and large stores. Furthermore, navigation that takes into account the user's emotional state is not provided, failing to reduce user stress and improve safety.
[0362] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0363] In this invention, the server includes means for receiving visual information acquired by an external imaging device, means for using an artificial intelligence algorithm to analyze the visual information in order to generate route guidance information, and means for analyzing the user's emotional state from their voice or facial expressions. This enables customized route guidance that takes into account the user's emotional state.
[0364] An "external device" is a terminal used to acquire visual information and transmit captured images and encoded data to a server.
[0365] "Visual information" refers to image data and encoded data acquired by external devices, and forms the basis for location information and path analysis.
[0366] An "artificial intelligence algorithm" is a computational method used to analyze acquired visual information and generate appropriate path guidance information.
[0367] "Emotional state" refers to the psychological state of a user as determined from their voice or facial expressions, and is an important element in customizing navigation.
[0368] "Route guidance information" refers to information that includes directions and routes for movement, generated based on analyzed visual information.
[0369] This invention provides a navigation system that takes emotional states into account, enabling users to move around shopping malls and large stores with peace of mind. The system comprises external devices, a server, an artificial intelligence algorithm, and an emotion analysis engine.
[0370] The server receives visual information acquired by an external device's vision sensor. Then, using an artificial intelligence algorithm, it analyzes this visual information to determine the user's location. Next, an emotion analysis engine analyzes the user's emotional state from their real-time voice and facial expressions. Based on this analysis, optimal route guidance information is generated.
[0371] External devices can be smart glasses or mobile terminals that acquire visual information and instantly transmit it to the server. Furthermore, the route guidance received by the user is perceived through voice and screen display. This allows the user to enjoy customized navigation tailored to their emotional state.
[0372] For example, if a user attempts to pass through a crowded area in a large facility, and the emotion analysis engine detects tension, the server calculates a quieter route or one that avoids the crowd and guides the user through an external device.
[0373] An example of a prompt message would be, "We want to develop a navigation system for smart glasses that analyzes the user's emotions and guides them along a safe and comfortable route."
[0374] This invention can be said to be a very useful system for visually impaired people and foreign travelers by providing flexible navigation that takes into account the user's emotional state.
[0375] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0376] Step 1:
[0377] The device uses a visual sensor to acquire visual information about the user's surroundings. Inputs include image data and encoded data (e.g., QR codes). This data is processed on the device and immediately sent to the server.
[0378] Step 2:
[0379] The server inputs the received visual information into an artificial intelligence algorithm. Based on this data, location information and surrounding environment are identified. The AI algorithm performs image analysis and generates output such as the user's current location and surrounding geographical information.
[0380] Step 3:
[0381] The device analyzes the user's voice and facial expressions using an emotion analysis engine. Biometric data (voice and facial expressions) serves as input. This analysis identifies the user's real-time emotional state and generates psychological state data as output.
[0382] Step 4:
[0383] The server combines acquired location information and emotional state to generate optimal route guidance information. This involves data calculations using a generative AI model, resulting in a customized travel route that takes the user's emotional state into consideration.
[0384] Step 5:
[0385] The terminal receives route guidance information transmitted from the server and provides guidance to the user through voice and screen displays. By receiving visually and audibly adjusted information, the user can travel to their destination with peace of mind.
[0386] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0387] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0388] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0389] [Third Embodiment]
[0390] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0391] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0392] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0393] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0394] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0395] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0396] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0397] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0398] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0399] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0400] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0401] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0402] This invention provides a navigation system for visually impaired people and foreign tourists to travel safely and efficiently, and mainly consists of an external device equipped with a camera, a server, and an artificial intelligence model.
[0403] First, the user uses an external device to acquire images of braille guidance paths and QR codes installed in the environment. This allows the actual movement environment information to be transmitted to the server as digital data.
[0404] The server utilizes an artificial intelligence model to analyze the received image data. This model is built on machine learning techniques and has the ability to accurately recognize braille guidance paths and QR codes. Based on the analysis results, the server calculates the optimal route from the user's current location to their destination and generates route guidance information.
[0405] The terminal receives route guidance information sent back from the server and provides it to the user as audio or screen display information. Audio guidance is primarily used for visually impaired users, while screen displays and audio guidance are provided to foreign travelers using a multi-language selection function. This system allows users to travel safely without relying on surrounding visual information.
[0406] As a concrete example, if a visually impaired person arrives at a train station in a city they are visiting for the first time, they can use their smartphone camera to photograph the tactile guidance path at their feet. This image is sent to a server, which analyzes it and calculates a safe route from the platform to the ticket gate. Then, the user receives instructions via voice message from their device, such as "Please go left. Turn right after 10 meters." This allows the visually impaired person to navigate independently.
[0407] The following describes the processing flow.
[0408] Step 1:
[0409] The user activates the camera on an external device and takes pictures of the tactile guidance paths on the ground and the installed QR codes. This image data is saved on the device.
[0410] Step 2:
[0411] The device immediately sends the acquired image data to the server. If a QR code is included, that data is also sent at the same time.
[0412] Step 3:
[0413] The server receives the transmitted image data and passes it to an artificial intelligence model. The model analyzes the braille guidance paths and QR codes to determine the current user's location.
[0414] Step 4:
[0415] The server calculates the optimal route to the destination based on the analysis results and generates route guidance information. This information includes the direction to travel and details about obstacles to be aware of.
[0416] Step 5:
[0417] The server sends the generated routing information back to the terminal.
[0418] Step 6:
[0419] The terminal provides the user with received route guidance information as voice guidance or on-screen display in the user's selected language. The terminal updates the guidance content according to the direction the user is heading.
[0420] Step 7:
[0421] The user travels to their destination according to the provided directions. They can also take additional photos as needed to ensure the accuracy of the directions.
[0422] (Example 1)
[0423] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0424] Navigating safely and efficiently in new environments is often challenging for visually impaired individuals and foreign visitors, due to their inability to rely on visual information. In particular, the lack of methods for reaching destinations without relying on visual guidance makes independent movement difficult. This project aims to address these issues and enable more people to navigate their surroundings with confidence.
[0425] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0426] In this invention, the server includes means for receiving acquired visual data, means for using a machine learning model to generate movement information based on the visual data, and means for transmitting the movement information to an external device as an acoustic or visual output. This enables visually impaired persons and foreign visitors to safely navigate the best route to their destination without relying on visual information.
[0427] "Acquired visual data" refers to digital data containing visual information acquired using external devices.
[0428] A "machine learning model for generating movement information based on visual data" is an algorithm that receives visual data as input, extracts the information necessary for movement based on the model's learning, and generates appropriate movement instructions.
[0429] "Means for transmitting movement information to external devices as acoustic or visual output" refers to technologies for providing the generated movement-related information to the user either as audio output or displayed on a screen such as a display.
[0430] The "best route" refers to the most efficient, safe, and accessible path for a user attempting to travel.
[0431] "External devices" refer to all devices that function as part of a system and are used for acquiring visual data, displaying movement information, and outputting audio.
[0432] This invention is a navigation system aimed at enabling visually impaired individuals and foreign visitors to move safely and efficiently. The system consists of external devices, a server, and a generated AI model.
[0433] Users use their smartphones or dedicated camera equipment to photograph the surrounding braille guidance paths and QR codes. The acquired visual data is transmitted to a server via Wi-Fi or a mobile network.
[0434] The server processes the received visual data using a generative AI model to analyze it. This AI model takes visual data as input and uses machine learning techniques to identify braille pathways and QR codes. Specific instructions for the model are given using prompts. For example, a prompt such as "Identify the pathways and QR codes contained in this image" might be input.
[0435] After analysis, the server calculates the optimal route from the user's current location to their destination and generates route guidance information. This guidance information includes a safe and efficient path for travel.
[0436] The terminal receives routing information from the server and provides guidance to the user via voice or screen display. For visually impaired users, speech synthesis technology is used to deliver the guidance messages. For foreign visitors, screen displays can also be provided in multiple languages.
[0437] As a concrete example, when a visually impaired person uses public transportation in a city they are visiting for the first time, they use their smartphone to photograph the tactile guidance paths within the station. This image data is sent to a server, where an AI model analyzes it and calculates a safe route. Then, the device provides specific instructions via voice, such as "Go straight. Turn right at the next corner."
[0438] Thus, based on the system of the present invention, it becomes possible to move autonomously and safely.
[0439] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0440] Step 1:
[0441] The user uses a smartphone as an external device to photograph surrounding braille guidance paths and QR codes. Visual data acquired from the real-world environment is obtained as input. In this operation, the user launches a camera app, frames the necessary objects, and takes a photograph. Visual data (image file) is generated as output.
[0442] Step 2:
[0443] The device transmits captured visual data to the server. The input is an image file stored on the device. The device then packets this data over the network and transmits it. The output is the digital image data sent to the server.
[0444] Step 3:
[0445] The server activates a generative AI model to analyze the received visual data. The input is the visual data received by the server. The server prompts the AI model with the command, "Identify the guidance paths and QR codes contained in this image," and the model analyzes the image. This analysis outputs information identifying the locations of the braille guidance paths and QR codes.
[0446] Step 4:
[0447] The server calculates the optimal route from the user's current location to their destination based on information obtained from the AI model. The input consists of environmental information obtained through analysis and the user's destination information. Using this data, the server applies a route calculation algorithm and outputs a safe and efficient travel route.
[0448] Step 5:
[0449] The server sends the generated routing information to the terminal. The input is the calculated routing information. The server encrypts this information and transmits it to the terminal over the network. The output is the routing information that reaches the terminal.
[0450] Step 6:
[0451] The terminal analyzes the received route guidance information and provides directions to the user. The input is route guidance information received from the server. The terminal uses speech synthesis technology to output specific instructions in voice, such as "Turn right. There is an elevator 50 meters ahead." It then initiates the guidance process, allowing the user to begin their journey to their destination.
[0452] (Application Example 1)
[0453] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0454] There is a need to alleviate the difficulties that visually impaired people and foreign tourists face in safely and smoothly reaching their destinations in commercial facilities and complex environments, and to provide more accurate and multilingual navigation. In particular, there is a need for technology that can efficiently provide route guidance without relying on visual information.
[0455] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0456] In this invention, the server includes means for receiving visual data acquired by the imaging function of an external device, means for using a machine learning model to analyze the visual data in order to generate route guidance information, means for transmitting the route guidance information to the external device as audio or visual information, and means for providing multilingual guidance to the user on a display terminal device. As a result, the user can move to their destination safely and efficiently without relying on visual guidance.
[0457] An "external device" is a device that is portable to the user and has the function of acquiring and displaying visual data.
[0458] "Shooting function" refers to a function that uses a camera or similar sensor to acquire images or videos.
[0459] "Visual data" refers to image and video information acquired by cameras and sensors.
[0460] "Means of receiving" refers to the function that allows a server to receive data transmitted from an external device.
[0461] A "machine learning model" is an algorithm trained to analyze visual data and extract meaningful information.
[0462] "Means of analysis" refers to the function of analyzing acquired visual data and extracting necessary information.
[0463] "Route guidance information" refers to instructions and navigation information that helps users reach their destination.
[0464] "Means of transmission" refers to the function of sending route guidance information as audio or visual information to an external device and presenting it to the user.
[0465] A "display terminal device" is a device used to provide audio or visual information to a user.
[0466] "Multilingual guidance" refers to a function that provides route guidance information in the language selected by the user.
[0467] The system for carrying out the invention comprises an external device, a server, and a display terminal device. The program for this system begins by acquiring visual data using the camera function of an external device carried by the user. The visual data is image data including identification codes and braille guidance paths.
[0468] The server uses a machine learning model to analyze the received visual data. This model extracts the user's location information from the visual data and generates optimal route guidance information to the destination. Specifically, a machine learning framework such as TensorFlow is used. The server also performs language processing to provide this route guidance information in multiple languages, translating it into the user's preferred language.
[0469] The terminal receives route guidance information transmitted from the server and presents it to the user as audio or visual information. Display terminal devices include smart glasses and smartphones, which display multilingual guidance. Libraries such as Zebra Crossing are used for QR code recognition.
[0470] To give a specific example, a visually impaired person can use smart glasses in a shopping mall to scan a QR code at the entrance. At this time, the server analyzes the route to the "food section" in real time and provides voice guidance such as "Go 10 meters and turn right."
[0471] An example of a prompt sentence to input into the generating AI model is, "Please write about a route guidance system that provides audio guidance to help visually impaired people move safely within a shopping mall."
[0472] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0473] Step 1:
[0474] The user captures the surrounding environment using the camera function of an external device. The input is image data acquired from the external device, and the output is the captured image being saved as data on the external device. In this process, the camera mounted on the external device captures the image and secures the necessary visual data.
[0475] Step 2:
[0476] The terminal sends the captured image data to the server. The input is image data stored on an external device, and the output is image data transferred to the server. This transmission operation uses data communication functions to deliver the data to the server quickly and securely.
[0477] Step 3:
[0478] The server analyzes the received image data using a machine learning model. The input is the image data sent to the server, and the output is the analyzed location and identification information. The server uses the TensorFlow library to recognize identification codes and braille guidance paths within the image and determine the user's current location.
[0479] Step 4:
[0480] The server generates optimal route guidance information to the destination based on the analysis results. The input is the analyzed location information, and the output is route guidance information. The server uses a generation AI model to generate detailed directions, such as "Go straight to your destination and turn right after 10 meters."
[0481] Step 5:
[0482] The server converts the generated route guidance information into multiple languages and sends it to the terminal. The input is route guidance information, and the output is multilingual guidance information. In this process, natural language processing technology is used to convert the guidance into a language that is easy for the user to understand.
[0483] Step 6:
[0484] The terminal provides the user with received route guidance information as audio or visual information. The input is multilingual guidance information sent from the server, and the output is user-recognizable audio or display. The terminal uses speech synthesis technology or display functionality to convey instructions to the user.
[0485] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0486] The present invention provides a navigation system that takes into account the emotional state of the user when visually impaired people or foreign tourists are traveling, and comprises an external device, a server, an artificial intelligence model, and an emotion engine.
[0487] The user uses an external device to photograph braille guidance paths or QR codes. This image data is immediately transmitted to the server. The server analyzes the received image data using an artificial intelligence model to identify the user's current location and surrounding geographical information. This analysis result serves as the basic data for generating route guidance information.
[0488] The emotion engine analyzes the user's emotional state in real time from their voice and facial expressions using an external device. The analyzed emotional information is used to customize navigation. For example, if the system determines that the user is stressed, the route is adjusted to select a quieter and safer path.
[0489] The server generates optimal route guidance information that also takes into account the user's emotional state and transmits it to an external device. The external device then provides this information to the user as voice guidance or screen display. In this process, the guidance is adjusted according to the user's emotional state to provide more reassuring guidance.
[0490] As a concrete example, when a user gets lost in an area with many signs and a lot of noise, the system senses the user's stress. The emotion engine selects reassuring voices and relaxing routes for the user, and the server calculates a travel route based on this. The terminal then provides a tailored message such as, "Let's try taking a quieter route." In this way, the system provides support to help users reach their destination with peace of mind.
[0491] The following describes the processing flow.
[0492] Step 1:
[0493] The user uses an external device's camera to photograph surrounding braille pathways and QR codes. This data is temporarily stored on the device.
[0494] Step 2:
[0495] The device sends the captured image data to the server. Because this transmission requires real-time processing, the communication is optimized.
[0496] Step 3:
[0497] The server passes the received image data to an artificial intelligence model, which analyzes the braille guidance paths and QR codes. The model then uses this to determine the user's current location and surrounding environment.
[0498] Step 4:
[0499] The device transmits data obtained from the user's voice input and facial recognition via the camera to the emotion engine. The emotion engine analyzes this data and estimates the user's emotional state.
[0500] Step 5:
[0501] The server generates optimal route guidance information by considering the analyzed location information and the user's emotional state obtained from the emotion engine. If the emotional state indicates stress, it prioritizes selecting a safer route.
[0502] Step 6:
[0503] The server sends the generated routing information back to the terminal.
[0504] Step 7:
[0505] The terminal provides the user with route guidance information received from the server as voice guidance or screen display. The guidance reflects the user's emotional state, employing relaxing voices and designs if necessary.
[0506] Step 8:
[0507] The user follows the provided instructions. If further shooting or emotion recognition is required, they return to step 1 as needed and repeat the process.
[0508] (Example 2)
[0509] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0510] When visually impaired individuals and foreign tourists navigate unfamiliar environments, there is a need for navigation that not only provides location information and routes, but also takes their emotional state into consideration to offer a more reassuring experience. However, current technology lacks navigation systems that consider the user's emotional state, which can increase the user's mental burden. Furthermore, real-time emotion analysis and the mechanism for reflecting it in navigation are complex and difficult to implement.
[0511] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0512] In this invention, the server includes means for receiving image data acquired by an external device's imaging device, means for using a machine learning model to analyze the image data in order to generate route guidance information, means for transmitting the route guidance information to the external device as voice or display data, means for using an emotion analysis engine in the external device to analyze the user's emotional state from their voice and facial expressions, and means for adjusting the route guidance information based on the emotional state. This makes it possible to provide safe and secure routes that are tailored to the user's emotional state.
[0513] An "external device" is a device that a user can carry with them and is used to acquire image data and analyze emotional states.
[0514] A "photography device" is a component mounted on an external device that acquires image data.
[0515] "Image data" refers to data containing visual information collected using a camera or imaging device.
[0516] "Means of receiving" refers to the technology used by a server to receive image data transmitted from an external device.
[0517] A "machine learning model for analysis" is an algorithm used to extract information from image data and generate path guidance information.
[0518] "Route guidance information" refers to information that includes recommended routes and methods for users to reach their destination.
[0519] "Voice or display data" refers to a format in which route guidance information is provided to the user, including voice guidance and visual displays.
[0520] "Means of transmission" refers to the technology for sending analyzed routing information to an external device and providing it to the user.
[0521] An "emotion analysis engine" is a system that identifies a user's emotional state in real time from their voice and facial expressions.
[0522] "Emotional state" refers to the user's mental and emotional state, including stress levels and feelings of security.
[0523] "Means of adjustment" refer to processes and techniques for optimizing path guidance information based on analyzed emotional states.
[0524] A "matrix barcode" is a two-dimensional code used to visually encode information, including formats such as QR codes.
[0525] This invention provides a system that offers navigation that takes into account the emotional state of visually impaired individuals and foreign tourists while they are traveling. The system comprises an external device, a server, a machine learning model, and an emotion analysis engine.
[0526] The user uses a portable external device to photograph braille guidance paths or matrix barcodes. The captured image data is immediately transmitted to a server via the network. The server receives this image data and analyzes it using a machine learning model. This analysis identifies the user's current location and surrounding geographical information, which then serves as the basis for generating route guidance information.
[0527] Furthermore, an emotion analysis engine installed in an external device analyzes the user's voice and facial expressions in real time to understand their emotional state. This information is used to customize navigation. Specifically, if the user is feeling stressed, the system adjusts the route guidance information to select a quieter and safer route.
[0528] The server generates optimal route guidance information that takes emotional states into account and transmits it to an external device as voice or display data. The external device uses speech synthesis technology or a display to provide the user with tailored navigation information. For example, if the user gets lost in a busy environment, it might display a message such as, "Let's try taking a quieter route," to provide reassuring guidance.
[0529] An example of a specific prompt for a generative AI model would be something like, "Analyze the user's current location and suggest the optimal route based on their emotional state." This prompt functions as an instruction for the AI model to generate an appropriate path.
[0530] In this way, the present invention realizes a navigation system that helps users reach their destination safely and comfortably.
[0531] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0532] Step 1:
[0533] The user uses an external camera to photograph braille guidance paths and matrix barcodes. The resulting input is visual data. This data is transmitted to a server via the network. The shooting process involves using a smartphone camera app to acquire visual data with clear focus.
[0534] Step 2:
[0535] The server utilizes a machine learning model to analyze the received visual data. The input is visual data, and the output is the user's current location and environmental information. In this process, the AI model reads matrix barcodes in the visual data and retrieves the corresponding location information from the database.
[0536] Step 3:
[0537] The device analyzes the user's voice and facial expressions in real time using an emotion analysis engine. Input is voice and video data, and output is the user's emotional state. Specifically, the device uses a microphone and camera to collect data and transmits it to the emotion analysis engine.
[0538] Step 4:
[0539] The server generates optimal route guidance information based on emotional state and geographical information. The inputs are the current location, environmental information, and the user's emotional state, and the output is the adjusted route guidance information. In this step, the server calculates and selects a quiet and safe route based on the emotional state.
[0540] Step 5:
[0541] The terminal provides the user with route guidance information received from the server as audio or display data. The input is the adjusted route guidance information, and the output is audio guidance or screen display to aid the user's understanding. As a concrete example, the terminal might output a voice message through its speaker saying, "Let's choose a slightly quieter route."
[0542] (Application Example 2)
[0543] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0544] There is a lack of safer and more comfortable route guidance for visually impaired people and foreign tourists navigating shopping malls and large stores. Furthermore, navigation that takes into account the user's emotional state is not provided, failing to reduce user stress and improve safety.
[0545] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0546] In this invention, the server includes means for receiving visual information acquired by an external imaging device, means for using an artificial intelligence algorithm to analyze the visual information in order to generate route guidance information, and means for analyzing the user's emotional state from their voice or facial expressions. This enables customized route guidance that takes into account the user's emotional state.
[0547] An "external device" is a terminal used to acquire visual information and transmit captured images and encoded data to a server.
[0548] "Visual information" refers to image data and encoded data acquired by external devices, and forms the basis for location information and path analysis.
[0549] An "artificial intelligence algorithm" is a computational method used to analyze acquired visual information and generate appropriate path guidance information.
[0550] "Emotional state" refers to the psychological state of a user, as determined from their voice or facial expressions, and is an important element in customizing navigation.
[0551] "Route guidance information" refers to information that includes directions and routes for movement, generated based on analyzed visual information.
[0552] This invention provides a navigation system that takes emotional states into account, enabling users to move around shopping malls and large stores with peace of mind. The system comprises external devices, a server, an artificial intelligence algorithm, and an emotion analysis engine.
[0553] The server receives visual information acquired by an external device's vision sensor. Then, using an artificial intelligence algorithm, it analyzes this visual information to determine the user's location. Next, an emotion analysis engine analyzes the user's emotional state from their real-time voice and facial expressions. Based on this analysis, optimal route guidance information is generated.
[0554] The external device can be smart glasses or a mobile terminal, which acquires visual information and instantly transmits it to the server. Furthermore, the route guidance received by the user is perceived through voice and screen display. This allows the user to enjoy customized navigation tailored to their emotional state.
[0555] For example, if a user attempts to pass through a crowded area in a large facility, and the emotion analysis engine detects tension, the server calculates a quieter route or one that avoids the crowd and guides the user through an external device.
[0556] An example of a prompt message would be, "We want to develop a navigation system for smart glasses that analyzes the user's emotions and guides them along a safe and comfortable route."
[0557] This invention can be said to be a very useful system for visually impaired people and foreign travelers by providing flexible navigation that takes into account the user's emotional state.
[0558] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0559] Step 1:
[0560] The device uses a visual sensor to acquire visual information about the user's surroundings. Inputs include image data and encoded data (e.g., QR codes). This data is processed on the device and immediately sent to the server.
[0561] Step 2:
[0562] The server inputs the received visual information into an artificial intelligence algorithm. Based on this data, location information and surrounding environment are identified. The AI algorithm performs image analysis and generates output such as the user's current location and surrounding geographical information.
[0563] Step 3:
[0564] The device analyzes the user's voice and facial expressions using an emotion analysis engine. Biometric data (voice and facial expressions) serves as input. This analysis identifies the user's real-time emotional state and generates psychological state data as output.
[0565] Step 4:
[0566] The server combines acquired location information and emotional state to generate optimal route guidance information. This involves data calculations using a generative AI model, resulting in a customized travel route that takes the user's emotional state into consideration.
[0567] Step 5:
[0568] The terminal receives route guidance information transmitted from the server and provides guidance to the user through voice and screen displays. By receiving visually and audibly adjusted information, the user can travel to their destination with peace of mind.
[0569] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0570] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0571] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0572] [Fourth Embodiment]
[0573] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0574] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0575] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0576] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0577] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0578] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0579] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0580] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0581] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0582] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0583] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0584] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0585] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0586] This invention provides a navigation system for visually impaired people and foreign tourists to travel safely and efficiently, and mainly consists of an external device equipped with a camera, a server, and an artificial intelligence model.
[0587] First, the user uses an external device to acquire images of braille guidance paths and QR codes installed in the environment. This allows the actual movement environment information to be transmitted to the server as digital data.
[0588] The server utilizes an artificial intelligence model to analyze the received image data. This model is built on machine learning techniques and has the ability to accurately recognize braille guidance paths and QR codes. Based on the analysis results, the server calculates the optimal route from the user's current location to their destination and generates route guidance information.
[0589] The terminal receives route guidance information sent back from the server and provides it to the user as audio or screen display information. Audio guidance is primarily used for visually impaired users, while screen displays and audio guidance are provided to foreign travelers using a multi-language selection function. This system allows users to travel safely without relying on surrounding visual information.
[0590] As a concrete example, if a visually impaired person arrives at a train station in a city they are visiting for the first time, they can use their smartphone camera to photograph the tactile guidance path at their feet. This image is sent to a server, which analyzes it and calculates a safe route from the platform to the ticket gate. Then, the user receives instructions via voice message from their device, such as "Please go left. Turn right after 10 meters." This allows the visually impaired person to navigate independently.
[0591] The following describes the processing flow.
[0592] Step 1:
[0593] The user activates the camera on an external device and takes pictures of the tactile guidance paths on the ground and the installed QR codes. This image data is saved on the device.
[0594] Step 2:
[0595] The device immediately sends the acquired image data to the server. If a QR code is included, that data is also sent at the same time.
[0596] Step 3:
[0597] The server receives the transmitted image data and passes it to an artificial intelligence model. The model analyzes the braille guidance paths and QR codes to determine the current user's location.
[0598] Step 4:
[0599] The server calculates the optimal route to the destination based on the analysis results and generates route guidance information. This information includes the direction to travel and details about obstacles to be aware of.
[0600] Step 5:
[0601] The server sends the generated routing information back to the terminal.
[0602] Step 6:
[0603] The terminal provides the user with received route guidance information as voice guidance or on-screen display in the user's selected language. The terminal updates the guidance content according to the direction the user is heading.
[0604] Step 7:
[0605] The user travels to their destination according to the provided directions. They can also take additional photos as needed to ensure the accuracy of the directions.
[0606] (Example 1)
[0607] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0608] Navigating safely and efficiently in new environments is often challenging for visually impaired individuals and foreign visitors, due to their inability to rely on visual information. In particular, the lack of methods for reaching destinations without relying on visual guidance makes independent movement difficult. This project aims to address these issues and enable more people to navigate their surroundings with confidence.
[0609] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0610] In this invention, the server includes means for receiving acquired visual data, means for using a machine learning model to generate movement information based on the visual data, and means for transmitting the movement information to an external device as an acoustic or visual output. This enables visually impaired persons and foreign visitors to safely navigate the best route to their destination without relying on visual information.
[0611] "Acquired visual data" refers to digital data containing visual information acquired using external devices.
[0612] A "machine learning model for generating movement information based on visual data" is an algorithm that receives visual data as input, extracts the information necessary for movement based on the model's learning, and generates appropriate movement instructions.
[0613] "Means for transmitting movement information to external devices as acoustic or visual output" refers to technologies for providing the generated movement-related information to the user either as audio output or displayed on a screen such as a display.
[0614] The "best route" refers to the most efficient, safe, and accessible path for a user attempting to travel.
[0615] "External devices" refer to all devices that function as part of a system and are used for acquiring visual data, displaying movement information, and outputting audio.
[0616] This invention is a navigation system aimed at enabling visually impaired individuals and foreign visitors to move safely and efficiently. The system consists of external devices, a server, and a generated AI model.
[0617] Users use their smartphones or dedicated camera equipment to photograph the surrounding braille guidance paths and QR codes. The acquired visual data is transmitted to a server via Wi-Fi or a mobile network.
[0618] The server processes the received visual data using a generative AI model to analyze it. This AI model takes visual data as input and uses machine learning techniques to identify braille pathways and QR codes. Specific instructions for the model are given using prompts. For example, a prompt such as "Identify the pathways and QR codes contained in this image" might be input.
[0619] After analysis, the server calculates the optimal route from the user's current location to their destination and generates route guidance information. This guidance information includes a safe and efficient path for travel.
[0620] The terminal receives routing information from the server and provides guidance to the user via voice or screen display. For visually impaired users, speech synthesis technology is used to deliver the guidance messages. For foreign visitors, screen displays can also be provided in multiple languages.
[0621] As a concrete example, when a visually impaired person uses public transportation in a city they are visiting for the first time, they use their smartphone to photograph the tactile guidance paths within the station. This image data is sent to a server, where an AI model analyzes it and calculates a safe route. Then, the device provides specific instructions via voice, such as "Go straight. Turn right at the next corner."
[0622] Thus, based on the system of the present invention, it becomes possible to move autonomously and safely.
[0623] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0624] Step 1:
[0625] The user uses a smartphone as an external device to photograph surrounding braille guidance paths and QR codes. Visual data acquired from the real-world environment is obtained as input. In this operation, the user launches a camera app, frames the necessary objects, and takes a photograph. Visual data (image file) is generated as output.
[0626] Step 2:
[0627] The device transmits captured visual data to the server. The input is an image file stored on the device. The device then packets this data over the network and transmits it. The output is the digital image data sent to the server.
[0628] Step 3:
[0629] The server activates a generative AI model to analyze the received visual data. The input is the visual data received by the server. The server prompts the AI model with the command, "Identify the guidance paths and QR codes contained in this image," and the model analyzes the image. This analysis outputs information identifying the locations of the braille guidance paths and QR codes.
[0630] Step 4:
[0631] The server calculates the optimal route from the user's current location to their destination based on information obtained from the AI model. The input consists of environmental information obtained through analysis and the user's destination information. Using this data, the server applies a route calculation algorithm and outputs a safe and efficient travel route.
[0632] Step 5:
[0633] The server sends the generated routing information to the terminal. The input is the calculated routing information. The server encrypts this information and transmits it to the terminal over the network. The output is the routing information that reaches the terminal.
[0634] Step 6:
[0635] The terminal analyzes the received route guidance information and provides directions to the user. The input is route guidance information received from the server. The terminal uses speech synthesis technology to output specific instructions in voice, such as "Turn right. There is an elevator 50 meters ahead." It then initiates the guidance process, allowing the user to begin their journey to their destination.
[0636] (Application Example 1)
[0637] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0638] There is a need to alleviate the difficulties that visually impaired people and foreign tourists face in safely and smoothly reaching their destinations in commercial facilities and complex environments, and to provide more accurate and multilingual navigation. In particular, there is a need for technology that can efficiently provide route guidance without relying on visual information.
[0639] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0640] In this invention, the server includes means for receiving visual data acquired by the imaging function of an external device, means for using a machine learning model to analyze the visual data in order to generate route guidance information, means for transmitting the route guidance information to the external device as audio or visual information, and means for providing multilingual guidance to the user on a display terminal device. As a result, the user can move to their destination safely and efficiently without relying on visual guidance.
[0641] An "external device" is a device that is portable to the user and has the function of acquiring and displaying visual data.
[0642] "Shooting function" refers to a function that uses a camera or similar sensor to acquire images or videos.
[0643] "Visual data" refers to image and video information acquired by cameras and sensors.
[0644] "Means of receiving" refers to the function that allows a server to receive data transmitted from an external device.
[0645] A "machine learning model" is an algorithm trained to analyze visual data and extract meaningful information.
[0646] "Means of analysis" refers to the function of analyzing acquired visual data and extracting necessary information.
[0647] "Route guidance information" refers to instructions and navigation information that helps users reach their destination.
[0648] "Means of transmission" refers to the function of sending route guidance information as audio or visual information to an external device and presenting it to the user.
[0649] A "display terminal device" is a device used to provide audio or visual information to a user.
[0650] "Multilingual guidance" refers to a function that provides route guidance information in the language selected by the user.
[0651] The system for carrying out the invention comprises an external device, a server, and a display terminal device. The program for this system begins by acquiring visual data using the camera function of an external device carried by the user. The visual data is image data including identification codes and braille guidance paths.
[0652] The server uses a machine learning model to analyze the received visual data. This model extracts the user's location information from the visual data and generates optimal route guidance information to the destination. Specifically, a machine learning framework such as TensorFlow is used. The server also performs language processing to provide this route guidance information in multiple languages, translating it into the user's preferred language.
[0653] The terminal receives route guidance information transmitted from the server and presents it to the user as audio or visual information. Display terminal devices include smart glasses and smartphones, which display multilingual guidance. Libraries such as Zebra Crossing are used for QR code recognition.
[0654] To give a specific example, a visually impaired person can use smart glasses in a shopping mall to scan a QR code at the entrance. At this time, the server analyzes the route to the "food section" in real time and provides voice guidance such as "Go 10 meters and turn right."
[0655] An example of a prompt sentence to input into the generating AI model is, "Please write about a route guidance system that provides audio guidance to help visually impaired people move safely within a shopping mall."
[0656] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0657] Step 1:
[0658] The user captures the surrounding environment using the camera function of an external device. The input is image data acquired from the external device, and the output is the captured image being saved as data on the external device. In this process, the camera mounted on the external device captures the image and secures the necessary visual data.
[0659] Step 2:
[0660] The terminal sends the captured image data to the server. The input is image data stored on an external device, and the output is image data transferred to the server. This transmission operation uses data communication functions to deliver the data to the server quickly and securely.
[0661] Step 3:
[0662] The server analyzes the received image data using a machine learning model. The input is the image data sent to the server, and the output is the analyzed location and identification information. The server uses the TensorFlow library to recognize identification codes and braille guidance paths within the image and determine the user's current location.
[0663] Step 4:
[0664] The server generates optimal route guidance information to the destination based on the analysis results. The input is the analyzed location information, and the output is route guidance information. The server uses a generation AI model to generate detailed directions, such as "Go straight to your destination and turn right after 10 meters."
[0665] Step 5:
[0666] The server converts the generated route guidance information into multiple languages and sends it to the terminal. The input is route guidance information, and the output is multilingual guidance information. In this process, natural language processing technology is used to convert the guidance into a language that is easy for the user to understand.
[0667] Step 6:
[0668] The terminal provides the user with received route guidance information as audio or visual information. The input is multilingual guidance information sent from the server, and the output is user-recognizable audio or display. The terminal uses speech synthesis technology or display functionality to convey instructions to the user.
[0669] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0670] The present invention provides a navigation system that takes into account the emotional state of the user when visually impaired people or foreign tourists are traveling, and comprises an external device, a server, an artificial intelligence model, and an emotion engine.
[0671] The user uses an external device to photograph braille guidance paths or QR codes. This image data is immediately transmitted to the server. The server analyzes the received image data using an artificial intelligence model to identify the user's current location and surrounding geographical information. This analysis result serves as the basic data for generating route guidance information.
[0672] The emotion engine analyzes the user's emotional state in real time from their voice and facial expressions using an external device. The analyzed emotional information is used to customize navigation. For example, if the system determines that the user is stressed, the route is adjusted to select a quieter and safer path.
[0673] The server generates optimal route guidance information that also takes into account the user's emotional state and transmits it to an external device. The external device then provides this information to the user as voice guidance or screen display. In this process, the guidance is adjusted according to the user's emotional state to provide more reassuring guidance.
[0674] As a concrete example, when a user gets lost in an area with many signs and a lot of noise, the system senses the user's stress. The emotion engine selects reassuring voices and relaxing routes for the user, and the server calculates a travel route based on this. The terminal then provides a tailored message such as, "Let's try taking a quieter route." In this way, the system provides support to help users reach their destination with peace of mind.
[0675] The following describes the processing flow.
[0676] Step 1:
[0677] The user uses an external device's camera to photograph surrounding braille pathways and QR codes. This data is temporarily stored on the device.
[0678] Step 2:
[0679] The device sends the captured image data to the server. Because this transmission requires real-time processing, the communication is optimized.
[0680] Step 3:
[0681] The server passes the received image data to an artificial intelligence model, which analyzes the braille guidance paths and QR codes. The model then uses this to determine the user's current location and surrounding environment.
[0682] Step 4:
[0683] The device transmits data obtained from the user's voice input and facial recognition via the camera to the emotion engine. The emotion engine analyzes this data and estimates the user's emotional state.
[0684] Step 5:
[0685] The server generates optimal route guidance information by considering the analyzed location information and the user's emotional state obtained from the emotion engine. If the emotional state indicates stress, it prioritizes selecting a safer route.
[0686] Step 6:
[0687] The server sends the generated routing information back to the terminal.
[0688] Step 7:
[0689] The terminal provides the user with route guidance information received from the server as voice guidance or screen display. The guidance reflects the user's emotional state, employing relaxing voices and designs if necessary.
[0690] Step 8:
[0691] The user follows the provided instructions. If further shooting or emotion recognition is required, they return to step 1 as needed and repeat the process.
[0692] (Example 2)
[0693] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0694] When visually impaired individuals and foreign tourists navigate unfamiliar environments, there is a need for navigation that not only provides location information and routes, but also takes their emotional state into consideration to offer a more reassuring experience. However, current technology lacks navigation systems that consider the user's emotional state, which can increase the user's mental burden. Furthermore, real-time emotion analysis and the mechanism for reflecting it in navigation are complex and difficult to implement.
[0695] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0696] In this invention, the server includes means for receiving image data acquired by an external device's imaging device, means for using a machine learning model to analyze the image data in order to generate route guidance information, means for transmitting the route guidance information to the external device as voice or display data, means for using an emotion analysis engine in the external device to analyze the user's emotional state from their voice and facial expressions, and means for adjusting the route guidance information based on the emotional state. This makes it possible to provide safe and secure routes that are tailored to the user's emotional state.
[0697] An "external device" is a device that a user can carry with them and is used to acquire image data and analyze emotional states.
[0698] A "photography device" is a component mounted on an external device that acquires image data.
[0699] "Image data" refers to data containing visual information collected using a camera or imaging device.
[0700] "Means of receiving" refers to the technology used by a server to receive image data transmitted from an external device.
[0701] A "machine learning model for analysis" is an algorithm used to extract information from image data and generate path guidance information.
[0702] "Route guidance information" refers to information that includes recommended routes and methods for users to reach their destination.
[0703] "Voice or display data" refers to a format in which route guidance information is provided to the user, including voice guidance and visual displays.
[0704] "Means of transmission" refers to the technology for sending analyzed routing information to an external device and providing it to the user.
[0705] An "emotion analysis engine" is a system that identifies a user's emotional state in real time from their voice and facial expressions.
[0706] "Emotional state" refers to the user's mental and emotional state, including stress levels and feelings of security.
[0707] "Means of adjustment" refer to processes and techniques for optimizing path guidance information based on analyzed emotional states.
[0708] A "matrix barcode" is a two-dimensional code used to visually encode information, including formats such as QR codes.
[0709] This invention provides a system that offers navigation that takes into account the emotional state of visually impaired individuals and foreign tourists while they are traveling. The system comprises an external device, a server, a machine learning model, and an emotion analysis engine.
[0710] The user uses a portable external device to photograph braille guidance paths or matrix barcodes. The captured image data is immediately transmitted to a server via the network. The server receives this image data and analyzes it using a machine learning model. This analysis identifies the user's current location and surrounding geographical information, which then serves as the basis for generating route guidance information.
[0711] Furthermore, an emotion analysis engine installed in an external device analyzes the user's voice and facial expressions in real time to understand their emotional state. This information is used to customize navigation. Specifically, if the user is feeling stressed, the system adjusts the route guidance information to select a quieter and safer route.
[0712] The server generates optimal route guidance information that takes emotional states into account and transmits it to an external device as voice or display data. The external device uses speech synthesis technology or a display to provide the user with tailored navigation information. For example, if the user gets lost in a busy environment, it might display a message such as, "Let's try taking a quieter route," to provide reassuring guidance.
[0713] An example of a specific prompt for a generative AI model would be something like, "Analyze the user's current location and suggest the optimal route based on their emotional state." This prompt functions as an instruction for the AI model to generate an appropriate path.
[0714] In this way, the present invention realizes a navigation system that helps users reach their destination safely and comfortably.
[0715] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0716] Step 1:
[0717] The user uses an external camera to photograph braille guidance paths and matrix barcodes. The resulting input is visual data. This data is transmitted to a server via the network. The shooting process involves using a smartphone camera app to acquire visual data with clear focus.
[0718] Step 2:
[0719] The server utilizes a machine learning model to analyze the received visual data. The input is visual data, and the output is the user's current location and environmental information. In this process, the AI model reads matrix barcodes in the visual data and retrieves the corresponding location information from the database.
[0720] Step 3:
[0721] The device analyzes the user's voice and facial expressions in real time using an emotion analysis engine. Input is voice and video data, and output is the user's emotional state. Specifically, the device uses a microphone and camera to collect data and transmits it to the emotion analysis engine.
[0722] Step 4:
[0723] The server generates optimal route guidance information based on emotional state and geographical information. The inputs are the current location, environmental information, and the user's emotional state, and the output is the adjusted route guidance information. In this step, the server calculates and selects a quiet and safe route based on the emotional state.
[0724] Step 5:
[0725] The terminal provides the user with route guidance information received from the server as audio or display data. The input is the adjusted route guidance information, and the output is audio guidance or screen display to aid the user's understanding. As a concrete example, the terminal might output a voice message through its speaker saying, "Let's choose a slightly quieter route."
[0726] (Application Example 2)
[0727] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0728] There is a lack of safer and more comfortable route guidance for visually impaired people and foreign tourists navigating shopping malls and large stores. Furthermore, navigation that takes into account the user's emotional state is not provided, failing to reduce user stress and improve safety.
[0729] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0730] In this invention, the server includes means for receiving visual information acquired by an external imaging device, means for using an artificial intelligence algorithm to analyze the visual information in order to generate route guidance information, and means for analyzing the user's emotional state from their voice or facial expressions. This enables customized route guidance that takes into account the user's emotional state.
[0731] An "external device" is a terminal used to acquire visual information and transmit captured images and encoded data to a server.
[0732] "Visual information" refers to image data and encoded data acquired by external devices, and forms the basis for location information and path analysis.
[0733] An "artificial intelligence algorithm" is a computational method used to analyze acquired visual information and generate appropriate path guidance information.
[0734] "Emotional state" refers to the psychological state of a user, as determined from their voice or facial expressions, and is an important element in customizing navigation.
[0735] "Route guidance information" refers to information that includes directions and routes for movement, generated based on analyzed visual information.
[0736] This invention provides a navigation system that takes emotional states into account, enabling users to move around shopping malls and large stores with peace of mind. The system comprises external devices, a server, an artificial intelligence algorithm, and an emotion analysis engine.
[0737] The server receives visual information acquired by an external device's vision sensor. Then, using an artificial intelligence algorithm, it analyzes this visual information to determine the user's location. Next, an emotion analysis engine analyzes the user's emotional state from their real-time voice and facial expressions. Based on this analysis, optimal route guidance information is generated.
[0738] The external device can be smart glasses or a mobile terminal, which acquires visual information and instantly transmits it to the server. Furthermore, the route guidance received by the user is perceived through voice and screen display. This allows the user to enjoy customized navigation tailored to their emotional state.
[0739] For example, if a user attempts to pass through a crowded area in a large facility, and the emotion analysis engine detects tension, the server calculates a quieter route or one that avoids the crowd and guides the user through an external device.
[0740] An example of a prompt message would be, "We want to develop a navigation system for smart glasses that analyzes the user's emotions and guides them along a safe and comfortable route."
[0741] This invention can be said to be a very useful system for visually impaired people and foreign travelers by providing flexible navigation that takes into account the user's emotional state.
[0742] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0743] Step 1:
[0744] The device uses a visual sensor to acquire visual information about the user's surroundings. Inputs include image data and encoded data (e.g., QR codes). This data is processed on the device and immediately sent to the server.
[0745] Step 2:
[0746] The server inputs the received visual information into an artificial intelligence algorithm. Based on this data, location information and surrounding environment are identified. The AI algorithm performs image analysis and generates output such as the user's current location and surrounding geographical information.
[0747] Step 3:
[0748] The device analyzes the user's voice and facial expressions using an emotion analysis engine. Biometric data (voice and facial expressions) serves as input. This analysis identifies the user's real-time emotional state and generates psychological state data as output.
[0749] Step 4:
[0750] The server combines acquired location information and emotional state to generate optimal route guidance information. This involves data calculations using a generative AI model, resulting in a customized travel route that takes the user's emotional state into consideration.
[0751] Step 5:
[0752] The terminal receives route guidance information transmitted from the server and provides guidance to the user through voice and screen displays. By receiving visually and audibly adjusted information, the user can travel to their destination with peace of mind.
[0753] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0754] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0755] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0756] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0757] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0758] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0759] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0760] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0761] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0762] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0763] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0764] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0765] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0766] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0767] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0768] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0769] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0770] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0771] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0772] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0773] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0774] The following is further disclosed regarding the embodiments described above.
[0775] (Claim 1)
[0776] Means for receiving image data acquired by an external imaging device,
[0777] A means of using an artificial intelligence model to analyze the aforementioned image data in order to generate route instruction information,
[0778] Means for transmitting the aforementioned route instruction information to an external device as audio or display information,
[0779] A system that includes this.
[0780] (Claim 2)
[0781] The system according to claim 1, characterized in that the image data includes a braille guidance path.
[0782] (Claim 3)
[0783] The system according to claim 1, characterized in that the code data acquired by an external device is a QR code containing location information.
[0784] "Example 1"
[0785] (Claim 1)
[0786] A means for receiving acquired visual data,
[0787] A means of using a machine learning model to generate movement information based on the aforementioned visual data,
[0788] Means for transmitting the aforementioned movement information to an external device as an acoustic or visual output,
[0789] A means for calculating the best route from the user's current location to their destination,
[0790] Means for generating specific directional instructions based on the aforementioned route,
[0791] A system that includes this.
[0792] (Claim 2)
[0793] The system according to claim 1, characterized in that the visual data includes a tactile guidance path.
[0794] (Claim 3)
[0795] The system according to claim 1, characterized in that the identification data acquired by an external device is a two-dimensional code containing location information.
[0796] "Application Example 1"
[0797] (Claim 1)
[0798] A means for receiving visual data acquired by the shooting function of an external device,
[0799] A means of using a machine learning model to analyze the aforementioned visual data in order to generate route guidance information,
[0800] Means for transmitting the aforementioned route guidance information to an external device as audio or visual information,
[0801] A means of providing multilingual guidance to the user on a display terminal device,
[0802] A system that includes this.
[0803] (Claim 2)
[0804] The system according to claim 1, characterized in that the visual data includes a braille guided path, and the user is guided along the route by voice guidance from a display terminal device.
[0805] (Claim 3)
[0806] The system according to claim 1, characterized in that the code information acquired by an external device is an identification code that includes location information.
[0807] "Example 2 of combining an emotion engine"
[0808] (Claim 1)
[0809] Means for receiving image data acquired by an external imaging device,
[0810] A means of using a machine learning model to analyze the aforementioned image data in order to generate route instruction information,
[0811] Means for transmitting the aforementioned route instruction information to an external device as voice or display data,
[0812] The external device includes means for using an emotion analysis engine that analyzes the user's emotional state from their voice and facial expressions,
[0813] Means for adjusting route instruction information based on the aforementioned emotional state,
[0814] A system that includes this.
[0815] (Claim 2)
[0816] The system according to claim 1, characterized in that the image data includes a tactile guidance pathway.
[0817] (Claim 3)
[0818] The system according to claim 1, characterized in that the code data acquired by an external device is a matrix barcode containing location information.
[0819] "Application example 2 when combining with an emotional engine"
[0820] (Claim 1)
[0821] A means for receiving visual information acquired by an external imaging device,
[0822] A means of using an artificial intelligence algorithm to analyze the aforementioned visual information in order to generate route guidance information,
[0823] A means of analyzing the emotional state from the user's voice or facial expressions,
[0824] Means for customizing route guidance information according to the aforementioned emotional state,
[0825] Means for transmitting the customized route guidance information to an external device as audio or visual display,
[0826] A system that includes this.
[0827] (Claim 2)
[0828] The system according to claim 1, characterized in that the visual information includes tactile guidance.
[0829] (Claim 3)
[0830] The system according to claim 1, characterized in that the encoded data acquired by an external device is a code that includes location information. [Explanation of Symbols]
[0831] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. Means for receiving image data acquired by an external imaging device, A means of using an artificial intelligence model to analyze the aforementioned image data in order to generate route instruction information, Means for transmitting the aforementioned route instruction information to an external device as audio or display information, A system that includes this.
2. The system according to claim 1, characterized in that the image data includes a braille guidance path.
3. The system according to claim 1, characterized in that the code data acquired by an external device is a two-dimensional code containing location information.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A