System

The system addresses the lack of detailed real-time information for diverse tourists by using location-based information retrieval and translation to enhance the experience of visually impaired, hearing impaired, and foreign tourists at tourist sites.

JP2026017295APending Publication Date: 2026-02-04SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024118077
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2026-02-04

AI Technical Summary

Technical Problem

Existing tourist guide systems fail to provide detailed, real-time information to visually impaired, hearing impaired, and foreign tourists, limiting their experience and understanding of tourist sites and cultural facilities.

Method used

A system that acquires a user's location, searches for relevant tourist attraction information, generates and plays audio guides, provides vibration notifications, displays text information, and translates into multiple languages, using a user terminal, GPS module, communication means, and databases.

Benefits of technology

Enables visually impaired, hearing impaired, and foreign tourists to receive detailed, real-time information, enhancing their experience by providing audio and text guides and vibration notifications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026017295000001_ABST
    Figure 2026017295000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system, comprising: means for acquiring a current location of a user; means for searching for relevant scenic spot information from a scenic spot database based on the acquired location information; means for sending the searched scenic spot information to a user terminal; and means for generating and playing voice guidance based on the sent scenic spot information.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] The present invention aims to solve the problem of insufficient information provided to visitors when visiting tourist sites and cultural facilities. It is particularly important to provide information in an accessible format to the visually impaired, the hearing impaired, and foreign tourists who speak different languages. This aims to improve the quality of the tourist experience and provide a deeper understanding of the local culture and history. [Means for solving the problem]

[0005] The present invention solves the above-mentioned problems by providing a system including a means for acquiring a user's current location. Specifically, the system includes a means for searching for related tourist attraction information from a tourist attraction database based on the acquired location information. The system also includes a means for transmitting the searched tourist attraction information to a user terminal and a means for generating and playing an audio guide based on the transmitted tourist attraction information. Furthermore, the system includes a vibration function for notifying the user of tourist attraction information in real time based on the user's current location, and a means for displaying tourist attraction information in text format for the hearing impaired.

[0006] Understood. Now, let's create definitions for the important words included in the claims.

[0007] A "user" is a person who uses the system to obtain tourist destination information and use that information to enhance their tourist experience.

[0008] "Current location" refers to the latitude and longitude information of the user's current location obtained by GPS or other location measurement means.

[0009] "Means for obtaining" refers to devices or software for obtaining the user's current location information, including, for example, a GPS module or location service.

[0010] A "tourist destination database" is a database that stores detailed information on tourist destinations and cultural facilities, as well as their location information.

[0011] "Searching means" refers to an algorithm or search engine that searches for relevant tourist destination information from a tourist destination database based on the acquired location information.

[0012] "Tourist destination information" is detailed data related to tourist destinations, such as descriptions, historical background, and attractions.

[0013] The "transmitting means" includes hardware and software for transmitting tourist destination information from the server to the user terminal using a communication protocol.

[0014] A "user terminal" is a device for receiving tourist destination information and displaying or playing audio, and includes smartphones and smart glasses.

[0015] The "means for generating" refers to the voice synthesis technology and software for generating an audio guide based on the transmitted tourist destination information.

[0016] "Playback means" refers to a device with built-in speakers or headphones that allows the user to listen to the generated audio guide.

[0017] The "vibration function" is a function that causes the device to vibrate to notify the user, and is used especially when approaching important tourist spots.

[0018] "Text format" refers to textual information, a format used by hearing-impaired people to obtain information visually. [Brief explanation of the drawings]

[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6]FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0021] First, the terms used in the following description will be explained.

[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0027] [First embodiment]

[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0040] The present invention relates to an audio guide system that provides detailed information in real time when visiting tourist spots and cultural facilities. This system aims to provide convenience to people with visual or hearing impairments, as well as foreign tourists.

[0041] System Configuration

[0042] The system mainly consists of the following components:

[0043] 1. User device: A mobile device such as a smartphone or smart glasses.

[0044] 2. Server: A server containing a tourist destination database.

[0045] 3. GPS module: A location information acquisition device for obtaining the user's current location.

[0046] 4. Communication means: A network connection for sending and receiving data between the user terminal and the server.

[0047] 5. Audio guide generation function: Software for converting tourist information into audio guides.

[0048] 6. Vibration function: Notification of important tourist spots.

[0049] 7. Text display function: A means of displaying text information for the hearing impaired.

[0050] Program processing

[0051] 1. Obtaining the current location

[0052] The user terminal periodically uses the GPS module to obtain the user's current location.

[0053] The acquired location information (latitude and longitude) is stored in the device and sent to the server.

[0054] 2. Search for tourist information

[0055] The server analyzes the received location information and searches a database for information on the nearest tourist attractions and cultural facilities.

[0056] 3. Sending and receiving tourist destination information

[0057] The server transmits the searched tourist spot information to the user terminal.

[0058] The user terminal parses the received information and converts it into an appropriate format.

[0059] 4. Generating and Providing Audio Guides

[0060] The user terminal converts the received text information into audio guidance using a speech synthesis engine.

[0061] The audio guide is played through headphones in the smart glasses worn by the user.

[0062] 5. Vibration notifications

[0063] When the user approaches an important point in a tourist destination, the user's device will notify them using a vibration function.

[0064] 6. Providing text information

[0065] For the hearing impaired, the server sends tourist information in text format to the terminal.

[0066] The user terminal displays text information on the smart glasses display.

[0067] 7. Automatic translation function

[0068] For foreign visitors to Japan, the server automatically translates tourist information and generates audio guides in multiple languages.

[0069] The translated audio guide is sent to the user's terminal and played back in the specified language.

[0070] Specific examples

[0071] Example 1: Visiting tourist spots in Kyoto (Japanese user)

[0072] The user arrives at Kyoto Station and starts the system.

[0073] The device uses GPS to obtain its current location (Kyoto Station) and sends it to the server.

[0074] The server searches a database for information about tourist attractions around Kyoto Station and sends it to the terminal.

[0075] The device converts the received tourist information into an audio guide, stating, "This is Kyoto Station, located in the center of Kyoto."

[0076] When the user approaches an important point on their way to Kyoto Tower, the device will vibrate to notify them.

[0077] Example 2: Sightseeing for foreign visitors (English users)

[0078] A foreign user arrives at Todaiji Temple in Nara and starts up the system.

[0079] The device obtains its current location and sends it to the server.

[0080] The server retrieves information about Todaiji Temple from a database and generates a multilingual audio guide.

[0081] The device plays the received English audio guide, stating, "This is the Great Buddha Hall of Todaiji Temple, built in the Nara period."

[0082] This system provides users with appropriate tourist information in real time, enriching their tourist experience. It is also suitable for people with visual and hearing impairments, as well as foreign tourists, and helps them gain a deeper understanding of the history and culture of the places they visit.

[0083] The processing flow will be explained below.

[0084] Step 1:

[0085] When the user arrives at a tourist destination, they launch the application on their device (smartphone or smart glasses).

[0086] Step 2:

[0087] The device uses GPS to obtain the user's current location (latitude and longitude), which is temporarily stored in the device.

[0088] Step 3:

[0089] The device sends the acquired location information to a server via the network. The transmitted data includes the user ID and device ID.

[0090] Step 4:

[0091] The server analyzes the received location information and uses it to search a database for information on the nearest tourist attractions and cultural facilities.

[0092] Step 5:

[0093] The server organizes the tourist destination information obtained as search results and returns it to the user's device. The returned data includes the tourist destination's description text, audio URL, and location information of important points.

[0094] Step 6:

[0095] The user terminal analyzes the received tourist destination information and extracts text data to be converted into machine voice.

[0096] Step 7:

[0097] The device's built-in voice synthesis engine converts text data into voice data and provides audio guidance through the smart glasses' headphones.

[0098] Step 8:

[0099] As the user continues to explore the tourist spot, the device compares the current location with the location information of important points, and notifies the user with a vibration function when the user approaches an important point.

[0100] Step 9:

[0101] For the hearing impaired, the server sends tourist information in text format to the user's terminal.

[0102] Step 10:

[0103] The user terminal displays the received text information on the smart glasses display, providing a visual guide.

[0104] Step 11:

[0105] For foreign users, the server detects the user's language setting and automatically translates tourist information.

[0106] Step 12:

[0107] The translated tourist information is then sent back to the device, where a multilingual audio guide is generated and played back in the specified language.

[0108] Through the above process, users can receive detailed tourist information in real time by voice or text, improving their travel experience.

[0109] Example 1

[0110] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0111] Conventional tourist guide systems are limited in the information they provide based on the user's current location, and therefore lack the ability to provide detailed, real-time tourist information in multiple languages ​​or support people with visual or hearing impairments. Furthermore, they lack the ability to notify users when they are approaching important tourist spots, limiting the user's sightseeing experience.

[0112] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0113] In this invention, the server includes means for acquiring the user's current location, means for searching for related tourist attraction information from a tourist attraction database based on the acquired location information, means for transmitting the searched tourist attraction information to the user terminal, means for generating and playing an audio guide based on the transmitted tourist attraction information, means for using a speech synthesis engine to play the audio guide, means for notifying the user by vibration of important points near the user's destination, means for displaying tourist attraction information in text format for the hearing impaired, and means for automatically translating the tourist attraction information into multiple languages. This allows the user to obtain detailed tourist information in real time, enabling a comprehensive tourist experience for people with visual or hearing impairments. Furthermore, the user can receive vibration notifications when approaching important points at the destination, preventing them from missing important tourist attractions.

[0114] Below are definitions of each important word.

[0115] The "current location of the user" is specific information of geographical latitude and longitude obtained using the terminal held by the user.

[0116] A "tourist destination database" is a collection of information that holds tourist information about specific areas and facilities, and is used for searches.

[0117] "Related tourist destination information" is detailed information about tourist destinations that is searched from a database based on the user's current location.

[0118] The "means for transmitting to the user terminal" is a communication method for transmitting tourist destination information in data format from the server to the user terminal.

[0119] The "means for generating and playing back audio guides" refers to a technology for converting tourist destination information into audio data and playing back that audio data for the user.

[0120] "Speech synthesis engine" is a general term for software and hardware used to convert text data into speech data.

[0121] "Means for notifying by vibration at important points" is a function that causes the device to vibrate to notify the user when the user approaches an important location in a specific tourist spot.

[0122] "Means for displaying tourist destination information in text format for the hearing impaired" is a function that provides information to hearing impaired users by visually displaying audio information in text.

[0123] "Means for automatically translating tourist destination information into multiple languages" refers to technology for converting information related to tourist destinations into different languages ​​and providing that information to users.

[0124] The present invention relates to an audio guide system that provides detailed information in real time when visiting tourist spots and cultural facilities. This system aims to provide convenience to people with visual or hearing impairments, as well as foreign tourists.

[0125] System Configuration

[0126] The system mainly consists of the following components:

[0127] 1. User device: A mobile device such as a smartphone or smart glasses.

[0128] 2. Server: A server containing a tourist destination database.

[0129] 3. GPS module: A location information acquisition device for obtaining the user's current location.

[0130] 4. Communication means: A network connection for sending and receiving data between the user terminal and the server.

[0131] 5. Audio guide generation function: Software for converting tourist information into audio guides (e.g., Google Text-to-Speech API).

[0132] 6. Vibration function: Notification of important tourist spots.

[0133] 7. Text display function: A means of displaying text information for the hearing impaired.

[0134] 8. Automatic translation function: Software for translating tourist information into different languages ​​(e.g. Microsoft Translator API).

[0135] Program Overview

[0136] The program for this voice guidance system performs the following processes.

[0137] Get current location:

[0138] The user device periodically acquires the user's current location using the GPS module. The acquired location information (latitude and longitude) is stored in the device and sent to the server.

[0139] Search for tourist information:

[0140] The server analyzes the received location information and searches a database for information on the nearest tourist attractions and cultural facilities. It queries the tourist attractions database using SQL queries and stores the results in a data structure (e.g., a list or dictionary).

[0141] Sending and receiving tourist destination information:

[0142] The server sends the searched tourist destination information to the user's terminal, which analyzes the received information and converts it into an appropriate format.

[0143] Generate and provide audio descriptions:

[0144] The user device converts the received text information into audio guidance using a speech synthesis engine, which is then played through the headphones in the smart glasses worn by the user.

[0145] Vibration notification:

[0146] When the user approaches an important point in a tourist destination, the user's device will notify them using a vibration function.

[0147] Text information provided:

[0148] For the hearing impaired, the server sends tourist information in text format to the device, which then displays the text on the smart glasses display.

[0149] Machine translation feature:

[0150] For foreign visitors to Japan, the server automatically translates tourist information and generates multilingual audio guides. The translated audio guides are sent to the user's device and played in the specified language.

[0151] Specific examples

[0152] Example 1: Visiting tourist spots in Kyoto (Japanese user)

[0153] 1. The user arrives at Kyoto Station and starts the system.

[0154] 2. The device uses GPS to obtain its current location (Kyoto Station) and sends it to the server.

[0155] 3. The server searches the database for information about tourist spots around Kyoto Station and sends it to the terminal.

[0156] 4. The device converts the received tourist information into an audio guide, stating, "This is Kyoto Station, located in the center of Kyoto."

[0157] 5. When the user approaches an important point on the way to Kyoto Tower, the device will vibrate to notify them.

[0158] Example 2: Sightseeing for foreign visitors (English users)

[0159] 1. A foreign user arrives at Todaiji Temple in Nara and starts up the system.

[0160] 2. The device obtains its current location and sends it to the server.

[0161] 3. The server retrieves information about Todaiji Temple from the database and generates a multilingual audio guide.

[0162] 4. The device plays the received English audio guide, stating, "This is the Great Buddha Hall of Todaiji Temple, built in the Nara period."

[0163] Prompt Sentence Examples

[0164] "Please explain the main tourist spots around Kyoto Station."

[0165] "Please tell me in English about the history and highlights of Todaiji Temple."

[0166] This system provides users with appropriate tourist information in real time, enriching their tourist experience. It is also suitable for people with visual and hearing impairments, as well as foreign tourists, and helps them gain a deeper understanding of the history and culture of the places they visit.

[0167] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0168] Step 1: Get current location

[0169] The user device uses a GPS module to obtain the user's current location. Specifically, the device's GPS module receives satellite signals and calculates latitude and longitude information. The obtained location information (e.g., latitude 35.0116, longitude 135.7681) is stored in internal memory and sent to the server. The input is satellite signal data from the GPS module, and the output is latitude and longitude coordinate information.

[0170] Step 2: Search for tourist information

[0171] The server analyzes the latitude and longitude location information received from the user device. Based on the received data, it uses an SQL query to search the tourist spot database and retrieves the relevant tourist spot information. For example, the query executed is "SELECT FROM tourist_spots WHERE latitude BETWEEN 34.9116 AND 35.1116 AND longitude BETWEEN 135.6681 AND 135.8681;". The input is latitude and longitude location information, and the output is a list of related tourist spot information.

[0172] Step 3: Submit tourist destination information

[0173] The server formats the tourist spot information obtained from the search into JSON format and sends it to the user's device over the network. For example, the JSON data sent is "{"spots": [{"name": "Kyoto Station", "description": "Central location in Kyoto"}]}". The input is a list of tourist spot information, and the output is the JSON data sent to the user's device.

[0174] Step 4: Receiving and analyzing tourist destination information

[0175] The user device parses the tourist attraction information in JSON format received from the server. Specifically, it uses a JSON parser to extract the data and saves it in its internal memory. For example, the information parsed is "{"name": "Kyoto Station", "description": "Central location in Kyoto"}". The input is the JSON data received from the server, and the output is the parsed tourist attraction information.

[0176] Step 5: Generate audio guide

[0177] The user device sends the received text information to a speech synthesis engine (e.g., Google Text-to-Speech API) and converts it into audio data. Specifically, it converts the text "This is Kyoto Station, located in the center of Kyoto" into an audio file. The input is the tourist destination information text, and the output is an audio file.

[0178] Step 6: Play the audio guide

[0179] The user device passes the generated audio file to the playback function, which plays the audio through the headphones of the user's smart glasses. This allows the user to hear the audio guidance, "This is Kyoto Station, located in the center of Kyoto." The input is the audio file, and the output is the audio guidance playback to the user.

[0180] Step 7: Vibration notifications at key points

[0181] The user device monitors the latitude and longitude of the current location and the tourist spot, and activates vibration when the distance is within a certain range. For example, the device vibrates when the user approaches Kyoto Tower. The input is the current location and the location information of the tourist spot, and the output is the activation of vibration.

[0182] Step 8: Displaying text information for the hearing impaired

[0183] The server sends tourist destination information in text format to the user's device. The user's device then calls an API to display the received text information on the smart glasses' display. For example, the text information displayed is "This is Kyoto Station, located in the center of Kyoto." The input is tourist destination information in text format, and the output is the text information displayed on the smart glasses' display.

[0184] Step 9: Automatic translation function

[0185] The server uses the Microsoft Translator API to automatically translate tourist information for foreign visitors to Japan. For example, it translates the Japanese phrase "This is Kyoto Station, located in the center of Kyoto" into English. The translated text is sent to a speech synthesis engine, which generates an audio guide. The input is the tourist information text, and the output is the translated audio guide.

[0186] In each processing step, detailed data processing and operations are coordinated to realize a system that provides appropriate tourist information to users.

[0187] (Application example 1)

[0188] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0189] In modern commercial facilities, the provision of information to diverse customers, such as users with visual or hearing impairments and foreign visitors to Japan who require multilingual support, is insufficient. In particular, it is difficult to provide real-time information in stores, such as product information and sales floor guidance, which reduces convenience for users. Given this background, it is necessary to provide consistent information to users with different disabilities and languages.

[0190] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0191] In this invention, the server includes means for acquiring the user's current location, means for searching for related tourist attraction information from a tourist attraction database based on the acquired location information, means for transmitting the searched tourist attraction information to the user terminal, means for generating and playing an audio guide based on the transmitted tourist attraction information, means for acquiring in-store location information, means for searching for related product information based on the acquired in-store location information, means for transmitting the searched product information to the user terminal, and means for generating and playing an audio guide based on the transmitted product information. This makes it possible to provide consistent real-time information to different customers, such as users with visual or hearing impairments and foreign visitors to Japan.

[0192] The "means for acquiring the user's current location" refers to a technical device for measuring and acquiring the latitude and longitude of the user's current location or a specific point within the store.

[0193] "Location information" is data indicating the latitude and longitude of the user's current location or a specific point within a store.

[0194] A "tourist destination database" is a collection of data that systematically stores information about tourist destinations.

[0195] The "means for searching tourist attraction information" is a technical device for searching for relevant information from a tourist attraction database based on the acquired location information.

[0196] A "user terminal" is an electronic device carried by a user, such as a smartphone or smart glasses.

[0197] "Means for generating and playing audio guides" refers to technical devices that convert text information into audio and play it back through the user terminal.

[0198] "Means for obtaining location information within the store" refers to a technical device that uses beacons, Wi-Fi, or other location measurement technologies to identify and obtain the user's specific location within the store.

[0199] The "means for searching for related product information" is a technical device for searching the database for the most suitable product information based on the acquired in-store location information.

[0200] The "vibration function" is a technical device that causes the user's terminal to vibrate to notify the user when the user approaches a specific location.

[0201] "Means for displaying in text format" refers to a technical device for displaying information as a string of characters on the display of a user terminal for the hearing impaired.

[0202] "Universal service" is a service that is consistently provided to all users, regardless of whether they have a disability or speak a different language.

[0203] The present invention is a voice guide system that provides users with detailed information in real time at tourist spots and commercial facilities, with the aim of providing convenience to people with visual or hearing impairments, as well as foreign tourists and shoppers.

[0204] System Configuration

[0205] The system mainly consists of the following components:

[0206] 1. User device: A mobile device such as a smartphone or smart glasses.

[0207] 2. Server: A server containing a tourist destination and product database.

[0208] 3. GPS module: A location information acquisition device for obtaining the user's current location.

[0209] 4. Communication means: A network connection for sending and receiving data between the user terminal and the server.

[0210] 5. Audio guide generation function: Software for converting tourist destination information and product information into audio guides.

[0211] 6. Vibration function: A means of notifying you of important tourist spots and product areas.

[0212] 7. Text display function: A means of displaying text information for the hearing impaired.

[0213] 8. In-store location information acquisition means: A device that acquires the user's location within the store using beacons or Wi-Fi access points.

[0214] Basic operation

[0215] The operation of the system is as follows.

[0216] The user terminal periodically acquires its current location using a GPS module or location information acquisition means within the store. This location information is stored in the user terminal and sent to the server.

[0217] The server analyzes the received location information and searches a database for information on the nearest tourist spots and products, and sends the searched information to the user's device in real time.

[0218] The user device analyzes the received information, converts it into an appropriate format, and generates audio guidance using a speech synthesis engine (e.g., Google Text-to-Speech API or Amazon Polly) that is then played through the user's headphones.

[0219] When the user approaches an important tourist spot or product area, the user's device will notify them using the vibration function.

[0220] For the hearing impaired, the server sends tourist destination and product information in text format to the user's device, and the text information is displayed on the display of smart glasses or a smartphone.

[0221] For foreign visitors to Japan, the server automatically translates tourist destination and product information, generates multilingual audio guides, and sends them to the user's device.

[0222] Specific examples

[0223] Example of a tourist destination application

[0224] When a user visits a tourist spot, they can activate this system, which will provide audio guidance with detailed tourist spot information and historical background based on their current location. For example, when a user arrives at Kyoto Station and activates the system, the device obtains their current location (Kyoto Station) and sends it to the server. The server searches a tourist spot database for information about tourist spots around Kyoto Station and sends it to the device. The device then converts this into an audio guide, stating, "This is Kyoto Station, located in the center of Kyoto."

[0225] Example of a physical store guidance application

[0226] When a user activates this system while shopping in a commercial facility, it acquires location information within the store and provides voice guidance on related product information. For example, if the user is in the electronics section of a department store, the device acquires the user's current location and sends it to the server. The server then searches for related product information from a product database and sends it to the device. The device then converts this into voice guidance, stating, "These are the latest earphone models." Additionally, when the user approaches a specific product area, the device vibrates to notify the user.

[0227] Prompt Sentence Examples

[0228] "I want to provide detailed product information based on the visitor's current location, tailored to their nationality and language preferences. The location information is coordinates (X: 100, Y: 200), and the product information is "high-sensitivity earphones" in Japanese. Please translate this into English and generate an audio guide."

[0229] This will enable consistent, real-time information to be provided to different customers, including users with visual or hearing impairments and foreign visitors to Japan.

[0230] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0231] Step 1:

[0232] Get current location

[0233] The user device periodically acquires the user's current location using a GPS module or in-store location information acquisition means (beacons or Wi-Fi). Specifically, when the user starts the system, the GPS module acquires latitude and longitude information, or receives signals from in-store beacons to identify the X and Y coordinates. This generates current location information (input) and stores it in the user device (output).

[0234] Step 2:

[0235] Sending current location information

[0236] The user device sends the acquired current location information to the server. Specifically, it uses a network communication method (Wi-Fi or mobile data communication) to send the location information data to the server. As a result, the location information data from the user device (input) is sent to the server (output).

[0237] Step 3:

[0238] Search for tourist information or product information

[0239] The server analyzes the received location information and searches for related tourist attraction information or product information from a tourist attraction database or product database. Specifically, it uses a database search algorithm to extract the nearest tourist attraction or product information from the location information (input). This process generates related tourist attraction information and product information (output).

[0240] Step 4:

[0241] Sending related information

[0242] The server sends the searched tourist spot information or product information to the user terminal. Specifically, the tourist spot information or product information (input) is sent to the user terminal via the network. This operation allows the tourist spot information or product information from the server to reach the user terminal (output).

[0243] Step 5:

[0244] Audio guide generation and provision

[0245] The user device analyzes the received information and converts it into an appropriate format. Specifically, it uses a speech synthesis engine (Google Text-to-Speech API or Amazon Polly) to convert text information (input) into voice data. The converted voice guide (output) is played through the user device's headphones.

[0246] Step 6:

[0247] Vibration notification

[0248] When a user approaches a particular tourist attraction or product area, the user's device will vibrate to notify them. This is done by calculating the distance to each set point based on the device's location data (input), and vibrating when the distance falls below a certain level (output).

[0249] Step 7:

[0250] Providing text information

[0251] For the hearing impaired, the server sends tourist destination information and product information in text format to the user's device. The user's device then displays this information on its display. Specifically, the text data (input) is screen-rendered for display. This provides the user with visual text information (output).

[0252] Step 8:

[0253] Machine translation and multilingual support

[0254] The server automatically translates tourist destination information and product information to generate multilingual audio guides for foreign visitors to Japan. Specifically, it translates text information (input) into the specified language using a multilingual translation model (e.g., Google Translate API). This translated text information is then converted into audio guides using a speech synthesis engine and sent to the user's device. This provides a translated multilingual audio guide (output).

[0255] Specific example prompts

[0256] "I want to provide detailed product information based on the visitor's current location, tailored to their nationality and language preferences. The location information is coordinates (X: 100, Y: 200), and the product information is "high-sensitivity earphones" in Japanese. Please translate this into English and generate an audio guide."

[0257] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0258] The present invention relates to an audio guide system that provides detailed information in real time when visiting tourist spots and cultural facilities. This system aims to provide convenience especially for people with visual or hearing impairments and foreign tourists. Furthermore, the present invention realizes a personalized sightseeing experience by recognizing the user's emotional state and providing information accordingly.

[0259] System Configuration

[0260] The system mainly consists of the following components:

[0261] 1. User device: A mobile device such as a smartphone or smart glasses.

[0262] 2. Server: A server containing a tourist destination database.

[0263] 3. GPS module: A location information acquisition device for obtaining the user's current location.

[0264] 4. Communication means: A network connection for sending and receiving data between the user terminal and the server.

[0265] 5. Audio guide generation function: Software for converting tourist information into audio guides.

[0266] 6. Vibration function: Notification of important tourist spots.

[0267] 7. Text display function: A means of displaying text information for the hearing impaired.

[0268] 8. Emotion engine: Software that recognizes emotional states and dynamically changes the content and format of the tourist destination information provided.

[0269] Program processing

[0270] 1. Obtaining the current location

[0271] The user device periodically acquires the user's current location (latitude and longitude) using the GPS module. This location information is temporarily stored in the device and then sent to the server.

[0272] 2. Search for tourist information

[0273] The server analyzes the received location information and uses it to search a database for information on the nearest tourist attractions and cultural facilities.

[0274] 3. Sending and receiving tourist destination information

[0275] The server organizes the tourist destination information obtained as search results and returns it to the user's device. The returned data includes the tourist destination's description text, audio URL, and location information of important points.

[0276] 4. Generating and Providing Audio Guides

[0277] The user device converts the received text information into audio guidance using a speech synthesis engine, which is then played through the headphones in the smart glasses worn by the user.

[0278] 5. Vibration notifications

[0279] When the user approaches an important point in a tourist destination, the user's device will notify them using a vibration function.

[0280] 6. Providing text information

[0281] For the hearing impaired, the server sends tourist information in text format to the user's device, which then displays the received text information on the smart glasses display, providing a visual guide.

[0282] 7. Automatic translation function

[0283] For foreign visitors to Japan, the server detects the user's language preference and automatically translates tourist information. The translated tourist information is then sent back to the device, which generates a multilingual audio guide. The device then plays the audio guide in the specified language.

[0284] 8. Emotion recognition and personalized information delivery

[0285] The emotion engine installed in the user device uses cameras and biometric sensors to recognize the user's emotional state from their facial expressions and biometric data. For example, if the user is excited, the emotion engine will recognize this and prioritize providing information on particularly attractive tourist spots and activity recommendations.

[0286] If the user is under stress, the system will provide information on relaxing spots and quiet places. By providing appropriate information according to the user's emotional state, the system can improve the user's satisfaction.

[0287] Specific examples

[0288] Example 1: Visiting tourist spots in Kyoto (Japanese user)

[0289] The user arrives at Kyoto Station and starts the system. The device uses GPS to obtain its current location (Kyoto Station) and sends it to the server. The server then searches a database for information about tourist spots around Kyoto Station and sends it to the device. The device converts the received tourist spot information into an audio guide, providing instructions such as, "This is Kyoto Station, located in the center of Kyoto." As the user approaches important points on their way to Kyoto Tower, the device vibrates to notify them. If the emotion engine detects the user's excitement, it provides additional information such as, "You can see a wonderful night view from the observation deck of Kyoto Tower."

[0290] Example 2: Sightseeing for foreign visitors (English users)

[0291] A foreign user arrives at Todaiji Temple in Nara and starts the system. The device obtains its current location and sends it to the server. The server retrieves information about Todaiji Temple from a database and generates a multilingual audio guide. The device plays the received English audio guide, stating, "This is the Great Buddha Hall of Todaiji Temple, built in the Nara period." If the emotion engine detects the user's stress level, it also provides information such as, "There is a quiet, relaxing park nearby."

[0292] This system can improve the travel experience by providing users with appropriate tourism information in real time via voice or text. Furthermore, the emotion engine enables information to be provided according to the user's emotional state, realizing a more personalized and satisfying tourism experience.

[0293] The processing flow will be explained below.

[0294] Step 1:

[0295] When the user arrives at a tourist destination, they launch the application on their device (smartphone or smart glasses).

[0296] Step 2:

[0297] The device uses GPS to obtain the user's current location (latitude and longitude), which is temporarily stored in the device.

[0298] Step 3:

[0299] The device sends the acquired location information to a server via the network. The transmitted data includes the user ID and device ID.

[0300] Step 4:

[0301] The server analyzes the received location information and uses it to search a database for information on the nearest tourist attractions and cultural facilities.

[0302] Step 5:

[0303] The server organizes the tourist destination information obtained as search results and returns it to the user's device. The returned data includes the tourist destination's description text, audio URL, and location information of important points.

[0304] Step 6:

[0305] The user terminal analyzes the received tourist destination information and extracts text data to be converted into machine voice.

[0306] Step 7:

[0307] The device's built-in voice synthesis engine converts text data into voice data and provides audio guidance through the smart glasses' headphones.

[0308] Step 8:

[0309] As the user continues to explore the tourist spot, the device compares the current location with the location information of important points, and notifies the user with a vibration function when the user approaches an important point.

[0310] Step 9:

[0311] For the hearing impaired, the server sends tourist information in text format to the user's terminal.

[0312] Step 10:

[0313] The user terminal displays the received text information on the smart glasses display, providing a visual guide.

[0314] Step 11:

[0315] For foreign users, the server detects the user's language setting and automatically translates tourist information.

[0316] Step 12:

[0317] The translated tourist information is then sent back to the device, where a multilingual audio guide is generated and played back in the specified language.

[0318] Step 13:

[0319] The emotion engine installed in the user terminal uses a camera and biometric sensors to recognize the user's emotional state from their facial expressions and biometric data.

[0320] Step 14:

[0321] The device transmits the recognized emotional state to the server, which then selects tourist destination information according to the emotion and retransmits it to the user device.

[0322] Step 15:

[0323] If the emotion engine detects the user's excited state, the server will prioritize providing information on tourist spots and activity guides that are more likely to interest the user.

[0324] Step 16:

[0325] If the emotion engine detects the user's stress state, the server will provide information on relaxation spots and quiet places.

[0326] Step 17:

[0327] The user terminal provides the newly received tourist spot information to the user by voice or text.

[0328] Example 2

[0329] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0330] Conventional tourist guide systems do not provide sufficient convenience for people with visual or hearing impairments or foreign tourists. Furthermore, they do not provide customized information based on the user's emotional state, making it difficult to realize a personalized tourist experience. Furthermore, they lack multilingual support, making it difficult to provide information appropriately to users with different cultural backgrounds.

[0331] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring the user's current location, means for searching for related tourist attraction information from a tourist attraction database based on the acquired location information, means for transmitting the searched tourist attraction information to the user terminal, means for generating and playing back an audio guide based on the transmitted tourist attraction information, means for recognizing the user's emotional state and dynamically changing the content and format of the tourist attraction information based on the recognition result, means for automatically translating the tourist attraction information and providing a multilingual audio guide, and means for displaying the tourist attraction information in text format for the hearing impaired. This provides convenience for people with visual or hearing impairments and foreign travelers, enables the provision of customized information according to the user's emotional state, and makes it possible to provide appropriate information to users with different cultural backgrounds by supporting multiple languages.

[0332] The "means for acquiring the user's current location" is a function for acquiring the user's current location using a GPS module or other location information acquisition device.

[0333] The "means for searching for related tourist attraction information from a tourist attraction database based on the acquired location information" is a function in which the server analyzes the location information received and uses that information to search for information on the nearest tourist attractions and cultural facilities from the tourist attraction database.

[0334] "Means for sending searched tourist spot information to user terminal" refers to the function by which the server returns tourist spot information obtained as search results to the user terminal. The returned data includes explanatory text of the tourist spot, audio URL, and location information of important points.

[0335] "Means for generating and playing back audio guides based on transmitted tourist attraction information" refers to a function that converts tourist attraction information received by a user terminal into audio guides using a voice synthesis engine and plays them back.

[0336] "Means for recognizing the user's emotional state and dynamically changing the content and format of tourist destination information based on the recognition results" refers to a function in which the emotion engine installed in the device uses the camera and biometric sensors to recognize the user's emotional state and customizes the information provided based on the results.

[0337] "Means for automatically translating tourist destination information and providing multilingual audio guides" refers to a function in which the server detects the user's set language, translates the tourist destination information into the corresponding language, and generates multilingual audio guides based on that translation information.

[0338] The "means for displaying tourist spot information in text format for the hearing impaired" is a function in which the server transmits tourist spot information in text format to the user terminal, and the user terminal displays it visually.

[0339] This invention relates to an audio guide system that provides detailed information in real time when visiting tourist spots and cultural facilities. This system aims to provide convenience especially for people with visual or hearing impairments, as well as foreign tourists. Furthermore, it realizes a personalized sightseeing experience by recognizing the user's emotional state and providing information according to it.

[0340] System Configuration

[0341] The system mainly consists of the following components:

[0342] 1. User device: Use a mobile device such as a smartphone or smart glasses.

[0343] 2. Server: A server that contains a tourist destination database, and the database uses MySQL or similar.

[0344] 3. GPS module: Includes a location information acquisition device for acquiring the user's current location.

[0345] 4. Communication method: A network connection for sending and receiving data between the user device and the server. Wi-Fi, 4G / 5G, etc. are used.

[0346] 5. Audio guide generation function: Google Cloud Text-to-Speech is used as software to convert tourist information into audio guides.

[0347] 6. Vibration function: The device's vibration function is used to notify users of important tourist spots.

[0348] 7. Text display function: Equipped with a display to display text information for the hearing impaired.

[0349] 8. Emotion engine: Cameras and biometric sensors are used as software to recognize emotional states and dynamically change the content and format of tourist information provided.

[0350] Program processing explanation

[0351] Get current location

[0352] The user device periodically acquires the user's current location (latitude and longitude) using the GPS module. This location information is temporarily stored in the device and then sent to the server.

[0353] Search for tourist information

[0354] The server analyzes the received location information and uses it to search a database for information on the nearest tourist attractions and cultural facilities.

[0355] Sending and receiving tourist information

[0356] The server organizes the tourist destination information obtained as search results and returns it to the user's device. The returned data includes the tourist destination's description text, audio URL, and location information of important points.

[0357] Generating and providing audio guides

[0358] The user device converts the received text information into audio guidance using a speech synthesis engine, which is then played through the headphones in the smart glasses worn by the user.

[0359] Vibration notification

[0360] When the user approaches an important point in a tourist destination, the user's device will notify them using a vibration function.

[0361] Providing text information

[0362] For the hearing impaired, the server sends tourist information in text format to the user's device, which then displays the received text information on the smart glasses display, providing a visual guide.

[0363] Machine translation function

[0364] For foreign visitors to Japan, the server detects the user's language preference and automatically translates tourist information. The translated tourist information is then sent back to the device, which generates a multilingual audio guide. The device then plays the audio guide in the specified language.

[0365] Emotion recognition and customized information provision

[0366] The emotion engine installed in the user device uses cameras and biometric sensors to recognize the user's emotional state from their facial expressions and biometric data. For example, if the user is excited, the emotion engine will recognize this and provide them with information on particularly attractive tourist spots and activity recommendations first. If the user is stressed, it will provide them with information on relaxing spots and quiet places. Providing appropriate information according to the user's emotional state can improve user satisfaction.

[0367] Specific examples

[0368] Example 1: Visiting tourist spots in Kyoto (Japanese user)

[0369] The user arrives at Kyoto Station and starts the system. The device uses GPS to obtain its current location (Kyoto Station) and sends it to the server. The server then searches a database for information about tourist spots around Kyoto Station and sends it to the device. The device converts the received tourist spot information into an audio guide, providing instructions such as, "This is Kyoto Station, located in the center of Kyoto." As the user approaches important points on their way to Kyoto Tower, the device vibrates to notify them. If the emotion engine detects the user's excitement, it provides additional information such as, "You can see a wonderful night view from the observation deck of Kyoto Tower."

[0370] Example 2: Sightseeing for foreign visitors (English users)

[0371] A foreign user arrives at Todaiji Temple in Nara and starts the system. The device obtains its current location and sends it to the server. The server retrieves information about Todaiji Temple from a database and generates a multilingual audio guide. The device plays the received English audio guide, stating, "This is the Great Buddha Hall of Todaiji Temple, built in the Nara period." If the emotion engine detects the user's stress level, it also provides information such as, "There is a quiet, relaxing park nearby."

[0372] Prompt Sentence Examples

[0373] As an example of an input prompt for a generative AI model, you could enter something like:

[0374] "Please explain in detail the process of the system that, when the user arrives at a tourist spot, acquires the user's current location using the smartphone, acquires tourist spot information from the server, and customizes the information according to the user's emotional state."

[0375]

[0376] "Describe the detailed process by which a real-time audio guide system uses a smartphone to acquire the current location, fetch tourist information from a server, and customize the information based on the user's emotional state."

[0377] This system can improve the travel experience by providing users with appropriate tourism information in real time via voice or text. Furthermore, the emotion engine enables information to be provided according to the user's emotional state, realizing a more personalized and satisfying tourism experience.

[0378] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0379] Step 1: Get current location

[0380] Input: The user launches the app and the GPS module obtains the current location.

[0381] Specific operation: A user arrives at a tourist destination and launches the app on their smartphone or smart glasses. The device immediately accesses the built-in GPS module to obtain the user's current location (latitude and longitude). The message "Retrieving user's current location..." is displayed.

[0382] Data processing or data calculation: Position information (e.g., 35.0116°N, 135.7681°E) is acquired and temporarily stored in the device's memory.

[0383] Output: The obtained location information is sent to the server.

[0384] Step 2: Search for tourist information

[0385] Input: The server receives the user's location information.

[0386] Specific operation: The server performs a database search based on the location information (latitude and longitude) received. The database stores tourist spot names, descriptions, audio guide URLs, etc.

[0387] Data processing or data calculation: The server analyzes the location information and searches for information on nearby tourist attractions in a database (e.g., MySQL).

[0388] Output: Tourist information (e.g., Kyoto Tower description text, audio URL, location information of important points) is obtained and sent to the user's device.

[0389] Step 3: Sending and receiving tourist destination information

[0390] Input: The server organizes tourist destination information from search results.

[0391] Specific operation: The server formats the tourist attraction information obtained as a search result, packages the necessary data (explanatory text, audio URL, location information of important points, etc.) and sends it to the user's terminal.

[0392] Data processing or data calculation: Convert tourist destination information data into a data structure such as JSON format, making it a format that can be transferred efficiently.

[0393] Output: The formatted tourist destination information is sent to the user's terminal.

[0394] Step 4: Generate and serve audio descriptions

[0395] Input: The user terminal receives tourist destination information from the server.

[0396] Specific operation: The user device passes the received text information to a speech synthesis engine (e.g., Google Cloud Text-to-Speech), which converts the information into audio data. The generated audio guide is played through the smart glasses headphones.

[0397] Data processing or data calculation: Converting text data into audio data and saving it in a playable format.

[0398] Output: The audio description is played to the user.

[0399] Step 5: Vibration Notification

[0400] Input: The user approaches a point of interest at a tourist destination.

[0401] Specific operation: When the user approaches a key point, the device receives location information from the built-in GPS and determines that it is a key point. The device then starts vibrating and notifies the user that "You are approaching a key point."

[0402] Data processing or data calculation: Compare the current location with the location of important points, and activate the vibration function when they come within a certain distance.

[0403] Output: A vibration will be activated to notify the user.

[0404] Step 6: Provide text information

[0405] Input: The server sends tourist information for the hearing impaired in text format.

[0406] Specific operation: The server sends text information to the user's terminal, which then displays it on the smart glasses display.

[0407] Data processing or data calculation: Converting text information into a displayable format and presenting it on a display.

[0408] Output: Text information is displayed on the smart glasses display.

[0409] Step 7: Automatic translation function

[0410] Input: The server detects the user's preferred language.

[0411] Specific operation: The server reads the language setting of the user's device, automatically translates the tourist attraction information based on that setting, and then sends the translated information back to the user's device.

[0412] Data processing or data calculation: Translate tourist destination information into the specified language using an automatic translation tool (e.g., Google Translate API).

[0413] Output: The translated tourist information is sent to the user's terminal, and an audio guide is played in the specified language.

[0414] Step 8: Emotion recognition and customized information delivery

[0415] Input: The user's emotional state is acquired by a camera or biometric sensors installed on the user's device.

[0416] How it works: The emotion engine analyzes the user's emotional state using facial recognition technology, heart rate measurement, etc. For example, if the user is excited, it will prioritize providing "attractive spots."

[0417] Data processing or data calculation: An algorithm is used to analyze the acquired biometric and facial expression data and determine the emotional state.

[0418] Output: Customized information according to the emotional state is provided to the user.

[0419] (Application example 2)

[0420] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0421] The present invention aims to improve convenience for people with visual or hearing impairments and foreign tourists by providing detailed information in real time when visiting tourist spots and cultural facilities through an audio guide system. It also aims to realize a personalized tourist experience by recognizing the user's emotional state and providing information accordingly. Furthermore, this technology can be applied to a factory work support system to provide workers with real-time work guides and important information, and to improve work efficiency and safety by detecting the user's fatigue and stress level and providing appropriate rest instructions.

[0422] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0423] In this invention, the server includes means for acquiring the user's current location, means for searching for related tourist attraction information from a tourist attraction database based on the acquired location information, means for transmitting the searched tourist attraction information to the user terminal, means for generating and playing back an audio guide based on the transmitted tourist attraction information, and means for recognizing the user's emotional state and dynamically changing the content and format of the tourist attraction information to be provided in accordance with that state. This makes it possible to provide the user with appropriate information in real time when visiting tourist attractions or working in a factory, improving the user experience and work efficiency.

[0424] "User's current location" refers to the geographic location of a particular user as determined by a GPS module or other location acquisition function.

[0425] A "tourist destination database" refers to an information system that stores information about tourist destinations and cultural facilities in a structured format.

[0426] "Tourist destination information" refers to detailed information such as descriptions, images, and audio guides about specific tourist destinations and cultural facilities.

[0427] "User terminal" refers to an electronic device that a user can carry around, such as a smartphone or smart glasses.

[0428] "Audio guide" refers to systems and content that provide audio information about tourist destinations and cultural facilities.

[0429] "Emotional state" refers to the user's emotional and mental state obtained by analyzing the user's facial expressions and biosensor data.

[0430] The "vibration function" refers to a function that notifies the user by vibrating the user terminal.

[0431] "Text format" refers to a format in which information is displayed in text.

[0432] "Real-time" refers to providing current information and status immediately without delay.

[0433] "Work guide" refers to information on instructions and operating procedures provided when working in a factory.

[0434] "Rest instructions" refer to instructions regarding when and where the user should take a rest.

[0435] "Personalized tourism experience" refers to customized tourism information and guidance provided based on the user's individual circumstances and interests.

[0436] "Multilingual" refers to the ability to provide information and guidance in multiple languages.

[0437] "Factory work support system" refers to technology and platforms that support work within a factory.

[0438] This invention is a voice guide system that provides users with detailed information in real time when visiting tourist spots and cultural facilities. This system aims to provide convenience, especially for people with visual or hearing impairments and foreign tourists. Furthermore, it realizes a personalized tourist experience by recognizing the user's emotional state and providing information accordingly. This technology can also be applied to a factory work support system, providing real-time work guidance and important information to workers, and improving work efficiency and safety by detecting the user's fatigue and stress level and providing appropriate rest instructions.

[0439] System Configuration

[0440] The system consists of the following components:

[0441] 1. User Device:

[0442] These are portable electronic devices such as smartphones and smart glasses that enable users to acquire and display information at tourist spots and in factories.

[0443] 2. Server:

[0444] This server contains a database of tourist attractions and data on factory work. The server provides relevant information based on requests from user terminals.

[0445] 3. GPS module:

[0446] This is a location information acquisition device for acquiring the current location of the user.

[0447] 4. Means of communication:

[0448] A network connection for sending and receiving data between a user terminal and a server.

[0449] 5. Audio guide generation function:

[0450] This is software for converting tourist information into audio guides. A typical example of this software is the TextToSpeech module.

[0451] 6. Vibration function:

[0452] It is a means of notifying important tourist spots and important locations within the factory.

[0453] 7. Text display function:

[0454] This is a way to display tourist information and work guides as text information for the hearing impaired. For this, we use the DisplayTextModule.

[0455] 8. Emotion Engine:

[0456] This software recognizes the user's emotional state and dynamically changes the content and format of the information provided. For example, the EmotionRecognition module can be used to analyze the user's emotional state.

[0457] Operating principle

[0458] The server includes a means for acquiring the user's current location, a means for searching for related information from a tourist spot database or a factory work database based on the acquired location information, a means for transmitting the searched information to the user terminal, a means for generating and playing back an audio guide based on the transmitted information, and a means for recognizing the user's emotional state and dynamically changing the content and format of the information to be provided depending on that state. This makes it possible to provide the user with appropriate information in real time while visiting tourist spots or working in a factory, improving the user experience and work efficiency.

[0459] Specific examples

[0460] Example 1: Factory work support system

[0461] Users wear smart glasses while moving around the factory to perform their work. The system uses a GPS module to obtain their current location, and the server searches for relevant work information based on the obtained location information. The received information is provided to the user as voice guidance and text display. The system also uses an EmotionRecognition module to recognize the user's emotional state, and if fatigue or stress is detected, it provides a vibration function and instructions to take a break.

[0462] Prompt Sentence Examples

[0463] "Please suggest an application program that updates the smart glasses feed in real time and displays work instructions based on the user's current location. Also, include a feature that automatically notifies the user to take a break if they are feeling stressed."

[0464] As described above, the present invention is a system for significantly improving work efficiency at tourist spots and factories. By implementing this system, it is expected that the user experience will be improved and work efficiency will increase.

[0465] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0466] Step 1:

[0467] The user terminal acquires its current location (latitude and longitude) using a GPS module.

[0468] input:

[0469] Location information obtained from the GPS module of the user's smart glasses or smartphone.

[0470] Data processing and calculation:

[0471] Obtain latitude and longitude data of your current location and temporarily store it on the device.

[0472] output:

[0473] Numerical data of latitude and longitude obtained.

[0474] Step 2:

[0475] The terminal transmits information about its current location to the server.

[0476] input:

[0477] Numerical latitude and longitude data obtained in step 1.

[0478] Specific behavior:

[0479] The device transmits GPS data to a server via a network.

[0480] output:

[0481] A request to send location information to the server.

[0482] Step 3:

[0483] The server analyzes the received location information and searches for related information from a tourist destination database or a factory work database.

[0484] input:

[0485] Latitude and longitude data of the current location sent from the device.

[0486] Data processing and calculation:

[0487] Query a database based on location information to find relevant tourist information or work procedures.

[0488] output:

[0489] Associated tourist attraction information and work procedure data.

[0490] Step 4:

[0491] The server organizes the obtained information and returns it to the user terminal.

[0492] input:

[0493] Tourist destination information and work procedure data searched in Step 3.

[0494] Specific behavior:

[0495] The server organizes the information, converts it into a specific format, and then transmits it over the network to the user terminal.

[0496] output:

[0497] Tourist destination information and work procedure data are sent to the device.

[0498] Step 5:

[0499] The user terminal converts the received information into audio guidance using a speech synthesis engine (TextToSpeech module) and plays it back.

[0500] input:

[0501] Received text data of tourist destination information and work procedures.

[0502] Data processing and calculation:

[0503] The process of converting text information into speech.

[0504] output:

[0505] Audio data played as audio guide.

[0506] Specific behavior:

[0507] The device uses the TextToSpeech module to convert the text into speech, which is played through the smart glasses' headphones.

[0508] Step 6:

[0509] When the user approaches an important point in a tourist spot or an important location in a factory, the device will notify them by vibrating.

[0510] input:

[0511] User location and location information of important points for tourist attractions and work procedures.

[0512] Data processing and calculation:

[0513] The current location is compared with the location of an important point and it is detected when it enters a certain range.

[0514] output:

[0515] Vibration notifications.

[0516] Specific behavior:

[0517] The user terminal activates the vibration function to notify the user.

[0518] Step 7:

[0519] The user terminal uses the EmotionRecognition module to recognize the user's emotional state and provide information according to the emotional state.

[0520] input:

[0521] Emotional state data from cameras and biometric sensors (e.g., facial expressions, heart rate).

[0522] Data processing and calculation:

[0523] The emotional state data is analyzed to recognize the emotional state of the user.

[0524] output:

[0525] Providing customized information according to the user's emotional state.

[0526] Specific behavior:

[0527] The EmotionRecognition module assesses the user's emotional state and provides information, including instructions to rest, if the user feels tired, for example.

[0528] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0529] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0530] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0531] [Second embodiment]

[0532] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0533] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0534] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0535] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0536] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0537] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0538] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0539] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0540] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0541] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0542] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0543] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0544] The present invention relates to an audio guide system that provides detailed information in real time when visiting tourist spots and cultural facilities. This system aims to provide convenience to people with visual or hearing impairments, as well as foreign tourists.

[0545] System Configuration

[0546] The system mainly consists of the following components:

[0547] 1. User device: A mobile device such as a smartphone or smart glasses.

[0548] 2. Server: A server containing a tourist destination database.

[0549] 3. GPS module: A location information acquisition device for obtaining the user's current location.

[0550] 4. Communication means: A network connection for sending and receiving data between the user terminal and the server.

[0551] 5. Audio guide generation function: Software for converting tourist information into audio guides.

[0552] 6. Vibration function: Notification of important tourist spots.

[0553] 7. Text display function: A means of displaying text information for the hearing impaired.

[0554] Program processing

[0555] 1. Obtaining the current location

[0556] The user terminal periodically uses the GPS module to obtain the user's current location.

[0557] The acquired location information (latitude and longitude) is stored in the device and sent to the server.

[0558] 2. Search for tourist information

[0559] The server analyzes the received location information and searches a database for information on the nearest tourist attractions and cultural facilities.

[0560] 3. Sending and receiving tourist destination information

[0561] The server transmits the searched tourist spot information to the user terminal.

[0562] The user terminal parses the received information and converts it into an appropriate format.

[0563] 4. Generating and Providing Audio Guides

[0564] The user terminal converts the received text information into audio guidance using a speech synthesis engine.

[0565] The audio guide is played through headphones in the smart glasses worn by the user.

[0566] 5. Vibration notifications

[0567] When the user approaches an important point in a tourist destination, the user's device will notify them using a vibration function.

[0568] 6. Providing text information

[0569] For the hearing impaired, the server sends tourist information in text format to the terminal.

[0570] The user terminal displays text information on the smart glasses display.

[0571] 7. Automatic translation function

[0572] For foreign visitors to Japan, the server automatically translates tourist information and generates audio guides in multiple languages.

[0573] The translated audio guide is sent to the user's terminal and played back in the specified language.

[0574] Specific examples

[0575] Example 1: Visiting tourist spots in Kyoto (Japanese user)

[0576] The user arrives at Kyoto Station and starts the system.

[0577] The device uses GPS to obtain its current location (Kyoto Station) and sends it to the server.

[0578] The server searches a database for information about tourist attractions around Kyoto Station and sends it to the terminal.

[0579] The device converts the received tourist information into an audio guide, stating, "This is Kyoto Station, located in the center of Kyoto."

[0580] When the user approaches an important point on their way to Kyoto Tower, the device will vibrate to notify them.

[0581] Example 2: Sightseeing for foreign visitors (English users)

[0582] A foreign user arrives at Todaiji Temple in Nara and starts up the system.

[0583] The device obtains its current location and sends it to the server.

[0584] The server retrieves information about Todaiji Temple from a database and generates a multilingual audio guide.

[0585] The device plays the received English audio guide, stating, "This is the Great Buddha Hall of Todaiji Temple, built in the Nara period."

[0586] This system provides users with appropriate tourist information in real time, enriching their tourist experience. It is also suitable for people with visual and hearing impairments, as well as foreign tourists, and helps them gain a deeper understanding of the history and culture of the places they visit.

[0587] The processing flow will be explained below.

[0588] Step 1:

[0589] When the user arrives at a tourist destination, they launch the application on their device (smartphone or smart glasses).

[0590] Step 2:

[0591] The device uses GPS to obtain the user's current location (latitude and longitude), which is temporarily stored in the device.

[0592] Step 3:

[0593] The device sends the acquired location information to a server via the network. The transmitted data includes the user ID and device ID.

[0594] Step 4:

[0595] The server analyzes the received location information and uses it to search a database for information on the nearest tourist attractions and cultural facilities.

[0596] Step 5:

[0597] The server organizes the tourist destination information obtained as search results and returns it to the user's device. The returned data includes the tourist destination's description text, audio URL, and location information of important points.

[0598] Step 6:

[0599] The user terminal analyzes the received tourist destination information and extracts text data to be converted into machine voice.

[0600] Step 7:

[0601] The device's built-in voice synthesis engine converts text data into voice data and provides audio guidance through the smart glasses' headphones.

[0602] Step 8:

[0603] As the user continues to explore the tourist spot, the device compares the current location with the location information of important points, and notifies the user with a vibration function when the user approaches an important point.

[0604] Step 9:

[0605] For the hearing impaired, the server sends tourist information in text format to the user's terminal.

[0606] Step 10:

[0607] The user terminal displays the received text information on the smart glasses display, providing a visual guide.

[0608] Step 11:

[0609] For foreign users, the server detects the user's language setting and automatically translates tourist information.

[0610] Step 12:

[0611] The translated tourist information is then sent back to the device, where a multilingual audio guide is generated and played back in the specified language.

[0612] Through the above process, users can receive detailed tourist information in real time by voice or text, improving their travel experience.

[0613] Example 1

[0614] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0615] Conventional tourist guide systems are limited in the information they provide based on the user's current location, and therefore lack the ability to provide detailed, real-time tourist information in multiple languages ​​or support people with visual or hearing impairments. Furthermore, they lack the ability to notify users when they are approaching important tourist spots, limiting the user's sightseeing experience.

[0616] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0617] In this invention, the server includes means for acquiring the user's current location, means for searching for related tourist attraction information from a tourist attraction database based on the acquired location information, means for transmitting the searched tourist attraction information to the user terminal, means for generating and playing an audio guide based on the transmitted tourist attraction information, means for using a speech synthesis engine to play the audio guide, means for notifying the user by vibration of important points near the user's destination, means for displaying tourist attraction information in text format for the hearing impaired, and means for automatically translating the tourist attraction information into multiple languages. This allows the user to obtain detailed tourist information in real time, enabling a comprehensive tourist experience for people with visual or hearing impairments. Furthermore, the user can receive vibration notifications when approaching important points at the destination, preventing them from missing important tourist attractions.

[0618] Below are definitions of each important word.

[0619] The "current location of the user" is specific information of geographical latitude and longitude obtained using the terminal held by the user.

[0620] A "tourist destination database" is a collection of information that holds tourist information about specific areas and facilities, and is used for searches.

[0621] "Related tourist destination information" is detailed information about tourist destinations that is searched from a database based on the user's current location.

[0622] The "means for transmitting to the user terminal" is a communication method for transmitting tourist destination information in data format from the server to the user terminal.

[0623] The "means for generating and playing back audio guides" refers to a technology for converting tourist destination information into audio data and playing back that audio data for the user.

[0624] "Speech synthesis engine" is a general term for software and hardware used to convert text data into speech data.

[0625] "Means for notifying by vibration at important points" is a function that causes the device to vibrate to notify the user when the user approaches an important location in a specific tourist spot.

[0626] "Means for displaying tourist destination information in text format for the hearing impaired" is a function that provides information to hearing impaired users by visually displaying audio information in text.

[0627] "Means for automatically translating tourist destination information into multiple languages" refers to technology for converting information related to tourist destinations into different languages ​​and providing that information to users.

[0628] The present invention relates to an audio guide system that provides detailed information in real time when visiting tourist spots and cultural facilities. This system aims to provide convenience to people with visual or hearing impairments, as well as foreign tourists.

[0629] System Configuration

[0630] The system mainly consists of the following components:

[0631] 1. User device: A mobile device such as a smartphone or smart glasses.

[0632] 2. Server: A server containing a tourist destination database.

[0633] 3. GPS module: A location information acquisition device for obtaining the user's current location.

[0634] 4. Communication means: A network connection for sending and receiving data between the user terminal and the server.

[0635] 5. Audio guide generation function: Software for converting tourist information into audio guides (e.g., Google Text-to-Speech API).

[0636] 6. Vibration function: Notification of important tourist spots.

[0637] 7. Text display function: A means of displaying text information for the hearing impaired.

[0638] 8. Automatic translation function: Software for translating tourist information into different languages ​​(e.g. Microsoft Translator API).

[0639] Program Overview

[0640] The program for this voice guidance system performs the following processes.

[0641] Get current location:

[0642] The user device periodically acquires the user's current location using the GPS module. The acquired location information (latitude and longitude) is stored in the device and sent to the server.

[0643] Search for tourist information:

[0644] The server analyzes the received location information and searches a database for information on the nearest tourist attractions and cultural facilities. It queries the tourist attractions database using SQL queries and stores the results in a data structure (e.g., a list or dictionary).

[0645] Sending and receiving tourist destination information:

[0646] The server sends the searched tourist destination information to the user's terminal, which analyzes the received information and converts it into an appropriate format.

[0647] Generate and provide audio descriptions:

[0648] The user device converts the received text information into audio guidance using a speech synthesis engine, which is then played through the headphones in the smart glasses worn by the user.

[0649] Vibration notification:

[0650] When the user approaches an important point in a tourist destination, the user's device will notify them using a vibration function.

[0651] Text information provided:

[0652] For the hearing impaired, the server sends tourist information in text format to the device, which then displays the text on the smart glasses display.

[0653] Machine translation feature:

[0654] For foreign visitors to Japan, the server automatically translates tourist information and generates multilingual audio guides. The translated audio guides are sent to the user's device and played in the specified language.

[0655] Specific examples

[0656] Example 1: Visiting tourist spots in Kyoto (Japanese user)

[0657] 1. The user arrives at Kyoto Station and starts the system.

[0658] 2. The device uses GPS to obtain its current location (Kyoto Station) and sends it to the server.

[0659] 3. The server searches the database for information about tourist spots around Kyoto Station and sends it to the terminal.

[0660] 4. The device converts the received tourist information into an audio guide, stating, "This is Kyoto Station, located in the center of Kyoto."

[0661] 5. When the user approaches an important point on the way to Kyoto Tower, the device will vibrate to notify them.

[0662] Example 2: Sightseeing for foreign visitors (English users)

[0663] 1. A foreign user arrives at Todaiji Temple in Nara and starts up the system.

[0664] 2. The device obtains its current location and sends it to the server.

[0665] 3. The server retrieves information about Todaiji Temple from the database and generates a multilingual audio guide.

[0666] 4. The device plays the received English audio guide, stating, "This is the Great Buddha Hall of Todaiji Temple, built in the Nara period."

[0667] Prompt Sentence Examples

[0668] "Please explain the main tourist spots around Kyoto Station."

[0669] "Please tell me in English about the history and highlights of Todaiji Temple."

[0670] This system provides users with appropriate tourist information in real time, enriching their tourist experience. It is also suitable for people with visual and hearing impairments, as well as foreign tourists, and helps them gain a deeper understanding of the history and culture of the places they visit.

[0671] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0672] Step 1: Get current location

[0673] The user device uses a GPS module to obtain the user's current location. Specifically, the device's GPS module receives satellite signals and calculates latitude and longitude information. The obtained location information (e.g., latitude 35.0116, longitude 135.7681) is stored in internal memory and sent to the server. The input is satellite signal data from the GPS module, and the output is latitude and longitude coordinate information.

[0674] Step 2: Search for tourist information

[0675] The server analyzes the latitude and longitude location information received from the user device. Based on the received data, it uses an SQL query to search the tourist spot database and retrieves the relevant tourist spot information. For example, the query executed is "SELECT FROM tourist_spots WHERE latitude BETWEEN 34.9116 AND 35.1116 AND longitude BETWEEN 135.6681 AND 135.8681;". The input is latitude and longitude location information, and the output is a list of related tourist spot information.

[0676] Step 3: Submit tourist destination information

[0677] The server formats the tourist spot information obtained from the search into JSON format and sends it to the user's device over the network. For example, the JSON data sent is "{"spots": [{"name": "Kyoto Station", "description": "Central location in Kyoto"}]}". The input is a list of tourist spot information, and the output is the JSON data sent to the user's device.

[0678] Step 4: Receiving and analyzing tourist destination information

[0679] The user device parses the tourist attraction information in JSON format received from the server. Specifically, it uses a JSON parser to extract the data and saves it in its internal memory. For example, the information parsed is "{"name": "Kyoto Station", "description": "Central location in Kyoto"}". The input is the JSON data received from the server, and the output is the parsed tourist attraction information.

[0680] Step 5: Generate audio guide

[0681] The user device sends the received text information to a speech synthesis engine (e.g., Google Text-to-Speech API) and converts it into audio data. Specifically, it converts the text "This is Kyoto Station, located in the center of Kyoto" into an audio file. The input is the tourist destination information text, and the output is an audio file.

[0682] Step 6: Play the audio guide

[0683] The user device passes the generated audio file to the playback function, which plays the audio through the headphones of the user's smart glasses. This allows the user to hear the audio guidance, "This is Kyoto Station, located in the center of Kyoto." The input is the audio file, and the output is the audio guidance playback to the user.

[0684] Step 7: Vibration notifications at key points

[0685] The user device monitors the latitude and longitude of the current location and the tourist spot, and activates vibration when the distance is within a certain range. For example, the device vibrates when the user approaches Kyoto Tower. The input is the current location and the location information of the tourist spot, and the output is the activation of vibration.

[0686] Step 8: Displaying text information for the hearing impaired

[0687] The server sends tourist destination information in text format to the user's device. The user's device then calls an API to display the received text information on the smart glasses' display. For example, the text information displayed is "This is Kyoto Station, located in the center of Kyoto." The input is tourist destination information in text format, and the output is the text information displayed on the smart glasses' display.

[0688] Step 9: Automatic translation function

[0689] The server uses the Microsoft Translator API to automatically translate tourist information for foreign visitors to Japan. For example, it translates the Japanese phrase "This is Kyoto Station, located in the center of Kyoto" into English. The translated text is sent to a speech synthesis engine, which generates an audio guide. The input is the tourist information text, and the output is the translated audio guide.

[0690] In each processing step, detailed data processing and operations are coordinated to realize a system that provides appropriate tourist information to users.

[0691] (Application example 1)

[0692] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0693] In modern commercial facilities, the provision of information to diverse customers, such as users with visual or hearing impairments and foreign visitors to Japan who require multilingual support, is insufficient. In particular, it is difficult to provide real-time information in stores, such as product information and sales floor guidance, which reduces convenience for users. Given this background, it is necessary to provide consistent information to users with different disabilities and languages.

[0694] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0695] In this invention, the server includes means for acquiring the user's current location, means for searching for related tourist attraction information from a tourist attraction database based on the acquired location information, means for transmitting the searched tourist attraction information to the user terminal, means for generating and playing an audio guide based on the transmitted tourist attraction information, means for acquiring in-store location information, means for searching for related product information based on the acquired in-store location information, means for transmitting the searched product information to the user terminal, and means for generating and playing an audio guide based on the transmitted product information. This makes it possible to provide consistent real-time information to different customers, such as users with visual or hearing impairments and foreign visitors to Japan.

[0696] The "means for acquiring the user's current location" refers to a technical device for measuring and acquiring the latitude and longitude of the user's current location or a specific point within the store.

[0697] "Location information" is data indicating the latitude and longitude of the user's current location or a specific point within a store.

[0698] A "tourist destination database" is a collection of data that systematically stores information about tourist destinations.

[0699] The "means for searching tourist attraction information" is a technical device for searching for relevant information from a tourist attraction database based on the acquired location information.

[0700] A "user terminal" is an electronic device carried by a user, such as a smartphone or smart glasses.

[0701] "Means for generating and playing audio guides" refers to technical devices that convert text information into audio and play it back through the user terminal.

[0702] "Means for obtaining location information within the store" refers to a technical device that uses beacons, Wi-Fi, or other location measurement technologies to identify and obtain the user's specific location within the store.

[0703] The "means for searching for related product information" is a technical device for searching the database for the most suitable product information based on the acquired in-store location information.

[0704] The "vibration function" is a technical device that causes the user's terminal to vibrate to notify the user when the user approaches a specific location.

[0705] "Means for displaying in text format" refers to a technical device for displaying information as a string of characters on the display of a user terminal for the hearing impaired.

[0706] "Universal service" is a service that is consistently provided to all users, regardless of whether they have a disability or speak a different language.

[0707] The present invention is a voice guide system that provides users with detailed information in real time at tourist spots and commercial facilities, with the aim of providing convenience to people with visual or hearing impairments, as well as foreign tourists and shoppers.

[0708] System Configuration

[0709] The system mainly consists of the following components:

[0710] 1. User device: A mobile device such as a smartphone or smart glasses.

[0711] 2. Server: A server containing a tourist destination and product database.

[0712] 3. GPS module: A location information acquisition device for obtaining the user's current location.

[0713] 4. Communication means: A network connection for sending and receiving data between the user terminal and the server.

[0714] 5. Audio guide generation function: Software for converting tourist destination information and product information into audio guides.

[0715] 6. Vibration function: A means of notifying you of important tourist spots and product areas.

[0716] 7. Text display function: A means of displaying text information for the hearing impaired.

[0717] 8. In-store location information acquisition means: A device that acquires the user's location within the store using beacons or Wi-Fi access points.

[0718] Basic operation

[0719] The operation of the system is as follows.

[0720] The user terminal periodically acquires its current location using a GPS module or location information acquisition means within the store. This location information is stored in the user terminal and sent to the server.

[0721] The server analyzes the received location information and searches a database for information on the nearest tourist spots and products, and sends the searched information to the user's device in real time.

[0722] The user device analyzes the received information, converts it into an appropriate format, and generates audio guidance using a speech synthesis engine (e.g., Google Text-to-Speech API or Amazon Polly) that is then played through the user's headphones.

[0723] When the user approaches an important tourist spot or product area, the user's device will notify them using the vibration function.

[0724] For the hearing impaired, the server sends tourist destination and product information in text format to the user's device, and the text information is displayed on the display of smart glasses or a smartphone.

[0725] For foreign visitors to Japan, the server automatically translates tourist destination and product information, generates multilingual audio guides, and sends them to the user's device.

[0726] Specific examples

[0727] Example of a tourist destination application

[0728] When a user visits a tourist spot, they can activate this system, which will provide audio guidance with detailed tourist spot information and historical background based on their current location. For example, when a user arrives at Kyoto Station and activates the system, the device obtains their current location (Kyoto Station) and sends it to the server. The server searches a tourist spot database for information about tourist spots around Kyoto Station and sends it to the device. The device then converts this into an audio guide, stating, "This is Kyoto Station, located in the center of Kyoto."

[0729] Example of a physical store guidance application

[0730] When a user activates this system while shopping in a commercial facility, it acquires location information within the store and provides voice guidance on related product information. For example, if the user is in the electronics section of a department store, the device acquires the user's current location and sends it to the server. The server then searches for related product information from a product database and sends it to the device. The device then converts this into voice guidance, stating, "These are the latest earphone models." Additionally, when the user approaches a specific product area, the device vibrates to notify the user.

[0731] Prompt Sentence Examples

[0732] "I want to provide detailed product information based on the visitor's current location, tailored to their nationality and language preferences. The location information is coordinates (X: 100, Y: 200), and the product information is "sensitive earphones" in Japanese. Please translate this into English and generate an audio guide."

[0733] This will enable consistent, real-time information to be provided to different customers, including users with visual or hearing impairments and foreign visitors to Japan.

[0734] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0735] Step 1:

[0736] Get current location

[0737] The user device periodically acquires the user's current location using a GPS module or in-store location information acquisition means (beacons or Wi-Fi). Specifically, when the user starts the system, the GPS module acquires latitude and longitude information, or receives signals from in-store beacons to identify the X and Y coordinates. This generates current location information (input) and stores it in the user device (output).

[0738] Step 2:

[0739] Sending current location information

[0740] The user device sends the acquired current location information to the server. Specifically, it uses a network communication method (Wi-Fi or mobile data communication) to send the location information data to the server. As a result, the location information data from the user device (input) is sent to the server (output).

[0741] Step 3:

[0742] Search for tourist information or product information

[0743] The server analyzes the received location information and searches for related tourist attraction information or product information from a tourist attraction database or product database. Specifically, it uses a database search algorithm to extract the nearest tourist attraction or product information from the location information (input). This process generates related tourist attraction information and product information (output).

[0744] Step 4:

[0745] Sending related information

[0746] The server sends the searched tourist spot information or product information to the user terminal. Specifically, the tourist spot information or product information (input) is sent to the user terminal via the network. This operation allows the tourist spot information or product information from the server to reach the user terminal (output).

[0747] Step 5:

[0748] Audio guide generation and provision

[0749] The user device analyzes the received information and converts it into an appropriate format. Specifically, it uses a speech synthesis engine (Google Text-to-Speech API or Amazon Polly) to convert text information (input) into voice data. The converted voice guide (output) is played through the user device's headphones.

[0750] Step 6:

[0751] Vibration notification

[0752] When a user approaches a particular tourist attraction or product area, the user's device will vibrate to notify them. This is done by calculating the distance to each set point based on the device's location data (input), and vibrating when the distance falls below a certain level (output).

[0753] Step 7:

[0754] Providing text information

[0755] For the hearing impaired, the server sends tourist destination information and product information in text format to the user's device. The user's device then displays this information on its display. Specifically, the text data (input) is screen-rendered for display. This provides the user with visual text information (output).

[0756] Step 8:

[0757] Machine translation and multilingual support

[0758] The server automatically translates tourist destination information and product information to generate multilingual audio guides for foreign visitors to Japan. Specifically, it translates text information (input) into the specified language using a multilingual translation model (e.g., Google Translate API). This translated text information is then converted into audio guides using a speech synthesis engine and sent to the user's device. This provides a translated multilingual audio guide (output).

[0759] Specific example prompts

[0760] "I want to provide detailed product information based on the visitor's current location, tailored to their nationality and language preferences. The location information is coordinates (X: 100, Y: 200), and the product information is "sensitive earphones" in Japanese. Please translate this into English and generate an audio guide."

[0761] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0762] The present invention relates to an audio guide system that provides detailed information in real time when visiting tourist spots and cultural facilities. This system aims to provide convenience especially for people with visual or hearing impairments and foreign tourists. Furthermore, the present invention realizes a personalized sightseeing experience by recognizing the user's emotional state and providing information accordingly.

[0763] System Configuration

[0764] The system mainly consists of the following components:

[0765] 1. User device: A mobile device such as a smartphone or smart glasses.

[0766] 2. Server: A server containing a tourist destination database.

[0767] 3. GPS module: A location information acquisition device for obtaining the user's current location.

[0768] 4. Communication means: A network connection for sending and receiving data between the user terminal and the server.

[0769] 5. Audio guide generation function: Software for converting tourist information into audio guides.

[0770] 6. Vibration function: Notification of important tourist spots.

[0771] 7. Text display function: A means of displaying text information for the hearing impaired.

[0772] 8. Emotion engine: Software that recognizes emotional states and dynamically changes the content and format of the tourist destination information provided.

[0773] Program processing

[0774] 1. Obtaining the current location

[0775] The user device periodically acquires the user's current location (latitude and longitude) using the GPS module. This location information is temporarily stored in the device and then sent to the server.

[0776] 2. Search for tourist information

[0777] The server analyzes the received location information and uses it to search a database for information on the nearest tourist attractions and cultural facilities.

[0778] 3. Sending and receiving tourist destination information

[0779] The server organizes the tourist destination information obtained as search results and returns it to the user's device. The returned data includes the tourist destination's description text, audio URL, and location information of important points.

[0780] 4. Generating and Providing Audio Guides

[0781] The user device converts the received text information into audio guidance using a speech synthesis engine, which is then played through the headphones in the smart glasses worn by the user.

[0782] 5. Vibration notifications

[0783] When the user approaches an important point in a tourist destination, the user's device will notify them using a vibration function.

[0784] 6. Providing text information

[0785] For the hearing impaired, the server sends tourist information in text format to the user's device, which then displays the received text information on the smart glasses display, providing a visual guide.

[0786] 7. Automatic translation function

[0787] For foreign visitors to Japan, the server detects the user's language preference and automatically translates tourist information. The translated tourist information is then sent back to the device, which generates a multilingual audio guide. The device then plays the audio guide in the specified language.

[0788] 8. Emotion recognition and personalized information delivery

[0789] The emotion engine installed in the user device uses cameras and biometric sensors to recognize the user's emotional state from their facial expressions and biometric data. For example, if the user is excited, the emotion engine will recognize this and prioritize providing information on particularly attractive tourist spots and activity recommendations.

[0790] If the user is under stress, the system will provide information on relaxing spots and quiet places. By providing appropriate information according to the user's emotional state, the system can improve the user's satisfaction.

[0791] Specific examples

[0792] Example 1: Visiting tourist spots in Kyoto (Japanese user)

[0793] The user arrives at Kyoto Station and starts the system. The device uses GPS to obtain its current location (Kyoto Station) and sends it to the server. The server then searches a database for information about tourist spots around Kyoto Station and sends it to the device. The device converts the received tourist spot information into an audio guide, providing instructions such as, "This is Kyoto Station, located in the center of Kyoto." As the user approaches important points on their way to Kyoto Tower, the device vibrates to notify them. If the emotion engine detects the user's excitement, it provides additional information such as, "You can see a wonderful night view from the observation deck of Kyoto Tower."

[0794] Example 2: Sightseeing for foreign visitors (English users)

[0795] A foreign user arrives at Todaiji Temple in Nara and starts the system. The device obtains its current location and sends it to the server. The server retrieves information about Todaiji Temple from a database and generates a multilingual audio guide. The device plays the received English audio guide, stating, "This is the Great Buddha Hall of Todaiji Temple, built in the Nara period." If the emotion engine detects the user's stress level, it also provides information such as, "There is a quiet, relaxing park nearby."

[0796] This system can improve the travel experience by providing users with appropriate tourism information in real time via voice or text. Furthermore, the emotion engine enables information to be provided according to the user's emotional state, realizing a more personalized and satisfying tourism experience.

[0797] The processing flow will be explained below.

[0798] Step 1:

[0799] When the user arrives at a tourist destination, they launch the application on their device (smartphone or smart glasses).

[0800] Step 2:

[0801] The device uses GPS to obtain the user's current location (latitude and longitude), which is temporarily stored in the device.

[0802] Step 3:

[0803] The device sends the acquired location information to a server via the network. The transmitted data includes the user ID and device ID.

[0804] Step 4:

[0805] The server analyzes the received location information and uses it to search a database for information on the nearest tourist attractions and cultural facilities.

[0806] Step 5:

[0807] The server organizes the tourist destination information obtained as search results and returns it to the user's device. The returned data includes the tourist destination's description text, audio URL, and location information of important points.

[0808] Step 6:

[0809] The user terminal analyzes the received tourist destination information and extracts text data to be converted into machine voice.

[0810] Step 7:

[0811] The device's built-in voice synthesis engine converts text data into voice data and provides audio guidance through the smart glasses' headphones.

[0812] Step 8:

[0813] As the user continues to explore the tourist spot, the device compares the current location with the location information of important points, and notifies the user with a vibration function when the user approaches an important point.

[0814] Step 9:

[0815] For the hearing impaired, the server sends tourist information in text format to the user's terminal.

[0816] Step 10:

[0817] The user terminal displays the received text information on the smart glasses display, providing a visual guide.

[0818] Step 11:

[0819] For foreign users, the server detects the user's language setting and automatically translates tourist information.

[0820] Step 12:

[0821] The translated tourist information is then sent back to the device, where a multilingual audio guide is generated and played back in the specified language.

[0822] Step 13:

[0823] The emotion engine installed in the user terminal uses a camera and biometric sensors to recognize the user's emotional state from their facial expressions and biometric data.

[0824] Step 14:

[0825] The device transmits the recognized emotional state to the server, which then selects tourist destination information according to the emotion and retransmits it to the user device.

[0826] Step 15:

[0827] If the emotion engine detects the user's excited state, the server will prioritize providing information on tourist spots and activity guides that are more likely to interest the user.

[0828] Step 16:

[0829] If the emotion engine detects the user's stress state, the server will provide information on relaxation spots and quiet places.

[0830] Step 17:

[0831] The user terminal provides the newly received tourist spot information to the user by voice or text.

[0832] Example 2

[0833] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0834] Conventional tourist guide systems do not provide sufficient convenience for people with visual or hearing impairments or foreign tourists. Furthermore, they do not provide customized information based on the user's emotional state, making it difficult to realize a personalized tourist experience. Furthermore, they lack multilingual support, making it difficult to provide information appropriately to users with different cultural backgrounds.

[0835] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring the user's current location, means for searching for related tourist attraction information from a tourist attraction database based on the acquired location information, means for transmitting the searched tourist attraction information to the user terminal, means for generating and playing back an audio guide based on the transmitted tourist attraction information, means for recognizing the user's emotional state and dynamically changing the content and format of the tourist attraction information based on the recognition result, means for automatically translating the tourist attraction information and providing a multilingual audio guide, and means for displaying the tourist attraction information in text format for the hearing impaired. This provides convenience for people with visual or hearing impairments and foreign travelers, enables the provision of customized information according to the user's emotional state, and makes it possible to provide appropriate information to users with different cultural backgrounds by supporting multiple languages.

[0836] The "means for acquiring the user's current location" is a function for acquiring the user's current location using a GPS module or other location information acquisition device.

[0837] The "means for searching for related tourist attraction information from a tourist attraction database based on the acquired location information" is a function in which the server analyzes the location information received and uses that information to search for information on the nearest tourist attractions and cultural facilities from the tourist attraction database.

[0838] "Means for sending searched tourist spot information to user terminal" refers to the function by which the server returns tourist spot information obtained as search results to the user terminal. The returned data includes explanatory text of the tourist spot, audio URL, and location information of important points.

[0839] "Means for generating and playing back audio guides based on transmitted tourist attraction information" refers to a function that converts tourist attraction information received by a user terminal into audio guides using a voice synthesis engine and plays them back.

[0840] "Means for recognizing the user's emotional state and dynamically changing the content and format of tourist destination information based on the recognition results" refers to a function in which the emotion engine installed in the device uses the camera and biometric sensors to recognize the user's emotional state and customizes the information provided based on the results.

[0841] "Means for automatically translating tourist destination information and providing multilingual audio guides" refers to a function in which the server detects the user's set language, translates the tourist destination information into the corresponding language, and generates multilingual audio guides based on that translation information.

[0842] The "means for displaying tourist spot information in text format for the hearing impaired" is a function in which the server transmits tourist spot information in text format to the user terminal, and the user terminal displays it visually.

[0843] This invention relates to an audio guide system that provides detailed information in real time when visiting tourist spots and cultural facilities. This system aims to provide convenience especially for people with visual or hearing impairments, as well as foreign tourists. Furthermore, it realizes a personalized sightseeing experience by recognizing the user's emotional state and providing information according to it.

[0844] System Configuration

[0845] The system mainly consists of the following components:

[0846] 1. User device: Use a mobile device such as a smartphone or smart glasses.

[0847] 2. Server: A server that contains a tourist destination database, and the database uses MySQL or similar.

[0848] 3. GPS module: Includes a location information acquisition device for acquiring the user's current location.

[0849] 4. Communication method: A network connection for sending and receiving data between the user device and the server. Wi-Fi, 4G / 5G, etc. are used.

[0850] 5. Audio guide generation function: Google Cloud Text-to-Speech is used as software to convert tourist information into audio guides.

[0851] 6. Vibration function: The device's vibration function is used to notify users of important tourist spots.

[0852] 7. Text display function: Equipped with a display to display text information for the hearing impaired.

[0853] 8. Emotion engine: Cameras and biometric sensors are used as software to recognize emotional states and dynamically change the content and format of tourist information provided.

[0854] Program processing explanation

[0855] Get current location

[0856] The user device periodically acquires the user's current location (latitude and longitude) using the GPS module. This location information is temporarily stored in the device and then sent to the server.

[0857] Search for tourist information

[0858] The server analyzes the received location information and uses it to search a database for information on the nearest tourist attractions and cultural facilities.

[0859] Sending and receiving tourist information

[0860] The server organizes the tourist destination information obtained as search results and returns it to the user's device. The returned data includes the tourist destination's description text, audio URL, and location information of important points.

[0861] Generating and providing audio guides

[0862] The user device converts the received text information into audio guidance using a speech synthesis engine, which is then played through the headphones in the smart glasses worn by the user.

[0863] Vibration notification

[0864] When the user approaches an important point in a tourist destination, the user's device will notify them using a vibration function.

[0865] Providing text information

[0866] For the hearing impaired, the server sends tourist information in text format to the user's device, which then displays the received text information on the smart glasses display, providing a visual guide.

[0867] Machine translation function

[0868] For foreign visitors to Japan, the server detects the user's language preference and automatically translates tourist information. The translated tourist information is then sent back to the device, which generates a multilingual audio guide. The device then plays the audio guide in the specified language.

[0869] Emotion recognition and customized information provision

[0870] The emotion engine installed in the user device uses cameras and biometric sensors to recognize the user's emotional state from their facial expressions and biometric data. For example, if the user is excited, the emotion engine will recognize this and provide them with information on particularly attractive tourist spots and activity recommendations first. If the user is stressed, it will provide them with information on relaxing spots and quiet places. Providing appropriate information according to the user's emotional state can improve user satisfaction.

[0871] Specific examples

[0872] Example 1: Visiting tourist spots in Kyoto (Japanese user)

[0873] The user arrives at Kyoto Station and starts the system. The device uses GPS to obtain its current location (Kyoto Station) and sends it to the server. The server then searches a database for information about tourist spots around Kyoto Station and sends it to the device. The device converts the received tourist spot information into an audio guide, providing instructions such as, "This is Kyoto Station, located in the center of Kyoto." As the user approaches important points on their way to Kyoto Tower, the device vibrates to notify them. If the emotion engine detects the user's excitement, it provides additional information such as, "You can see a wonderful night view from the observation deck of Kyoto Tower."

[0874] Example 2: Sightseeing for foreign visitors (English users)

[0875] A foreign user arrives at Todaiji Temple in Nara and starts the system. The device obtains its current location and sends it to the server. The server retrieves information about Todaiji Temple from a database and generates a multilingual audio guide. The device plays the received English audio guide, stating, "This is the Great Buddha Hall of Todaiji Temple, built in the Nara period." If the emotion engine detects the user's stress level, it also provides information such as, "There is a quiet, relaxing park nearby."

[0876] Prompt Sentence Examples

[0877] As an example of an input prompt for a generative AI model, you could enter something like:

[0878] "Please explain in detail the process of the system that, when the user arrives at a tourist spot, acquires the user's current location using the smartphone, acquires tourist spot information from the server, and customizes the information according to the user's emotional state."

[0879]

[0880] "Describe the detailed process by which a real-time audio guide system uses a smartphone to acquire the current location, fetch tourist information from a server, and customize the information based on the user's emotional state."

[0881] This system can improve the travel experience by providing users with appropriate tourism information in real time via voice or text. Furthermore, the emotion engine enables information to be provided according to the user's emotional state, realizing a more personalized and satisfying tourism experience.

[0882] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0883] Step 1: Get current location

[0884] Input: The user launches the app and the GPS module obtains the current location.

[0885] Specific operation: A user arrives at a tourist destination and launches the app on their smartphone or smart glasses. The device immediately accesses the built-in GPS module to obtain the user's current location (latitude and longitude). The message "Retrieving user's current location..." is displayed.

[0886] Data processing or data calculation: Position information (e.g., 35.0116°N, 135.7681°E) is acquired and temporarily stored in the device's memory.

[0887] Output: The obtained location information is sent to the server.

[0888] Step 2: Search for tourist information

[0889] Input: The server receives the user's location information.

[0890] Specific operation: The server performs a database search based on the location information (latitude and longitude) received. The database stores tourist spot names, descriptions, audio guide URLs, etc.

[0891] Data processing or data calculation: The server analyzes the location information and searches for information on nearby tourist attractions in a database (e.g., MySQL).

[0892] Output: Tourist information (e.g., Kyoto Tower description text, audio URL, location information of important points) is obtained and sent to the user's device.

[0893] Step 3: Sending and receiving tourist destination information

[0894] Input: The server organizes tourist destination information from search results.

[0895] Specific operation: The server formats the tourist attraction information obtained as a search result, packages the necessary data (explanatory text, audio URL, location information of important points, etc.) and sends it to the user's terminal.

[0896] Data processing or data calculation: Convert tourist destination information data into a data structure such as JSON format, making it a format that can be transferred efficiently.

[0897] Output: The formatted tourist destination information is sent to the user's terminal.

[0898] Step 4: Generate and serve audio descriptions

[0899] Input: The user terminal receives tourist destination information from the server.

[0900] Specific operation: The user device passes the received text information to a speech synthesis engine (e.g., Google Cloud Text-to-Speech), which converts the information into audio data. The generated audio guide is played through the smart glasses headphones.

[0901] Data processing or data calculation: Converting text data into audio data and saving it in a playable format.

[0902] Output: The audio description is played to the user.

[0903] Step 5: Vibration Notification

[0904] Input: The user approaches a point of interest at a tourist destination.

[0905] Specific operation: When the user approaches a key point, the device receives location information from the built-in GPS and determines that it is a key point. The device then starts vibrating and notifies the user that "You are approaching a key point."

[0906] Data processing or data calculation: Compare the current location with the location of important points, and activate the vibration function when they come within a certain distance.

[0907] Output: A vibration will be activated to notify the user.

[0908] Step 6: Provide text information

[0909] Input: The server sends tourist information for the hearing impaired in text format.

[0910] Specific operation: The server sends text information to the user's terminal, which then displays it on the smart glasses display.

[0911] Data processing or data calculation: Converting text information into a displayable format and presenting it on a display.

[0912] Output: Text information is displayed on the smart glasses display.

[0913] Step 7: Automatic translation function

[0914] Input: The server detects the user's preferred language.

[0915] Specific operation: The server reads the language setting of the user's device, automatically translates the tourist attraction information based on that setting, and then sends the translated information back to the user's device.

[0916] Data processing or data calculation: Translate tourist destination information into the specified language using an automatic translation tool (e.g., Google Translate API).

[0917] Output: The translated tourist information is sent to the user's terminal, and an audio guide is played in the specified language.

[0918] Step 8: Emotion recognition and customized information delivery

[0919] Input: The user's emotional state is acquired by a camera or biometric sensors installed on the user's device.

[0920] How it works: The emotion engine analyzes the user's emotional state using facial recognition technology, heart rate measurement, etc. For example, if the user is excited, it will prioritize providing "attractive spots."

[0921] Data processing or data calculation: An algorithm is used to analyze the acquired biometric and facial expression data and determine the emotional state.

[0922] Output: Customized information according to the emotional state is provided to the user.

[0923] (Application example 2)

[0924] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0925] The present invention aims to improve convenience for people with visual or hearing impairments and foreign tourists by providing detailed information in real time when visiting tourist spots and cultural facilities through an audio guide system. It also aims to realize a personalized tourist experience by recognizing the user's emotional state and providing information accordingly. Furthermore, this technology can be applied to a factory work support system to provide workers with real-time work guides and important information, and to improve work efficiency and safety by detecting the user's fatigue and stress level and providing appropriate rest instructions.

[0926] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0927] In this invention, the server includes means for acquiring the user's current location, means for searching for related tourist attraction information from a tourist attraction database based on the acquired location information, means for transmitting the searched tourist attraction information to the user terminal, means for generating and playing back an audio guide based on the transmitted tourist attraction information, and means for recognizing the user's emotional state and dynamically changing the content and format of the tourist attraction information to be provided in accordance with that state. This makes it possible to provide the user with appropriate information in real time when visiting tourist attractions or working in a factory, improving the user experience and work efficiency.

[0928] "User's current location" refers to the geographic location of a particular user as determined by a GPS module or other location acquisition function.

[0929] A "tourist destination database" refers to an information system that stores information about tourist destinations and cultural facilities in a structured format.

[0930] "Tourist destination information" refers to detailed information such as descriptions, images, and audio guides about specific tourist destinations and cultural facilities.

[0931] "User terminal" refers to an electronic device that a user can carry around, such as a smartphone or smart glasses.

[0932] "Audio guide" refers to systems and content that provide audio information about tourist destinations and cultural facilities.

[0933] "Emotional state" refers to the user's emotional and mental state obtained by analyzing the user's facial expressions and biosensor data.

[0934] The "vibration function" refers to a function that notifies the user by vibrating the user terminal.

[0935] "Text format" refers to a format in which information is displayed in text.

[0936] "Real-time" refers to providing current information and status immediately without delay.

[0937] "Work guide" refers to information on instructions and operating procedures provided when working in a factory.

[0938] "Rest instructions" refer to instructions regarding when and where the user should take a rest.

[0939] "Personalized tourism experience" refers to customized tourism information and guidance provided based on the user's individual circumstances and interests.

[0940] "Multilingual" refers to the ability to provide information and guidance in multiple languages.

[0941] "Factory work support system" refers to technology and platforms that support work within a factory.

[0942] This invention is a voice guide system that provides users with detailed information in real time when visiting tourist spots and cultural facilities. This system aims to provide convenience, especially for people with visual or hearing impairments and foreign tourists. Furthermore, it realizes a personalized tourist experience by recognizing the user's emotional state and providing information accordingly. This technology can also be applied to a factory work support system, providing real-time work guidance and important information to workers, and improving work efficiency and safety by detecting the user's fatigue and stress level and providing appropriate rest instructions.

[0943] System Configuration

[0944] The system consists of the following components:

[0945] 1. User Device:

[0946] These are portable electronic devices such as smartphones and smart glasses that enable users to acquire and display information at tourist spots and in factories.

[0947] 2. Server:

[0948] This server contains a database of tourist attractions and data on factory work. The server provides relevant information based on requests from user terminals.

[0949] 3. GPS module:

[0950] This is a location information acquisition device for acquiring the current location of the user.

[0951] 4. Means of communication:

[0952] A network connection for sending and receiving data between a user terminal and a server.

[0953] 5. Audio guide generation function:

[0954] This is software for converting tourist information into audio guides. A typical example of this software is the TextToSpeech module.

[0955] 6. Vibration function:

[0956] It is a means of notifying important tourist spots and important locations within the factory.

[0957] 7. Text display function:

[0958] This is a way to display tourist information and work guides as text information for the hearing impaired. For this, we use the DisplayTextModule.

[0959] 8. Emotion Engine:

[0960] This software recognizes the user's emotional state and dynamically changes the content and format of the information provided. For example, the EmotionRecognition module can be used to analyze the user's emotional state.

[0961] Operating principle

[0962] The server includes a means for acquiring the user's current location, a means for searching for related information from a tourist spot database or a factory work database based on the acquired location information, a means for transmitting the searched information to the user terminal, a means for generating and playing back an audio guide based on the transmitted information, and a means for recognizing the user's emotional state and dynamically changing the content and format of the information to be provided depending on that state. This makes it possible to provide the user with appropriate information in real time while visiting tourist spots or working in a factory, improving the user experience and work efficiency.

[0963] Specific examples

[0964] Example 1: Factory work support system

[0965] Users wear smart glasses while moving around the factory to perform their work. The system uses a GPS module to obtain their current location, and the server searches for relevant work information based on the obtained location information. The received information is provided to the user as voice guidance and text display. The system also uses an EmotionRecognition module to recognize the user's emotional state, and if fatigue or stress is detected, it provides a vibration function and instructions to take a break.

[0966] Prompt Sentence Examples

[0967] "Please suggest an application program that updates the smart glasses feed in real time and displays work instructions based on the user's current location. Also, include a feature that automatically notifies the user to take a break if they are feeling stressed."

[0968] As described above, the present invention is a system for significantly improving work efficiency at tourist spots and factories. By implementing this system, it is expected that the user experience will be improved and work efficiency will increase.

[0969] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0970] Step 1:

[0971] The user terminal acquires its current location (latitude and longitude) using a GPS module.

[0972] input:

[0973] Location information obtained from the GPS module of the user's smart glasses or smartphone.

[0974] Data processing and calculation:

[0975] Obtain latitude and longitude data of your current location and temporarily store it on the device.

[0976] output:

[0977] Numerical data of latitude and longitude obtained.

[0978] Step 2:

[0979] The terminal transmits information about its current location to the server.

[0980] input:

[0981] Numerical latitude and longitude data obtained in step 1.

[0982] Specific behavior:

[0983] The device transmits GPS data to a server via a network.

[0984] output:

[0985] A request to send location information to the server.

[0986] Step 3:

[0987] The server analyzes the received location information and searches for related information from a tourist destination database or a factory work database.

[0988] input:

[0989] Latitude and longitude data of the current location sent from the device.

[0990] Data processing and calculation:

[0991] Query a database based on location information to find relevant tourist information or work procedures.

[0992] output:

[0993] Associated tourist attraction information and work procedure data.

[0994] Step 4:

[0995] The server organizes the obtained information and returns it to the user terminal.

[0996] input:

[0997] Tourist destination information and work procedure data searched in Step 3.

[0998] Specific behavior:

[0999] The server organizes the information, converts it into a specific format, and then transmits it over the network to the user terminal.

[1000] output:

[1001] Tourist destination information and work procedure data are sent to the device.

[1002] Step 5:

[1003] The user terminal converts the received information into audio guidance using a speech synthesis engine (TextToSpeech module) and plays it back.

[1004] input:

[1005] Received text data of tourist destination information and work procedures.

[1006] Data processing and calculation:

[1007] The process of converting text information into speech.

[1008] output:

[1009] Audio data played as audio guide.

[1010] Specific behavior:

[1011] The device uses the TextToSpeech module to convert the text into speech, which is played through the smart glasses' headphones.

[1012] Step 6:

[1013] When the user approaches an important point in a tourist spot or an important location in a factory, the device will notify them by vibrating.

[1014] input:

[1015] User location and location information of important points for tourist attractions and work procedures.

[1016] Data processing and calculation:

[1017] The current location is compared with the location of an important point and it is detected when it enters a certain range.

[1018] output:

[1019] Vibration notifications.

[1020] Specific behavior:

[1021] The user terminal activates the vibration function to notify the user.

[1022] Step 7:

[1023] The user terminal uses the EmotionRecognition module to recognize the user's emotional state and provide information according to the emotional state.

[1024] input:

[1025] Emotional state data from cameras and biometric sensors (e.g., facial expressions, heart rate).

[1026] Data processing and calculation:

[1027] The emotional state data is analyzed to recognize the emotional state of the user.

[1028] output:

[1029] Providing customized information according to the user's emotional state.

[1030] Specific behavior:

[1031] The EmotionRecognition module assesses the user's emotional state and provides information, including instructions to rest, if the user feels tired, for example.

[1032] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1033] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1034] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1035] [Third embodiment]

[1036] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1037] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1038] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1039] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1040] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1041] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1042] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1043] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1044] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1045] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1046] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1047] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1048] The present invention relates to an audio guide system that provides detailed information in real time when visiting tourist spots and cultural facilities. This system aims to provide convenience to people with visual or hearing impairments, as well as foreign tourists.

[1049] System Configuration

[1050] The system mainly consists of the following components:

[1051] 1. User device: A mobile device such as a smartphone or smart glasses.

[1052] 2. Server: A server containing a tourist destination database.

[1053] 3. GPS module: A location information acquisition device for obtaining the user's current location.

[1054] 4. Communication means: A network connection for sending and receiving data between the user terminal and the server.

[1055] 5. Audio guide generation function: Software for converting tourist information into audio guides.

[1056] 6. Vibration function: Notification of important tourist spots.

[1057] 7. Text display function: A means of displaying text information for the hearing impaired.

[1058] Program processing

[1059] 1. Obtaining the current location

[1060] The user terminal periodically uses the GPS module to obtain the user's current location.

[1061] The acquired location information (latitude and longitude) is stored in the device and sent to the server.

[1062] 2. Search for tourist information

[1063] The server analyzes the received location information and searches a database for information on the nearest tourist attractions and cultural facilities.

[1064] 3. Sending and receiving tourist destination information

[1065] The server transmits the searched tourist spot information to the user terminal.

[1066] The user terminal parses the received information and converts it into an appropriate format.

[1067] 4. Generating and Providing Audio Guides

[1068] The user terminal converts the received text information into audio guidance using a speech synthesis engine.

[1069] The audio guide is played through headphones in the smart glasses worn by the user.

[1070] 5. Vibration notifications

[1071] When the user approaches an important point in a tourist destination, the user's device will notify them using a vibration function.

[1072] 6. Providing text information

[1073] For the hearing impaired, the server sends tourist information in text format to the terminal.

[1074] The user terminal displays text information on the smart glasses display.

[1075] 7. Automatic translation function

[1076] For foreign visitors to Japan, the server automatically translates tourist information and generates audio guides in multiple languages.

[1077] The translated audio guide is sent to the user's terminal and played back in the specified language.

[1078] Specific examples

[1079] Example 1: Visiting tourist spots in Kyoto (Japanese user)

[1080] The user arrives at Kyoto Station and starts the system.

[1081] The device uses GPS to obtain its current location (Kyoto Station) and sends it to the server.

[1082] The server searches a database for information about tourist attractions around Kyoto Station and sends it to the terminal.

[1083] The device converts the received tourist information into an audio guide, stating, "This is Kyoto Station, located in the center of Kyoto."

[1084] When the user approaches an important point on their way to Kyoto Tower, the device will vibrate to notify them.

[1085] Example 2: Sightseeing for foreign visitors (English users)

[1086] A foreign user arrives at Todaiji Temple in Nara and starts up the system.

[1087] The device obtains its current location and sends it to the server.

[1088] The server retrieves information about Todaiji Temple from a database and generates a multilingual audio guide.

[1089] The device plays the received English audio guide, stating, "This is the Great Buddha Hall of Todaiji Temple, built in the Nara period."

[1090] This system provides users with appropriate tourist information in real time, enriching their tourist experience. It is also suitable for people with visual and hearing impairments, as well as foreign tourists, and helps them gain a deeper understanding of the history and culture of the places they visit.

[1091] The processing flow will be explained below.

[1092] Step 1:

[1093] When the user arrives at a tourist destination, they launch the application on their device (smartphone or smart glasses).

[1094] Step 2:

[1095] The device uses GPS to obtain the user's current location (latitude and longitude), which is temporarily stored in the device.

[1096] Step 3:

[1097] The device sends the acquired location information to a server via the network. The transmitted data includes the user ID and device ID.

[1098] Step 4:

[1099] The server analyzes the received location information and uses it to search a database for information on the nearest tourist attractions and cultural facilities.

[1100] Step 5:

[1101] The server organizes the tourist destination information obtained as search results and returns it to the user's device. The returned data includes the tourist destination's description text, audio URL, and location information of important points.

[1102] Step 6:

[1103] The user terminal analyzes the received tourist destination information and extracts text data to be converted into machine voice.

[1104] Step 7:

[1105] The device's built-in voice synthesis engine converts text data into voice data and provides audio guidance through the smart glasses' headphones.

[1106] Step 8:

[1107] As the user continues to explore the tourist spot, the device compares the current location with the location information of important points, and notifies the user with a vibration function when the user approaches an important point.

[1108] Step 9:

[1109] For the hearing impaired, the server sends tourist information in text format to the user's terminal.

[1110] Step 10:

[1111] The user terminal displays the received text information on the smart glasses display, providing a visual guide.

[1112] Step 11:

[1113] For foreign users, the server detects the user's language setting and automatically translates tourist information.

[1114] Step 12:

[1115] The translated tourist information is then sent back to the device, where a multilingual audio guide is generated and played back in the specified language.

[1116] Through the above process, users can receive detailed tourist information in real time by voice or text, improving their travel experience.

[1117] Example 1

[1118] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1119] Conventional tourist guide systems are limited in the information they provide based on the user's current location, and therefore lack the ability to provide detailed, real-time tourist information in multiple languages ​​or support people with visual or hearing impairments. Furthermore, they lack the ability to notify users when they are approaching important tourist spots, limiting the user's sightseeing experience.

[1120] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1121] In this invention, the server includes means for acquiring the user's current location, means for searching for related tourist attraction information from a tourist attraction database based on the acquired location information, means for transmitting the searched tourist attraction information to the user terminal, means for generating and playing an audio guide based on the transmitted tourist attraction information, means for using a speech synthesis engine to play the audio guide, means for notifying the user by vibration of important points near the user's destination, means for displaying tourist attraction information in text format for the hearing impaired, and means for automatically translating the tourist attraction information into multiple languages. This allows the user to obtain detailed tourist information in real time, enabling a comprehensive tourist experience for people with visual or hearing impairments. Furthermore, the user can receive vibration notifications when approaching important points at the destination, preventing them from missing important tourist attractions.

[1122] Below are definitions of each important word.

[1123] The "current location of the user" is specific information of geographical latitude and longitude obtained using the terminal held by the user.

[1124] A "tourist destination database" is a collection of information that holds tourist information about specific areas and facilities, and is used for searches.

[1125] "Related tourist destination information" is detailed information about tourist destinations that is searched from a database based on the user's current location.

[1126] The "means for transmitting to the user terminal" is a communication method for transmitting tourist destination information in data format from the server to the user terminal.

[1127] The "means for generating and playing back audio guides" refers to a technology for converting tourist destination information into audio data and playing back that audio data for the user.

[1128] "Speech synthesis engine" is a general term for software and hardware used to convert text data into speech data.

[1129] "Means for notifying by vibration at important points" is a function that causes the device to vibrate to notify the user when the user approaches an important location in a specific tourist spot.

[1130] "Means for displaying tourist destination information in text format for the hearing impaired" is a function that provides information to hearing impaired users by visually displaying audio information in text.

[1131] "Means for automatically translating tourist destination information into multiple languages" refers to technology for converting information related to tourist destinations into different languages ​​and providing that information to users.

[1132] The present invention relates to an audio guide system that provides detailed information in real time when visiting tourist spots and cultural facilities. This system aims to provide convenience to people with visual or hearing impairments, as well as foreign tourists.

[1133] System Configuration

[1134] The system mainly consists of the following components:

[1135] 1. User device: A mobile device such as a smartphone or smart glasses.

[1136] 2. Server: A server containing a tourist destination database.

[1137] 3. GPS module: A location information acquisition device for obtaining the user's current location.

[1138] 4. Communication means: A network connection for sending and receiving data between the user terminal and the server.

[1139] 5. Audio guide generation function: Software for converting tourist information into audio guides (e.g., Google Text-to-Speech API).

[1140] 6. Vibration function: Notification of important tourist spots.

[1141] 7. Text display function: A means of displaying text information for the hearing impaired.

[1142] 8. Automatic translation function: Software for translating tourist information into different languages ​​(e.g. Microsoft Translator API).

[1143] Program Overview

[1144] The program for this voice guidance system performs the following processes.

[1145] Get current location:

[1146] The user device periodically acquires the user's current location using the GPS module. The acquired location information (latitude and longitude) is stored in the device and sent to the server.

[1147] Search for tourist information:

[1148] The server analyzes the received location information and searches a database for information on the nearest tourist attractions and cultural facilities. It queries the tourist attractions database using SQL queries and stores the results in a data structure (e.g., a list or dictionary).

[1149] Sending and receiving tourist destination information:

[1150] The server sends the searched tourist destination information to the user's terminal, which analyzes the received information and converts it into an appropriate format.

[1151] Generate and provide audio descriptions:

[1152] The user device converts the received text information into audio guidance using a speech synthesis engine, which is then played through the headphones in the smart glasses worn by the user.

[1153] Vibration notification:

[1154] When the user approaches an important point in a tourist destination, the user's device will notify them using a vibration function.

[1155] Text information provided:

[1156] For the hearing impaired, the server sends tourist information in text format to the device, which then displays the text on the smart glasses display.

[1157] Machine translation feature:

[1158] For foreign visitors to Japan, the server automatically translates tourist information and generates multilingual audio guides. The translated audio guides are sent to the user's device and played in the specified language.

[1159] Specific examples

[1160] Example 1: Visiting tourist spots in Kyoto (Japanese user)

[1161] 1. The user arrives at Kyoto Station and starts the system.

[1162] 2. The device uses GPS to obtain its current location (Kyoto Station) and sends it to the server.

[1163] 3. The server searches the database for information about tourist spots around Kyoto Station and sends it to the terminal.

[1164] 4. The device converts the received tourist information into an audio guide, stating, "This is Kyoto Station, located in the center of Kyoto."

[1165] 5. When the user approaches an important point on the way to Kyoto Tower, the device will vibrate to notify them.

[1166] Example 2: Sightseeing for foreign visitors (English users)

[1167] 1. A foreign user arrives at Todaiji Temple in Nara and starts up the system.

[1168] 2. The device obtains its current location and sends it to the server.

[1169] 3. The server retrieves information about Todaiji Temple from the database and generates a multilingual audio guide.

[1170] 4. The device plays the received English audio guide, stating, "This is the Great Buddha Hall of Todaiji Temple, built in the Nara period."

[1171] Prompt Sentence Examples

[1172] "Please explain the main tourist spots around Kyoto Station."

[1173] "Please tell me in English about the history and highlights of Todaiji Temple."

[1174] This system provides users with appropriate tourist information in real time, enriching their tourist experience. It is also suitable for people with visual and hearing impairments, as well as foreign tourists, and helps them gain a deeper understanding of the history and culture of the places they visit.

[1175] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1176] Step 1: Get current location

[1177] The user device uses a GPS module to obtain the user's current location. Specifically, the device's GPS module receives satellite signals and calculates latitude and longitude information. The obtained location information (e.g., latitude 35.0116, longitude 135.7681) is stored in internal memory and sent to the server. The input is satellite signal data from the GPS module, and the output is latitude and longitude coordinate information.

[1178] Step 2: Search for tourist information

[1179] The server analyzes the latitude and longitude location information received from the user device. Based on the received data, it uses an SQL query to search the tourist spot database and retrieves the relevant tourist spot information. For example, the query executed is "SELECT FROM tourist_spots WHERE latitude BETWEEN 34.9116 AND 35.1116 AND longitude BETWEEN 135.6681 AND 135.8681;". The input is latitude and longitude location information, and the output is a list of related tourist spot information.

[1180] Step 3: Submit tourist destination information

[1181] The server formats the tourist spot information obtained from the search into JSON format and sends it to the user's device over the network. For example, the JSON data sent is "{"spots": [{"name": "Kyoto Station", "description": "Central location in Kyoto"}]}". The input is a list of tourist spot information, and the output is the JSON data sent to the user's device.

[1182] Step 4: Receiving and analyzing tourist destination information

[1183] The user device parses the tourist attraction information in JSON format received from the server. Specifically, it uses a JSON parser to extract the data and saves it in its internal memory. For example, the information parsed is "{"name": "Kyoto Station", "description": "Central location in Kyoto"}". The input is the JSON data received from the server, and the output is the parsed tourist attraction information.

[1184] Step 5: Generate audio guide

[1185] The user device sends the received text information to a speech synthesis engine (e.g., Google Text-to-Speech API) and converts it into audio data. Specifically, it converts the text "This is Kyoto Station, located in the center of Kyoto" into an audio file. The input is the tourist destination information text, and the output is an audio file.

[1186] Step 6: Play the audio guide

[1187] The user device passes the generated audio file to the playback function, which plays the audio through the headphones of the user's smart glasses. This allows the user to hear the audio guidance, "This is Kyoto Station, located in the center of Kyoto." The input is the audio file, and the output is the audio guidance playback to the user.

[1188] Step 7: Vibration notifications at key points

[1189] The user device monitors the latitude and longitude of the current location and the tourist spot, and activates vibration when the distance is within a certain range. For example, the device vibrates when the user approaches Kyoto Tower. The input is the current location and the location information of the tourist spot, and the output is the activation of vibration.

[1190] Step 8: Displaying text information for the hearing impaired

[1191] The server sends tourist destination information in text format to the user's device. The user's device then calls an API to display the received text information on the smart glasses' display. For example, the text information displayed is "This is Kyoto Station, located in the center of Kyoto." The input is tourist destination information in text format, and the output is the text information displayed on the smart glasses' display.

[1192] Step 9: Automatic translation function

[1193] The server uses the Microsoft Translator API to automatically translate tourist information for foreign visitors to Japan. For example, it translates the Japanese phrase "This is Kyoto Station, located in the center of Kyoto" into English. The translated text is sent to a speech synthesis engine, which generates an audio guide. The input is the tourist information text, and the output is the translated audio guide.

[1194] In each processing step, detailed data processing and operations are coordinated to realize a system that provides appropriate tourist information to users.

[1195] (Application example 1)

[1196] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1197] In modern commercial facilities, the provision of information to diverse customers, such as users with visual or hearing impairments and foreign visitors to Japan who require multilingual support, is insufficient. In particular, it is difficult to provide real-time information in stores, such as product information and sales floor guidance, which reduces convenience for users. Given this background, it is necessary to provide consistent information to users with different disabilities and languages.

[1198] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1199] In this invention, the server includes means for acquiring the user's current location, means for searching for related tourist attraction information from a tourist attraction database based on the acquired location information, means for transmitting the searched tourist attraction information to the user terminal, means for generating and playing an audio guide based on the transmitted tourist attraction information, means for acquiring in-store location information, means for searching for related product information based on the acquired in-store location information, means for transmitting the searched product information to the user terminal, and means for generating and playing an audio guide based on the transmitted product information. This makes it possible to provide consistent real-time information to different customers, such as users with visual or hearing impairments and foreign visitors to Japan.

[1200] The "means for acquiring the user's current location" refers to a technical device for measuring and acquiring the latitude and longitude of the user's current location or a specific point within the store.

[1201] "Location information" is data indicating the latitude and longitude of the user's current location or a specific point within a store.

[1202] A "tourist destination database" is a collection of data that systematically stores information about tourist destinations.

[1203] The "means for searching tourist attraction information" is a technical device for searching for relevant information from a tourist attraction database based on the acquired location information.

[1204] A "user terminal" is an electronic device carried by a user, such as a smartphone or smart glasses.

[1205] "Means for generating and playing audio guides" refers to technical devices that convert text information into audio and play it back through the user terminal.

[1206] "Means for obtaining location information within the store" refers to a technical device that uses beacons, Wi-Fi, or other location measurement technologies to identify and obtain the user's specific location within the store.

[1207] The "means for searching for related product information" is a technical device for searching the database for the most suitable product information based on the acquired in-store location information.

[1208] The "vibration function" is a technical device that causes the user's terminal to vibrate to notify the user when the user approaches a specific location.

[1209] "Means for displaying in text format" refers to a technical device for displaying information as a string of characters on the display of a user terminal for the hearing impaired.

[1210] "Universal service" is a service that is consistently provided to all users, regardless of whether they have a disability or speak a different language.

[1211] The present invention is a voice guide system that provides users with detailed information in real time at tourist spots and commercial facilities, with the aim of providing convenience to people with visual or hearing impairments, as well as foreign tourists and shoppers.

[1212] System Configuration

[1213] The system mainly consists of the following components:

[1214] 1. User device: A mobile device such as a smartphone or smart glasses.

[1215] 2. Server: A server containing a tourist destination and product database.

[1216] 3. GPS module: A location information acquisition device for obtaining the user's current location.

[1217] 4. Communication means: A network connection for sending and receiving data between the user terminal and the server.

[1218] 5. Audio guide generation function: Software for converting tourist destination information and product information into audio guides.

[1219] 6. Vibration function: A means of notifying you of important tourist spots and product areas.

[1220] 7. Text display function: A means of displaying text information for the hearing impaired.

[1221] 8. In-store location information acquisition means: A device that acquires the user's location within the store using beacons or Wi-Fi access points.

[1222] Basic operation

[1223] The operation of the system is as follows.

[1224] The user terminal periodically acquires its current location using a GPS module or location information acquisition means within the store. This location information is stored in the user terminal and sent to the server.

[1225] The server analyzes the received location information and searches a database for information on the nearest tourist spots and products, and sends the searched information to the user's device in real time.

[1226] The user device analyzes the received information, converts it into an appropriate format, and generates audio guidance using a speech synthesis engine (e.g., Google Text-to-Speech API or Amazon Polly) that is then played through the user's headphones.

[1227] When the user approaches an important tourist spot or product area, the user's device will notify them using the vibration function.

[1228] For the hearing impaired, the server sends tourist destination and product information in text format to the user's device, and the text information is displayed on the display of smart glasses or a smartphone.

[1229] For foreign visitors to Japan, the server automatically translates tourist destination and product information, generates multilingual audio guides, and sends them to the user's device.

[1230] Specific examples

[1231] Example of a tourist destination application

[1232] When a user visits a tourist spot, they can activate this system, which will provide audio guidance with detailed tourist spot information and historical background based on their current location. For example, when a user arrives at Kyoto Station and activates the system, the device obtains their current location (Kyoto Station) and sends it to the server. The server searches a tourist spot database for information about tourist spots around Kyoto Station and sends it to the device. The device then converts this into an audio guide, stating, "This is Kyoto Station, located in the center of Kyoto."

[1233] Example of a physical store guidance application

[1234] When a user activates this system while shopping in a commercial facility, it acquires location information within the store and provides voice guidance on related product information. For example, if the user is in the electronics section of a department store, the device acquires the user's current location and sends it to the server. The server then searches for related product information from a product database and sends it to the device. The device then converts this into voice guidance, stating, "These are the latest earphone models." Additionally, when the user approaches a specific product area, the device vibrates to notify the user.

[1235] Prompt Sentence Examples

[1236] "I want to provide detailed product information based on the visitor's current location, tailored to their nationality and language preferences. The location information is coordinates (X: 100, Y: 200), and the product information is "sensitive earphones" in Japanese. Please translate this into English and generate an audio guide."

[1237] This will enable consistent, real-time information to be provided to different customers, including users with visual or hearing impairments and foreign visitors to Japan.

[1238] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1239] Step 1:

[1240] Get current location

[1241] The user device periodically acquires the user's current location using a GPS module or in-store location information acquisition means (beacons or Wi-Fi). Specifically, when the user starts the system, the GPS module acquires latitude and longitude information, or receives signals from in-store beacons to identify the X and Y coordinates. This generates current location information (input) and stores it in the user device (output).

[1242] Step 2:

[1243] Sending current location information

[1244] The user device sends the acquired current location information to the server. Specifically, it uses a network communication method (Wi-Fi or mobile data communication) to send the location information data to the server. As a result, the location information data from the user device (input) is sent to the server (output).

[1245] Step 3:

[1246] Search for tourist information or product information

[1247] The server analyzes the received location information and searches for related tourist attraction information or product information from a tourist attraction database or product database. Specifically, it uses a database search algorithm to extract the nearest tourist attraction or product information from the location information (input). This process generates related tourist attraction information and product information (output).

[1248] Step 4:

[1249] Sending related information

[1250] The server sends the searched tourist spot information or product information to the user terminal. Specifically, the tourist spot information or product information (input) is sent to the user terminal via the network. This operation allows the tourist spot information or product information from the server to reach the user terminal (output).

[1251] Step 5:

[1252] Audio guide generation and provision

[1253] The user device analyzes the received information and converts it into an appropriate format. Specifically, it uses a speech synthesis engine (Google Text-to-Speech API or Amazon Polly) to convert text information (input) into voice data. The converted voice guide (output) is played through the user device's headphones.

[1254] Step 6:

[1255] Vibration notification

[1256] When a user approaches a particular tourist attraction or product area, the user's device will vibrate to notify them. This is done by calculating the distance to each set point based on the device's location data (input), and vibrating when the distance falls below a certain level (output).

[1257] Step 7:

[1258] Providing text information

[1259] For the hearing impaired, the server sends tourist destination information and product information in text format to the user's device. The user's device then displays this information on its display. Specifically, the text data (input) is screen-rendered for display. This provides the user with visual text information (output).

[1260] Step 8:

[1261] Machine translation and multilingual support

[1262] The server automatically translates tourist destination information and product information to generate multilingual audio guides for foreign visitors to Japan. Specifically, it translates text information (input) into the specified language using a multilingual translation model (e.g., Google Translate API). This translated text information is then converted into audio guides using a speech synthesis engine and sent to the user's device. This provides a translated multilingual audio guide (output).

[1263] Specific example prompts

[1264] "I want to provide detailed product information based on the visitor's current location, tailored to their nationality and language preferences. The location information is coordinates (X: 100, Y: 200), and the product information is "sensitive earphones" in Japanese. Please translate this into English and generate an audio guide."

[1265] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1266] The present invention relates to an audio guide system that provides detailed information in real time when visiting tourist spots and cultural facilities. This system aims to provide convenience especially for people with visual or hearing impairments and foreign tourists. Furthermore, the present invention realizes a personalized sightseeing experience by recognizing the user's emotional state and providing information accordingly.

[1267] System Configuration

[1268] The system mainly consists of the following components:

[1269] 1. User device: A mobile device such as a smartphone or smart glasses.

[1270] 2. Server: A server containing a tourist destination database.

[1271] 3. GPS module: A location information acquisition device for obtaining the user's current location.

[1272] 4. Communication means: A network connection for sending and receiving data between the user terminal and the server.

[1273] 5. Audio guide generation function: Software for converting tourist information into audio guides.

[1274] 6. Vibration function: Notification of important tourist spots.

[1275] 7. Text display function: A means of displaying text information for the hearing impaired.

[1276] 8. Emotion engine: Software that recognizes emotional states and dynamically changes the content and format of the tourist destination information provided.

[1277] Program processing

[1278] 1. Obtaining the current location

[1279] The user device periodically acquires the user's current location (latitude and longitude) using the GPS module. This location information is temporarily stored in the device and then sent to the server.

[1280] 2. Search for tourist information

[1281] The server analyzes the received location information and uses it to search a database for information on the nearest tourist attractions and cultural facilities.

[1282] 3. Sending and receiving tourist destination information

[1283] The server organizes the tourist destination information obtained as search results and returns it to the user's device. The returned data includes the tourist destination's description text, audio URL, and location information of important points.

[1284] 4. Generating and Providing Audio Guides

[1285] The user device converts the received text information into audio guidance using a speech synthesis engine, which is then played through the headphones in the smart glasses worn by the user.

[1286] 5. Vibration notifications

[1287] When the user approaches an important point in a tourist destination, the user's device will notify them using a vibration function.

[1288] 6. Providing text information

[1289] For the hearing impaired, the server sends tourist information in text format to the user's device, which then displays the received text information on the smart glasses display, providing a visual guide.

[1290] 7. Automatic translation function

[1291] For foreign visitors to Japan, the server detects the user's language preference and automatically translates tourist information. The translated tourist information is then sent back to the device, which generates a multilingual audio guide. The device then plays the audio guide in the specified language.

[1292] 8. Emotion recognition and personalized information delivery

[1293] The emotion engine installed in the user device uses cameras and biometric sensors to recognize the user's emotional state from their facial expressions and biometric data. For example, if the user is excited, the emotion engine will recognize this and prioritize providing information on particularly attractive tourist spots and activity recommendations.

[1294] If the user is under stress, the system will provide information on relaxing spots and quiet places. By providing appropriate information according to the user's emotional state, the system can improve the user's satisfaction.

[1295] Specific examples

[1296] Example 1: Visiting tourist spots in Kyoto (Japanese user)

[1297] The user arrives at Kyoto Station and starts the system. The device uses GPS to obtain its current location (Kyoto Station) and sends it to the server. The server then searches a database for information about tourist spots around Kyoto Station and sends it to the device. The device converts the received tourist spot information into an audio guide, providing instructions such as, "This is Kyoto Station, located in the center of Kyoto." As the user approaches important points on their way to Kyoto Tower, the device vibrates to notify them. If the emotion engine detects the user's excitement, it provides additional information such as, "You can see a wonderful night view from the observation deck of Kyoto Tower."

[1298] Example 2: Sightseeing for foreign visitors (English users)

[1299] A foreign user arrives at Todaiji Temple in Nara and starts the system. The device obtains its current location and sends it to the server. The server retrieves information about Todaiji Temple from a database and generates a multilingual audio guide. The device plays the received English audio guide, stating, "This is the Great Buddha Hall of Todaiji Temple, built in the Nara period." If the emotion engine detects the user's stress level, it also provides information such as, "There is a quiet, relaxing park nearby."

[1300] This system can improve the travel experience by providing users with appropriate tourism information in real time via voice or text. Furthermore, the emotion engine enables information to be provided according to the user's emotional state, realizing a more personalized and satisfying tourism experience.

[1301] The processing flow will be explained below.

[1302] Step 1:

[1303] When the user arrives at a tourist destination, they launch the application on their device (smartphone or smart glasses).

[1304] Step 2:

[1305] The device uses GPS to obtain the user's current location (latitude and longitude), which is temporarily stored in the device.

[1306] Step 3:

[1307] The device sends the acquired location information to a server via the network. The transmitted data includes the user ID and device ID.

[1308] Step 4:

[1309] The server analyzes the received location information and uses it to search a database for information on the nearest tourist attractions and cultural facilities.

[1310] Step 5:

[1311] The server organizes the tourist destination information obtained as search results and returns it to the user's device. The returned data includes the tourist destination's description text, audio URL, and location information of important points.

[1312] Step 6:

[1313] The user terminal analyzes the received tourist destination information and extracts text data to be converted into machine voice.

[1314] Step 7:

[1315] The device's built-in voice synthesis engine converts text data into voice data and provides audio guidance through the smart glasses' headphones.

[1316] Step 8:

[1317] As the user continues to explore the tourist spot, the device compares the current location with the location information of important points, and notifies the user with a vibration function when the user approaches an important point.

[1318] Step 9:

[1319] For the hearing impaired, the server sends tourist information in text format to the user's terminal.

[1320] Step 10:

[1321] The user terminal displays the received text information on the smart glasses display, providing a visual guide.

[1322] Step 11:

[1323] For foreign users, the server detects the user's language setting and automatically translates tourist information.

[1324] Step 12:

[1325] The translated tourist information is then sent back to the device, where a multilingual audio guide is generated and played back in the specified language.

[1326] Step 13:

[1327] The emotion engine installed in the user terminal uses a camera and biometric sensors to recognize the user's emotional state from their facial expressions and biometric data.

[1328] Step 14:

[1329] The device transmits the recognized emotional state to the server, which then selects tourist destination information according to the emotion and retransmits it to the user device.

[1330] Step 15:

[1331] If the emotion engine detects the user's excited state, the server will prioritize providing information on tourist spots and activity guides that are more likely to interest the user.

[1332] Step 16:

[1333] If the emotion engine detects the user's stress state, the server will provide information on relaxation spots and quiet places.

[1334] Step 17:

[1335] The user terminal provides the newly received tourist spot information to the user by voice or text.

[1336] Example 2

[1337] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1338] Conventional tourist guide systems do not provide sufficient convenience for people with visual or hearing impairments or foreign tourists. Furthermore, they do not provide customized information based on the user's emotional state, making it difficult to realize a personalized tourist experience. Furthermore, they lack multilingual support, making it difficult to provide information appropriately to users with different cultural backgrounds.

[1339] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring the user's current location, means for searching for related tourist attraction information from a tourist attraction database based on the acquired location information, means for transmitting the searched tourist attraction information to the user terminal, means for generating and playing back an audio guide based on the transmitted tourist attraction information, means for recognizing the user's emotional state and dynamically changing the content and format of the tourist attraction information based on the recognition result, means for automatically translating the tourist attraction information and providing a multilingual audio guide, and means for displaying the tourist attraction information in text format for the hearing impaired. This provides convenience for people with visual or hearing impairments and foreign travelers, enables the provision of customized information according to the user's emotional state, and makes it possible to provide appropriate information to users with different cultural backgrounds by supporting multiple languages.

[1340] The "means for acquiring the user's current location" is a function for acquiring the user's current location using a GPS module or other location information acquisition device.

[1341] The "means for searching for related tourist attraction information from a tourist attraction database based on the acquired location information" is a function in which the server analyzes the location information received and uses that information to search for information on the nearest tourist attractions and cultural facilities from the tourist attraction database.

[1342] "Means for sending searched tourist spot information to user terminal" refers to the function by which the server returns tourist spot information obtained as search results to the user terminal. The returned data includes explanatory text of the tourist spot, audio URL, and location information of important points.

[1343] "Means for generating and playing back audio guides based on transmitted tourist attraction information" refers to a function that converts tourist attraction information received by a user terminal into audio guides using a voice synthesis engine and plays them back.

[1344] "Means for recognizing the user's emotional state and dynamically changing the content and format of tourist destination information based on the recognition results" refers to a function in which the emotion engine installed in the device uses the camera and biometric sensors to recognize the user's emotional state and customizes the information provided based on the results.

[1345] "Means for automatically translating tourist destination information and providing multilingual audio guides" refers to a function in which the server detects the user's set language, translates the tourist destination information into the corresponding language, and generates multilingual audio guides based on that translation information.

[1346] The "means for displaying tourist spot information in text format for the hearing impaired" is a function in which the server transmits tourist spot information in text format to the user terminal, and the user terminal displays it visually.

[1347] This invention relates to an audio guide system that provides detailed information in real time when visiting tourist spots and cultural facilities. This system aims to provide convenience especially for people with visual or hearing impairments, as well as foreign tourists. Furthermore, it realizes a personalized sightseeing experience by recognizing the user's emotional state and providing information according to it.

[1348] System Configuration

[1349] The system mainly consists of the following components:

[1350] 1. User device: Use a mobile device such as a smartphone or smart glasses.

[1351] 2. Server: A server that contains a tourist destination database, and the database uses MySQL or similar.

[1352] 3. GPS module: Includes a location information acquisition device for acquiring the user's current location.

[1353] 4. Communication method: A network connection for sending and receiving data between the user device and the server. Wi-Fi, 4G / 5G, etc. are used.

[1354] 5. Audio guide generation function: Google Cloud Text-to-Speech is used as software to convert tourist information into audio guides.

[1355] 6. Vibration function: The device's vibration function is used to notify users of important tourist spots.

[1356] 7. Text display function: Equipped with a display to display text information for the hearing impaired.

[1357] 8. Emotion engine: Cameras and biometric sensors are used as software to recognize emotional states and dynamically change the content and format of tourist information provided.

[1358] Program processing explanation

[1359] Get current location

[1360] The user device periodically acquires the user's current location (latitude and longitude) using the GPS module. This location information is temporarily stored in the device and then sent to the server.

[1361] Search for tourist information

[1362] The server analyzes the received location information and uses it to search a database for information on the nearest tourist attractions and cultural facilities.

[1363] Sending and receiving tourist information

[1364] The server organizes the tourist destination information obtained as search results and returns it to the user's device. The returned data includes the tourist destination's description text, audio URL, and location information of important points.

[1365] Generating and providing audio guides

[1366] The user device converts the received text information into audio guidance using a speech synthesis engine, which is then played through the headphones in the smart glasses worn by the user.

[1367] Vibration notification

[1368] When the user approaches an important point in a tourist destination, the user's device will notify them using a vibration function.

[1369] Providing text information

[1370] For the hearing impaired, the server sends tourist information in text format to the user's device, which then displays the received text information on the smart glasses display, providing a visual guide.

[1371] Machine translation function

[1372] For foreign visitors to Japan, the server detects the user's language preference and automatically translates tourist information. The translated tourist information is then sent back to the device, which generates a multilingual audio guide. The device then plays the audio guide in the specified language.

[1373] Emotion recognition and customized information provision

[1374] The emotion engine installed in the user device uses cameras and biometric sensors to recognize the user's emotional state from their facial expressions and biometric data. For example, if the user is excited, the emotion engine will recognize this and provide them with information on particularly attractive tourist spots and activity recommendations first. If the user is stressed, it will provide them with information on relaxing spots and quiet places. Providing appropriate information according to the user's emotional state can improve user satisfaction.

[1375] Specific examples

[1376] Example 1: Visiting tourist spots in Kyoto (Japanese user)

[1377] The user arrives at Kyoto Station and starts the system. The device uses GPS to obtain its current location (Kyoto Station) and sends it to the server. The server then searches a database for information about tourist spots around Kyoto Station and sends it to the device. The device converts the received tourist spot information into an audio guide, providing instructions such as, "This is Kyoto Station, located in the center of Kyoto." As the user approaches important points on their way to Kyoto Tower, the device vibrates to notify them. If the emotion engine detects the user's excitement, it provides additional information such as, "You can see a wonderful night view from the observation deck of Kyoto Tower."

[1378] Example 2: Sightseeing for foreign visitors (English users)

[1379] A foreign user arrives at Todaiji Temple in Nara and starts the system. The device obtains its current location and sends it to the server. The server retrieves information about Todaiji Temple from a database and generates a multilingual audio guide. The device plays the received English audio guide, stating, "This is the Great Buddha Hall of Todaiji Temple, built in the Nara period." If the emotion engine detects the user's stress level, it also provides information such as, "There is a quiet, relaxing park nearby."

[1380] Prompt Sentence Examples

[1381] As an example of an input prompt for a generative AI model, you could enter something like:

[1382] "Please explain in detail the process of the system that, when the user arrives at a tourist spot, acquires the user's current location using the smartphone, acquires tourist spot information from the server, and customizes the information according to the user's emotional state."

[1383]

[1384] "Describe the detailed process by which a real-time audio guide system uses a smartphone to acquire the current location, fetch tourist information from a server, and customize the information based on the user's emotional state."

[1385] This system can improve the travel experience by providing users with appropriate tourism information in real time via voice or text. Furthermore, the emotion engine enables information to be provided according to the user's emotional state, realizing a more personalized and satisfying tourism experience.

[1386] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1387] Step 1: Get current location

[1388] Input: The user launches the app and the GPS module obtains the current location.

[1389] Specific operation: A user arrives at a tourist destination and launches the app on their smartphone or smart glasses. The device immediately accesses the built-in GPS module to obtain the user's current location (latitude and longitude). The message "Retrieving user's current location..." is displayed.

[1390] Data processing or data calculation: Position information (e.g., 35.0116°N, 135.7681°E) is acquired and temporarily stored in the device's memory.

[1391] Output: The obtained location information is sent to the server.

[1392] Step 2: Search for tourist information

[1393] Input: The server receives the user's location information.

[1394] Specific operation: The server performs a database search based on the location information (latitude and longitude) received. The database stores tourist spot names, descriptions, audio guide URLs, etc.

[1395] Data processing or data calculation: The server analyzes the location information and searches for information on nearby tourist attractions in a database (e.g., MySQL).

[1396] Output: Tourist information (e.g., Kyoto Tower description text, audio URL, location information of important points) is obtained and sent to the user's device.

[1397] Step 3: Sending and receiving tourist destination information

[1398] Input: The server organizes tourist destination information from search results.

[1399] Specific operation: The server formats the tourist attraction information obtained as a search result, packages the necessary data (explanatory text, audio URL, location information of important points, etc.) and sends it to the user's terminal.

[1400] Data processing or data calculation: Convert tourist destination information data into a data structure such as JSON format, making it a format that can be transferred efficiently.

[1401] Output: The formatted tourist destination information is sent to the user's terminal.

[1402] Step 4: Generate and serve audio descriptions

[1403] Input: The user terminal receives tourist destination information from the server.

[1404] Specific operation: The user device passes the received text information to a speech synthesis engine (e.g., Google Cloud Text-to-Speech), which converts the information into audio data. The generated audio guide is played through the smart glasses headphones.

[1405] Data processing or data calculation: Converting text data into audio data and saving it in a playable format.

[1406] Output: The audio description is played to the user.

[1407] Step 5: Vibration Notification

[1408] Input: The user approaches a point of interest at a tourist destination.

[1409] Specific operation: When the user approaches a key point, the device receives location information from the built-in GPS and determines that it is a key point. The device then starts vibrating and notifies the user that "You are approaching a key point."

[1410] Data processing or data calculation: Compare the current location with the location of important points, and activate the vibration function when they come within a certain distance.

[1411] Output: A vibration will be activated to notify the user.

[1412] Step 6: Provide text information

[1413] Input: The server sends tourist information for the hearing impaired in text format.

[1414] Specific operation: The server sends text information to the user's terminal, which then displays it on the smart glasses display.

[1415] Data processing or data calculation: Converting text information into a displayable format and presenting it on a display.

[1416] Output: Text information is displayed on the smart glasses display.

[1417] Step 7: Automatic translation function

[1418] Input: The server detects the user's preferred language.

[1419] Specific operation: The server reads the language setting of the user's device, automatically translates the tourist attraction information based on that setting, and then sends the translated information back to the user's device.

[1420] Data processing or data calculation: Translate tourist destination information into the specified language using an automatic translation tool (e.g., Google Translate API).

[1421] Output: The translated tourist information is sent to the user's terminal, and an audio guide is played in the specified language.

[1422] Step 8: Emotion recognition and customized information delivery

[1423] Input: The user's emotional state is acquired by a camera or biometric sensors installed on the user's device.

[1424] How it works: The emotion engine analyzes the user's emotional state using facial recognition technology, heart rate measurement, etc. For example, if the user is excited, it will prioritize providing "attractive spots."

[1425] Data processing or data calculation: An algorithm is used to analyze the acquired biometric and facial expression data and determine the emotional state.

[1426] Output: Customized information according to the emotional state is provided to the user.

[1427] (Application example 2)

[1428] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1429] The present invention aims to improve convenience for people with visual or hearing impairments and foreign tourists by providing detailed information in real time when visiting tourist spots and cultural facilities through an audio guide system. It also aims to realize a personalized tourist experience by recognizing the user's emotional state and providing information accordingly. Furthermore, this technology can be applied to a factory work support system to provide workers with real-time work guides and important information, and to improve work efficiency and safety by detecting the user's fatigue and stress level and providing appropriate rest instructions.

[1430] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1431] In this invention, the server includes means for acquiring the user's current location, means for searching for related tourist attraction information from a tourist attraction database based on the acquired location information, means for transmitting the searched tourist attraction information to the user terminal, means for generating and playing back an audio guide based on the transmitted tourist attraction information, and means for recognizing the user's emotional state and dynamically changing the content and format of the tourist attraction information to be provided in accordance with that state. This makes it possible to provide the user with appropriate information in real time when visiting tourist attractions or working in a factory, improving the user experience and work efficiency.

[1432] "User's current location" refers to the geographic location of a particular user as determined by a GPS module or other location acquisition function.

[1433] A "tourist destination database" refers to an information system that stores information about tourist destinations and cultural facilities in a structured format.

[1434] "Tourist destination information" refers to detailed information such as descriptions, images, and audio guides about specific tourist destinations and cultural facilities.

[1435] "User terminal" refers to an electronic device that a user can carry around, such as a smartphone or smart glasses.

[1436] "Audio guide" refers to systems and content that provide audio information about tourist destinations and cultural facilities.

[1437] "Emotional state" refers to the user's emotional and mental state obtained by analyzing the user's facial expressions and biosensor data.

[1438] The "vibration function" refers to a function that notifies the user by vibrating the user terminal.

[1439] "Text format" refers to a format in which information is displayed in text.

[1440] "Real-time" refers to providing current information and status immediately without delay.

[1441] "Work guide" refers to information on instructions and operating procedures provided when working in a factory.

[1442] "Rest instructions" refer to instructions regarding when and where the user should take a rest.

[1443] "Personalized tourism experience" refers to customized tourism information and guidance provided based on the user's individual circumstances and interests.

[1444] "Multilingual" refers to the ability to provide information and guidance in multiple languages.

[1445] "Factory work support system" refers to technology and platforms that support work within a factory.

[1446] This invention is a voice guide system that provides users with detailed information in real time when visiting tourist spots and cultural facilities. This system aims to provide convenience, especially for people with visual or hearing impairments and foreign tourists. Furthermore, it realizes a personalized tourist experience by recognizing the user's emotional state and providing information accordingly. This technology can also be applied to a factory work support system, providing real-time work guidance and important information to workers, and improving work efficiency and safety by detecting the user's fatigue and stress level and providing appropriate rest instructions.

[1447] System Configuration

[1448] The system consists of the following components:

[1449] 1. User Device:

[1450] These are portable electronic devices such as smartphones and smart glasses that enable users to acquire and display information at tourist spots and in factories.

[1451] 2. Server:

[1452] This server contains a database of tourist attractions and data on factory work. The server provides relevant information based on requests from user terminals.

[1453] 3. GPS module:

[1454] This is a location information acquisition device for acquiring the current location of the user.

[1455] 4. Means of communication:

[1456] A network connection for sending and receiving data between a user terminal and a server.

[1457] 5. Audio guide generation function:

[1458] This is software for converting tourist information into audio guides. A typical example of this software is the TextToSpeech module.

[1459] 6. Vibration function:

[1460] It is a means of notifying important tourist spots and important locations within the factory.

[1461] 7. Text display function:

[1462] This is a way to display tourist information and work guides as text information for the hearing impaired. For this, we use the DisplayTextModule.

[1463] 8. Emotion Engine:

[1464] This software recognizes the user's emotional state and dynamically changes the content and format of the information provided. For example, the EmotionRecognition module can be used to analyze the user's emotional state.

[1465] Operating principle

[1466] The server includes a means for acquiring the user's current location, a means for searching for related information from a tourist spot database or a factory work database based on the acquired location information, a means for transmitting the searched information to the user terminal, a means for generating and playing back an audio guide based on the transmitted information, and a means for recognizing the user's emotional state and dynamically changing the content and format of the information to be provided depending on that state. This makes it possible to provide the user with appropriate information in real time while visiting tourist spots or working in a factory, improving the user experience and work efficiency.

[1467] Specific examples

[1468] Example 1: Factory work support system

[1469] Users wear smart glasses while moving around the factory to perform their work. The system uses a GPS module to obtain their current location, and the server searches for relevant work information based on the obtained location information. The received information is provided to the user as voice guidance and text display. The system also uses an EmotionRecognition module to recognize the user's emotional state, and if fatigue or stress is detected, it provides a vibration function and instructions to take a break.

[1470] Prompt Sentence Examples

[1471] "Please suggest an application program that updates the smart glasses feed in real time and displays work instructions based on the user's current location. Also, include a feature that automatically notifies the user to take a break if they are feeling stressed."

[1472] As described above, the present invention is a system for significantly improving work efficiency at tourist spots and factories. By implementing this system, it is expected that the user experience will be improved and work efficiency will increase.

[1473] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1474] Step 1:

[1475] The user terminal acquires its current location (latitude and longitude) using a GPS module.

[1476] input:

[1477] Location information obtained from the GPS module of the user's smart glasses or smartphone.

[1478] Data processing and calculation:

[1479] Obtain latitude and longitude data of your current location and temporarily store it on the device.

[1480] output:

[1481] Numerical data of latitude and longitude obtained.

[1482] Step 2:

[1483] The terminal transmits information about its current location to the server.

[1484] input:

[1485] Numerical latitude and longitude data obtained in step 1.

[1486] Specific behavior:

[1487] The device transmits GPS data to a server via a network.

[1488] output:

[1489] A request to send location information to the server.

[1490] Step 3:

[1491] The server analyzes the received location information and searches for related information from a tourist destination database or a factory work database.

[1492] input:

[1493] Latitude and longitude data of the current location sent from the device.

[1494] Data processing and calculation:

[1495] Query a database based on location information to find relevant tourist information or work procedures.

[1496] output:

[1497] Associated tourist attraction information and work procedure data.

[1498] Step 4:

[1499] The server organizes the obtained information and returns it to the user terminal.

[1500] input:

[1501] Tourist destination information and work procedure data searched in Step 3.

[1502] Specific behavior:

[1503] The server organizes the information, converts it into a specific format, and then transmits it over the network to the user terminal.

[1504] output:

[1505] Tourist destination information and work procedure data are sent to the device.

[1506] Step 5:

[1507] The user terminal converts the received information into audio guidance using a speech synthesis engine (TextToSpeech module) and plays it back.

[1508] input:

[1509] Received text data of tourist destination information and work procedures.

[1510] Data processing and calculation:

[1511] The process of converting text information into speech.

[1512] output:

[1513] Audio data played as audio guide.

[1514] Specific behavior:

[1515] The device uses the TextToSpeech module to convert the text into speech, which is played through the smart glasses' headphones.

[1516] Step 6:

[1517] When the user approaches an important point in a tourist spot or an important location in a factory, the device will notify them by vibrating.

[1518] input:

[1519] User location and location information of important points for tourist attractions and work procedures.

[1520] Data processing and calculation:

[1521] The current location is compared with the location of an important point and it is detected when it enters a certain range.

[1522] output:

[1523] Vibration notifications.

[1524] Specific behavior:

[1525] The user terminal activates the vibration function to notify the user.

[1526] Step 7:

[1527] The user terminal uses the EmotionRecognition module to recognize the user's emotional state and provide information according to the emotional state.

[1528] input:

[1529] Emotional state data from cameras and biometric sensors (e.g., facial expressions, heart rate).

[1530] Data processing and calculation:

[1531] The emotional state data is analyzed to recognize the emotional state of the user.

[1532] output:

[1533] Providing customized information according to the user's emotional state.

[1534] Specific behavior:

[1535] The EmotionRecognition module assesses the user's emotional state and provides information, including instructions to rest, if the user feels tired, for example.

[1536] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1537] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1538] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1539] [Fourth embodiment]

[1540] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1541] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1542] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1543] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1544] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1545] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1546] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1547] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1548] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1549] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1550] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1551] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1552] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1553] The present invention relates to an audio guide system that provides detailed information in real time when visiting tourist spots and cultural facilities. This system aims to provide convenience to people with visual or hearing impairments, as well as foreign tourists.

[1554] System Configuration

[1555] The system mainly consists of the following components:

[1556] 1. User device: A mobile device such as a smartphone or smart glasses.

[1557] 2. Server: A server containing a tourist destination database.

[1558] 3. GPS module: A location information acquisition device for obtaining the user's current location.

[1559] 4. Communication means: A network connection for sending and receiving data between the user terminal and the server.

[1560] 5. Audio guide generation function: Software for converting tourist information into audio guides.

[1561] 6. Vibration function: Notification of important tourist spots.

[1562] 7. Text display function: A means of displaying text information for the hearing impaired.

[1563] Program processing

[1564] 1. Obtaining the current location

[1565] The user terminal periodically uses the GPS module to obtain the user's current location.

[1566] The acquired location information (latitude and longitude) is stored in the device and sent to the server.

[1567] 2. Search for tourist information

[1568] The server analyzes the received location information and searches a database for information on the nearest tourist attractions and cultural facilities.

[1569] 3. Sending and receiving tourist destination information

[1570] The server transmits the searched tourist spot information to the user terminal.

[1571] The user terminal parses the received information and converts it into an appropriate format.

[1572] 4. Generating and Providing Audio Guides

[1573] The user terminal converts the received text information into audio guidance using a speech synthesis engine.

[1574] The audio guide is played through headphones in the smart glasses worn by the user.

[1575] 5. Vibration notifications

[1576] When the user approaches an important point in a tourist destination, the user's device will notify them using a vibration function.

[1577] 6. Providing text information

[1578] For the hearing impaired, the server sends tourist information in text format to the terminal.

[1579] The user terminal displays text information on the smart glasses display.

[1580] 7. Automatic translation function

[1581] For foreign visitors to Japan, the server automatically translates tourist information and generates audio guides in multiple languages.

[1582] The translated audio guide is sent to the user's terminal and played back in the specified language.

[1583] Specific examples

[1584] Example 1: Visiting tourist spots in Kyoto (Japanese user)

[1585] The user arrives at Kyoto Station and starts the system.

[1586] The device uses GPS to obtain its current location (Kyoto Station) and sends it to the server.

[1587] The server searches a database for information about tourist attractions around Kyoto Station and sends it to the terminal.

[1588] The device converts the received tourist information into an audio guide, stating, "This is Kyoto Station, located in the center of Kyoto."

[1589] When the user approaches an important point on their way to Kyoto Tower, the device will vibrate to notify them.

[1590] Example 2: Sightseeing for foreign visitors (English users)

[1591] A foreign user arrives at Todaiji Temple in Nara and starts up the system.

[1592] The device obtains its current location and sends it to the server.

[1593] The server retrieves information about Todaiji Temple from a database and generates a multilingual audio guide.

[1594] The device plays the received English audio guide, stating, "This is the Great Buddha Hall of Todaiji Temple, built in the Nara period."

[1595] This system provides users with appropriate tourist information in real time, enriching their tourist experience. It is also suitable for people with visual and hearing impairments, as well as foreign tourists, and helps them gain a deeper understanding of the history and culture of the places they visit.

[1596] The processing flow will be explained below.

[1597] Step 1:

[1598] When the user arrives at a tourist destination, they launch the application on their device (smartphone or smart glasses).

[1599] Step 2:

[1600] The device uses GPS to obtain the user's current location (latitude and longitude), which is temporarily stored in the device.

[1601] Step 3:

[1602] The device sends the acquired location information to a server via the network. The transmitted data includes the user ID and device ID.

[1603] Step 4:

[1604] The server analyzes the received location information and uses it to search a database for information on the nearest tourist attractions and cultural facilities.

[1605] Step 5:

[1606] The server organizes the tourist destination information obtained as search results and returns it to the user's device. The returned data includes the tourist destination's description text, audio URL, and location information of important points.

[1607] Step 6:

[1608] The user terminal analyzes the received tourist destination information and extracts text data to be converted into machine voice.

[1609] Step 7:

[1610] The device's built-in voice synthesis engine converts text data into voice data and provides audio guidance through the smart glasses' headphones.

[1611] Step 8:

[1612] As the user continues to explore the tourist spot, the device compares the current location with the location information of important points, and notifies the user with a vibration function when the user approaches an important point.

[1613] Step 9:

[1614] For the hearing impaired, the server sends tourist information in text format to the user's terminal.

[1615] Step 10:

[1616] The user terminal displays the received text information on the smart glasses display, providing a visual guide.

[1617] Step 11:

[1618] For foreign users, the server detects the user's language setting and automatically translates tourist information.

[1619] Step 12:

[1620] The translated tourist information is then sent back to the device, where a multilingual audio guide is generated and played back in the specified language.

[1621] Through the above process, users can receive detailed tourist information in real time by voice or text, improving their travel experience.

[1622] Example 1

[1623] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1624] Conventional tourist guide systems are limited in the information they provide based on the user's current location, and therefore lack the ability to provide detailed, real-time tourist information in multiple languages ​​or support people with visual or hearing impairments. Furthermore, they lack the ability to notify users when they are approaching important tourist spots, limiting the user's sightseeing experience.

[1625] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1626] In this invention, the server includes means for acquiring the user's current location, means for searching for related tourist attraction information from a tourist attraction database based on the acquired location information, means for transmitting the searched tourist attraction information to the user terminal, means for generating and playing an audio guide based on the transmitted tourist attraction information, means for using a speech synthesis engine to play the audio guide, means for notifying the user by vibration of important points near the user's destination, means for displaying tourist attraction information in text format for the hearing impaired, and means for automatically translating the tourist attraction information into multiple languages. This allows the user to obtain detailed tourist information in real time, enabling a comprehensive tourist experience for people with visual or hearing impairments. Furthermore, the user can receive vibration notifications when approaching important points at the destination, preventing them from missing important tourist attractions.

[1627] Below are definitions of each important word.

[1628] The "current location of the user" is specific information of geographical latitude and longitude obtained using the terminal held by the user.

[1629] A "tourist destination database" is a collection of information that holds tourist information about specific areas and facilities, and is used for searches.

[1630] "Related tourist destination information" is detailed information about tourist destinations that is searched from a database based on the user's current location.

[1631] The "means for transmitting to the user terminal" is a communication method for transmitting tourist destination information in data format from the server to the user terminal.

[1632] The "means for generating and playing back audio guides" refers to a technology for converting tourist destination information into audio data and playing back that audio data for the user.

[1633] "Speech synthesis engine" is a general term for software and hardware used to convert text data into speech data.

[1634] "Means for notifying by vibration at important points" is a function that causes the device to vibrate to notify the user when the user approaches an important location in a specific tourist spot.

[1635] "Means for displaying tourist destination information in text format for the hearing impaired" is a function that provides information to hearing impaired users by visually displaying audio information in text.

[1636] "Means for automatically translating tourist destination information into multiple languages" refers to technology for converting information related to tourist destinations into different languages ​​and providing that information to users.

[1637] The present invention relates to an audio guide system that provides detailed information in real time when visiting tourist spots and cultural facilities. This system aims to provide convenience to people with visual or hearing impairments, as well as foreign tourists.

[1638] System Configuration

[1639] The system mainly consists of the following components:

[1640] 1. User device: A mobile device such as a smartphone or smart glasses.

[1641] 2. Server: A server containing a tourist destination database.

[1642] 3. GPS module: A location information acquisition device for obtaining the user's current location.

[1643] 4. Communication means: A network connection for sending and receiving data between the user terminal and the server.

[1644] 5. Audio guide generation function: Software for converting tourist information into audio guides (e.g., Google Text-to-Speech API).

[1645] 6. Vibration function: Notification of important tourist spots.

[1646] 7. Text display function: A means of displaying text information for the hearing impaired.

[1647] 8. Automatic translation function: Software for translating tourist information into different languages ​​(e.g. Microsoft Translator API).

[1648] Program Overview

[1649] The program for this voice guidance system performs the following processes.

[1650] Get current location:

[1651] The user device periodically acquires the user's current location using the GPS module. The acquired location information (latitude and longitude) is stored in the device and sent to the server.

[1652] Search for tourist information:

[1653] The server analyzes the received location information and searches a database for information on the nearest tourist attractions and cultural facilities. It queries the tourist attractions database using SQL queries and stores the results in a data structure (e.g., a list or dictionary).

[1654] Sending and receiving tourist destination information:

[1655] The server sends the searched tourist destination information to the user's terminal, which analyzes the received information and converts it into an appropriate format.

[1656] Generate and provide audio descriptions:

[1657] The user device converts the received text information into audio guidance using a speech synthesis engine, which is then played through the headphones in the smart glasses worn by the user.

[1658] Vibration notification:

[1659] When the user approaches an important point in a tourist destination, the user's device will notify them using a vibration function.

[1660] Text information provided:

[1661] For the hearing impaired, the server sends tourist information in text format to the device, which then displays the text on the smart glasses display.

[1662] Machine translation feature:

[1663] For foreign visitors to Japan, the server automatically translates tourist information and generates multilingual audio guides. The translated audio guides are sent to the user's device and played in the specified language.

[1664] Specific examples

[1665] Example 1: Visiting tourist spots in Kyoto (Japanese user)

[1666] 1. The user arrives at Kyoto Station and starts the system.

[1667] 2. The device uses GPS to obtain its current location (Kyoto Station) and sends it to the server.

[1668] 3. The server searches the database for information about tourist spots around Kyoto Station and sends it to the terminal.

[1669] 4. The device converts the received tourist information into an audio guide, stating, "This is Kyoto Station, located in the center of Kyoto."

[1670] 5. When the user approaches an important point on the way to Kyoto Tower, the device will vibrate to notify them.

[1671] Example 2: Sightseeing for foreign visitors (English users)

[1672] 1. A foreign user arrives at Todaiji Temple in Nara and starts up the system.

[1673] 2. The device obtains its current location and sends it to the server.

[1674] 3. The server retrieves information about Todaiji Temple from the database and generates a multilingual audio guide.

[1675] 4. The device plays the received English audio guide, stating, "This is the Great Buddha Hall of Todaiji Temple, built in the Nara period."

[1676] Prompt Sentence Examples

[1677] "Please explain the main tourist spots around Kyoto Station."

[1678] "Please tell me in English about the history and highlights of Todaiji Temple."

[1679] This system provides users with appropriate tourist information in real time, enriching their tourist experience. It is also suitable for people with visual and hearing impairments, as well as foreign tourists, and helps them gain a deeper understanding of the history and culture of the places they visit.

[1680] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1681] Step 1: Get current location

[1682] The user device uses a GPS module to obtain the user's current location. Specifically, the device's GPS module receives satellite signals and calculates latitude and longitude information. The obtained location information (e.g., latitude 35.0116, longitude 135.7681) is stored in internal memory and sent to the server. The input is satellite signal data from the GPS module, and the output is latitude and longitude coordinate information.

[1683] Step 2: Search for tourist information

[1684] The server analyzes the latitude and longitude location information received from the user device. Based on the received data, it uses an SQL query to search the tourist spot database and retrieves the relevant tourist spot information. For example, the query executed is "SELECT FROM tourist_spots WHERE latitude BETWEEN 34.9116 AND 35.1116 AND longitude BETWEEN 135.6681 AND 135.8681;". The input is latitude and longitude location information, and the output is a list of related tourist spot information.

[1685] Step 3: Submit tourist destination information

[1686] The server formats the tourist spot information obtained from the search into JSON format and sends it to the user's device over the network. For example, the JSON data sent is "{"spots": [{"name": "Kyoto Station", "description": "Central location in Kyoto"}]}". The input is a list of tourist spot information, and the output is the JSON data sent to the user's device.

[1687] Step 4: Receiving and analyzing tourist destination information

[1688] The user device parses the tourist attraction information in JSON format received from the server. Specifically, it uses a JSON parser to extract the data and saves it in its internal memory. For example, the information parsed is "{"name": "Kyoto Station", "description": "Central location in Kyoto"}". The input is the JSON data received from the server, and the output is the parsed tourist attraction information.

[1689] Step 5: Generate audio guide

[1690] The user device sends the received text information to a speech synthesis engine (e.g., Google Text-to-Speech API) and converts it into audio data. Specifically, it converts the text "This is Kyoto Station, located in the center of Kyoto" into an audio file. The input is the tourist destination information text, and the output is an audio file.

[1691] Step 6: Play the audio guide

[1692] The user device passes the generated audio file to the playback function, which plays the audio through the headphones of the user's smart glasses. This allows the user to hear the audio guidance, "This is Kyoto Station, located in the center of Kyoto." The input is the audio file, and the output is the audio guidance playback to the user.

[1693] Step 7: Vibration notifications at key points

[1694] The user device monitors the latitude and longitude of the current location and the tourist spot, and activates vibration when the distance is within a certain range. For example, the device vibrates when the user approaches Kyoto Tower. The input is the current location and the location information of the tourist spot, and the output is the activation of vibration.

[1695] Step 8: Displaying text information for the hearing impaired

[1696] The server sends tourist destination information in text format to the user's device. The user's device then calls an API to display the received text information on the smart glasses' display. For example, the text information displayed is "This is Kyoto Station, located in the center of Kyoto." The input is tourist destination information in text format, and the output is the text information displayed on the smart glasses' display.

[1697] Step 9: Automatic translation function

[1698] The server uses the Microsoft Translator API to automatically translate tourist information for foreign visitors to Japan. For example, it translates the Japanese phrase "This is Kyoto Station, located in the center of Kyoto" into English. The translated text is sent to a speech synthesis engine, which generates an audio guide. The input is the tourist information text, and the output is the translated audio guide.

[1699] In each processing step, detailed data processing and operations are coordinated to realize a system that provides appropriate tourist information to users.

[1700] (Application example 1)

[1701] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1702] In modern commercial facilities, the provision of information to diverse customers, such as users with visual or hearing impairments and foreign visitors to Japan who require multilingual support, is insufficient. In particular, it is difficult to provide real-time information in stores, such as product information and sales floor guidance, which reduces convenience for users. Given this background, it is necessary to provide consistent information to users with different disabilities and languages.

[1703] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1704] In this invention, the server includes means for acquiring the user's current location, means for searching for related tourist attraction information from a tourist attraction database based on the acquired location information, means for transmitting the searched tourist attraction information to the user terminal, means for generating and playing an audio guide based on the transmitted tourist attraction information, means for acquiring in-store location information, means for searching for related product information based on the acquired in-store location information, means for transmitting the searched product information to the user terminal, and means for generating and playing an audio guide based on the transmitted product information. This makes it possible to provide consistent real-time information to different customers, such as users with visual or hearing impairments and foreign visitors to Japan.

[1705] The "means for acquiring the user's current location" refers to a technical device for measuring and acquiring the latitude and longitude of the user's current location or a specific point within the store.

[1706] "Location information" is data indicating the latitude and longitude of the user's current location or a specific point within a store.

[1707] A "tourist destination database" is a collection of data that systematically stores information about tourist destinations.

[1708] The "means for searching tourist attraction information" is a technical device for searching for relevant information from a tourist attraction database based on the acquired location information.

[1709] A "user terminal" is an electronic device carried by a user, such as a smartphone or smart glasses.

[1710] "Means for generating and playing audio guides" refers to technical devices that convert text information into audio and play it back through the user terminal.

[1711] "Means for obtaining location information within the store" refers to a technical device that uses beacons, Wi-Fi, or other location measurement technologies to identify and obtain the user's specific location within the store.

[1712] The "means for searching for related product information" is a technical device for searching the database for the most suitable product information based on the acquired in-store location information.

[1713] The "vibration function" is a technical device that causes the user's terminal to vibrate to notify the user when the user approaches a specific location.

[1714] "Means for displaying in text format" refers to a technical device for displaying information as a string of characters on the display of a user terminal for the hearing impaired.

[1715] "Universal service" is a service that is consistently provided to all users, regardless of whether they have a disability or speak a different language.

[1716] The present invention is a voice guide system that provides users with detailed information in real time at tourist spots and commercial facilities, with the aim of providing convenience to people with visual or hearing impairments, as well as foreign tourists and shoppers.

[1717] System Configuration

[1718] The system mainly consists of the following components:

[1719] 1. User device: A mobile device such as a smartphone or smart glasses.

[1720] 2. Server: A server containing a tourist destination and product database.

[1721] 3. GPS module: A location information acquisition device for obtaining the user's current location.

[1722] 4. Communication means: A network connection for sending and receiving data between the user terminal and the server.

[1723] 5. Audio guide generation function: Software for converting tourist destination information and product information into audio guides.

[1724] 6. Vibration function: A means of notifying you of important tourist spots and product areas.

[1725] 7. Text display function: A means of displaying text information for the hearing impaired.

[1726] 8. In-store location information acquisition means: A device that acquires the user's location within the store using beacons or Wi-Fi access points.

[1727] Basic operation

[1728] The operation of the system is as follows.

[1729] The user terminal periodically acquires its current location using a GPS module or location information acquisition means within the store. This location information is stored in the user terminal and sent to the server.

[1730] The server analyzes the received location information and searches a database for information on the nearest tourist spots and products, and sends the searched information to the user's device in real time.

[1731] The user device analyzes the received information, converts it into an appropriate format, and generates audio guidance using a speech synthesis engine (e.g., Google Text-to-Speech API or Amazon Polly) that is then played through the user's headphones.

[1732] When the user approaches an important tourist spot or product area, the user's device will notify them using the vibration function.

[1733] For the hearing impaired, the server sends tourist destination and product information in text format to the user's device, and the text information is displayed on the display of smart glasses or a smartphone.

[1734] For foreign visitors to Japan, the server automatically translates tourist destination and product information, generates multilingual audio guides, and sends them to the user's device.

[1735] Specific examples

[1736] Example of a tourist destination application

[1737] When a user visits a tourist spot, they can activate this system, which will provide audio guidance with detailed tourist spot information and historical background based on their current location. For example, when a user arrives at Kyoto Station and activates the system, the device obtains their current location (Kyoto Station) and sends it to the server. The server searches a tourist spot database for information about tourist spots around Kyoto Station and sends it to the device. The device then converts this into an audio guide, stating, "This is Kyoto Station, located in the center of Kyoto."

[1738] Example of a physical store guidance application

[1739] When a user activates this system while shopping in a commercial facility, it acquires location information within the store and provides voice guidance on related product information. For example, if the user is in the electronics section of a department store, the device acquires the user's current location and sends it to the server. The server then searches for related product information from a product database and sends it to the device. The device then converts this into voice guidance, stating, "These are the latest earphone models." Additionally, when the user approaches a specific product area, the device vibrates to notify the user.

[1740] Prompt Sentence Examples

[1741] "I want to provide detailed product information based on the visitor's current location, tailored to their nationality and language preferences. The location information is coordinates (X: 100, Y: 200), and the product information is "sensitive earphones" in Japanese. Please translate this into English and generate an audio guide."

[1742] This will enable consistent, real-time information to be provided to different customers, including users with visual or hearing impairments and foreign visitors to Japan.

[1743] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1744] Step 1:

[1745] Get current location

[1746] The user device periodically acquires the user's current location using a GPS module or in-store location information acquisition means (beacons or Wi-Fi). Specifically, when the user starts the system, the GPS module acquires latitude and longitude information, or receives signals from in-store beacons to identify the X and Y coordinates. This generates current location information (input) and stores it in the user device (output).

[1747] Step 2:

[1748] Sending current location information

[1749] The user device sends the acquired current location information to the server. Specifically, it uses a network communication method (Wi-Fi or mobile data communication) to send the location information data to the server. As a result, the location information data from the user device (input) is sent to the server (output).

[1750] Step 3:

[1751] Search for tourist information or product information

[1752] The server analyzes the received location information and searches for related tourist attraction information or product information from a tourist attraction database or product database. Specifically, it uses a database search algorithm to extract the nearest tourist attraction or product information from the location information (input). This process generates related tourist attraction information and product information (output).

[1753] Step 4:

[1754] Sending related information

[1755] The server sends the searched tourist spot information or product information to the user terminal. Specifically, the tourist spot information or product information (input) is sent to the user terminal via the network. This operation allows the tourist spot information or product information from the server to reach the user terminal (output).

[1756] Step 5:

[1757] Audio guide generation and provision

[1758] The user device analyzes the received information and converts it into an appropriate format. Specifically, it uses a speech synthesis engine (Google Text-to-Speech API or Amazon Polly) to convert text information (input) into voice data. The converted voice guide (output) is played through the user device's headphones.

[1759] Step 6:

[1760] Vibration notification

[1761] When a user approaches a particular tourist attraction or product area, the user's device will vibrate to notify them. This is done by calculating the distance to each set point based on the device's location data (input), and vibrating when the distance falls below a certain level (output).

[1762] Step 7:

[1763] Providing text information

[1764] For the hearing impaired, the server sends tourist destination information and product information in text format to the user's device. The user's device then displays this information on its display. Specifically, the text data (input) is screen-rendered for display. This provides the user with visual text information (output).

[1765] Step 8:

[1766] Machine translation and multilingual support

[1767] The server automatically translates tourist destination information and product information to generate multilingual audio guides for foreign visitors to Japan. Specifically, it translates text information (input) into the specified language using a multilingual translation model (e.g., Google Translate API). This translated text information is then converted into audio guides using a speech synthesis engine and sent to the user's device. This provides a translated multilingual audio guide (output).

[1768] Specific example prompts

[1769] "I want to provide detailed product information based on the visitor's current location, tailored to their nationality and language preferences. The location information is coordinates (X: 100, Y: 200), and the product information is "sensitive earphones" in Japanese. Please translate this into English and generate an audio guide."

[1770] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1771] The present invention relates to an audio guide system that provides detailed information in real time when visiting tourist spots and cultural facilities. This system aims to provide convenience especially for people with visual or hearing impairments and foreign tourists. Furthermore, the present invention realizes a personalized sightseeing experience by recognizing the user's emotional state and providing information accordingly.

[1772] System Configuration

[1773] The system mainly consists of the following components:

[1774] 1. User device: A mobile device such as a smartphone or smart glasses.

[1775] 2. Server: A server containing a tourist destination database.

[1776] 3. GPS module: A location information acquisition device for obtaining the user's current location.

[1777] 4. Communication means: A network connection for sending and receiving data between the user terminal and the server.

[1778] 5. Audio guide generation function: Software for converting tourist information into audio guides.

[1779] 6. Vibration function: Notification of important tourist spots.

[1780] 7. Text display function: A means of displaying text information for the hearing impaired.

[1781] 8. Emotion engine: Software that recognizes emotional states and dynamically changes the content and format of the tourist destination information provided.

[1782] Program processing

[1783] 1. Obtaining the current location

[1784] The user device periodically acquires the user's current location (latitude and longitude) using the GPS module. This location information is temporarily stored in the device and then sent to the server.

[1785] 2. Search for tourist information

[1786] The server analyzes the received location information and uses it to search a database for information on the nearest tourist attractions and cultural facilities.

[1787] 3. Sending and receiving tourist destination information

[1788] The server organizes the tourist destination information obtained as search results and returns it to the user's device. The returned data includes the tourist destination's description text, audio URL, and location information of important points.

[1789] 4. Generating and Providing Audio Guides

[1790] The user device converts the received text information into audio guidance using a speech synthesis engine, which is then played through the headphones in the smart glasses worn by the user.

[1791] 5. Vibration notifications

[1792] When the user approaches an important point in a tourist destination, the user's device will notify them using a vibration function.

[1793] 6. Providing text information

[1794] For the hearing impaired, the server sends tourist information in text format to the user's device, which then displays the received text information on the smart glasses display, providing a visual guide.

[1795] 7. Automatic translation function

[1796] For foreign visitors to Japan, the server detects the user's language preference and automatically translates tourist information. The translated tourist information is then sent back to the device, which generates a multilingual audio guide. The device then plays the audio guide in the specified language.

[1797] 8. Emotion recognition and personalized information delivery

[1798] The emotion engine installed in the user device uses cameras and biometric sensors to recognize the user's emotional state from their facial expressions and biometric data. For example, if the user is excited, the emotion engine will recognize this and prioritize providing information on particularly attractive tourist spots and activity recommendations.

[1799] If the user is under stress, the system will provide information on relaxing spots and quiet places. By providing appropriate information according to the user's emotional state, the system can improve the user's satisfaction.

[1800] Specific examples

[1801] Example 1: Visiting tourist spots in Kyoto (Japanese user)

[1802] The user arrives at Kyoto Station and starts the system. The device uses GPS to obtain its current location (Kyoto Station) and sends it to the server. The server then searches a database for information about tourist spots around Kyoto Station and sends it to the device. The device converts the received tourist spot information into an audio guide, providing instructions such as, "This is Kyoto Station, located in the center of Kyoto." As the user approaches important points on their way to Kyoto Tower, the device vibrates to notify them. If the emotion engine detects the user's excitement, it provides additional information such as, "You can see a wonderful night view from the observation deck of Kyoto Tower."

[1803] Example 2: Sightseeing for foreign visitors (English users)

[1804] A foreign user arrives at Todaiji Temple in Nara and starts the system. The device obtains its current location and sends it to the server. The server retrieves information about Todaiji Temple from a database and generates a multilingual audio guide. The device plays the received English audio guide, stating, "This is the Great Buddha Hall of Todaiji Temple, built in the Nara period." If the emotion engine detects the user's stress level, it also provides information such as, "There is a quiet, relaxing park nearby."

[1805] This system can improve the travel experience by providing users with appropriate tourism information in real time via voice or text. Furthermore, the emotion engine enables information to be provided according to the user's emotional state, realizing a more personalized and satisfying tourism experience.

[1806] The processing flow will be explained below.

[1807] Step 1:

[1808] When the user arrives at a tourist destination, they launch the application on their device (smartphone or smart glasses).

[1809] Step 2:

[1810] The device uses GPS to obtain the user's current location (latitude and longitude), which is temporarily stored in the device.

[1811] Step 3:

[1812] The device sends the acquired location information to a server via the network. The transmitted data includes the user ID and device ID.

[1813] Step 4:

[1814] The server analyzes the received location information and uses it to search a database for information on the nearest tourist attractions and cultural facilities.

[1815] Step 5:

[1816] The server organizes the tourist destination information obtained as search results and returns it to the user's device. The returned data includes the tourist destination's description text, audio URL, and location information of important points.

[1817] Step 6:

[1818] The user terminal analyzes the received tourist destination information and extracts text data to be converted into machine voice.

[1819] Step 7:

[1820] The device's built-in voice synthesis engine converts text data into voice data and provides audio guidance through the smart glasses' headphones.

[1821] Step 8:

[1822] As the user continues to explore the tourist spot, the device compares the current location with the location information of important points, and notifies the user with a vibration function when the user approaches an important point.

[1823] Step 9:

[1824] For the hearing impaired, the server sends tourist information in text format to the user's terminal.

[1825] Step 10:

[1826] The user terminal displays the received text information on the smart glasses display, providing a visual guide.

[1827] Step 11:

[1828] For foreign users, the server detects the user's language setting and automatically translates tourist information.

[1829] Step 12:

[1830] The translated tourist information is then sent back to the device, where a multilingual audio guide is generated and played back in the specified language.

[1831] Step 13:

[1832] The emotion engine installed in the user terminal uses a camera and biometric sensors to recognize the user's emotional state from their facial expressions and biometric data.

[1833] Step 14:

[1834] The device transmits the recognized emotional state to the server, which then selects tourist destination information according to the emotion and retransmits it to the user device.

[1835] Step 15:

[1836] If the emotion engine detects the user's excited state, the server will prioritize providing information on tourist spots and activity guides that are more likely to interest the user.

[1837] Step 16:

[1838] If the emotion engine detects the user's stress state, the server will provide information on relaxation spots and quiet places.

[1839] Step 17:

[1840] The user terminal provides the newly received tourist spot information to the user by voice or text.

[1841] Example 2

[1842] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1843] Conventional tourist guide systems do not provide sufficient convenience for people with visual or hearing impairments or foreign tourists. Furthermore, they do not provide customized information based on the user's emotional state, making it difficult to realize a personalized tourist experience. Furthermore, they lack multilingual support, making it difficult to provide information appropriately to users with different cultural backgrounds.

[1844] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring the user's current location, means for searching for related tourist attraction information from a tourist attraction database based on the acquired location information, means for transmitting the searched tourist attraction information to the user terminal, means for generating and playing back an audio guide based on the transmitted tourist attraction information, means for recognizing the user's emotional state and dynamically changing the content and format of the tourist attraction information based on the recognition result, means for automatically translating the tourist attraction information and providing a multilingual audio guide, and means for displaying the tourist attraction information in text format for the hearing impaired. This provides convenience for people with visual or hearing impairments and foreign travelers, enables the provision of customized information according to the user's emotional state, and makes it possible to provide appropriate information to users with different cultural backgrounds by supporting multiple languages.

[1845] The "means for acquiring the user's current location" is a function for acquiring the user's current location using a GPS module or other location information acquisition device.

[1846] The "means for searching for related tourist attraction information from a tourist attraction database based on the acquired location information" is a function in which the server analyzes the location information received and uses that information to search for information on the nearest tourist attractions and cultural facilities from the tourist attraction database.

[1847] "Means for sending searched tourist spot information to user terminal" refers to the function by which the server returns tourist spot information obtained as search results to the user terminal. The returned data includes explanatory text of the tourist spot, audio URL, and location information of important points.

[1848] "Means for generating and playing back audio guides based on transmitted tourist attraction information" refers to a function that converts tourist attraction information received by a user terminal into audio guides using a voice synthesis engine and plays them back.

[1849] "Means for recognizing the user's emotional state and dynamically changing the content and format of tourist destination information based on the recognition results" refers to a function in which the emotion engine installed in the device uses the camera and biometric sensors to recognize the user's emotional state and customizes the information provided based on the results.

[1850] "Means for automatically translating tourist destination information and providing multilingual audio guides" refers to a function in which the server detects the user's set language, translates the tourist destination information into the corresponding language, and generates multilingual audio guides based on that translation information.

[1851] The "means for displaying tourist spot information in text format for the hearing impaired" is a function in which the server transmits tourist spot information in text format to the user terminal, and the user terminal displays it visually.

[1852] This invention relates to an audio guide system that provides detailed information in real time when visiting tourist spots and cultural facilities. This system aims to provide convenience especially for people with visual or hearing impairments, as well as foreign tourists. Furthermore, it realizes a personalized sightseeing experience by recognizing the user's emotional state and providing information according to it.

[1853] System Configuration

[1854] The system mainly consists of the following components:

[1855] 1. User device: Use a mobile device such as a smartphone or smart glasses.

[1856] 2. Server: A server that contains a tourist destination database, and the database uses MySQL or similar.

[1857] 3. GPS module: Includes a location information acquisition device for acquiring the user's current location.

[1858] 4. Communication method: A network connection for sending and receiving data between the user device and the server. Wi-Fi, 4G / 5G, etc. are used.

[1859] 5. Audio guide generation function: Google Cloud Text-to-Speech is used as software to convert tourist information into audio guides.

[1860] 6. Vibration function: The device's vibration function is used to notify users of important tourist spots.

[1861] 7. Text display function: Equipped with a display to display text information for the hearing impaired.

[1862] 8. Emotion engine: Cameras and biometric sensors are used as software to recognize emotional states and dynamically change the content and format of tourist information provided.

[1863] Program processing explanation

[1864] Get current location

[1865] The user device periodically acquires the user's current location (latitude and longitude) using the GPS module. This location information is temporarily stored in the device and then sent to the server.

[1866] Search for tourist information

[1867] The server analyzes the received location information and uses it to search a database for information on the nearest tourist attractions and cultural facilities.

[1868] Sending and receiving tourist information

[1869] The server organizes the tourist destination information obtained as search results and returns it to the user's device. The returned data includes the tourist destination's description text, audio URL, and location information of important points.

[1870] Generating and providing audio guides

[1871] The user device converts the received text information into audio guidance using a speech synthesis engine, which is then played through the headphones in the smart glasses worn by the user.

[1872] Vibration notification

[1873] When the user approaches an important point in a tourist destination, the user's device will notify them using a vibration function.

[1874] Providing text information

[1875] For the hearing impaired, the server sends tourist information in text format to the user's device, which then displays the received text information on the smart glasses display, providing a visual guide.

[1876] Machine translation function

[1877] For foreign visitors to Japan, the server detects the user's language preference and automatically translates tourist information. The translated tourist information is then sent back to the device, which generates a multilingual audio guide. The device then plays the audio guide in the specified language.

[1878] Emotion recognition and customized information provision

[1879] The emotion engine installed in the user device uses cameras and biometric sensors to recognize the user's emotional state from their facial expressions and biometric data. For example, if the user is excited, the emotion engine will recognize this and provide them with information on particularly attractive tourist spots and activity recommendations first. If the user is stressed, it will provide them with information on relaxing spots and quiet places. Providing appropriate information according to the user's emotional state can improve user satisfaction.

[1880] Specific examples

[1881] Example 1: Visiting tourist spots in Kyoto (Japanese user)

[1882] The user arrives at Kyoto Station and starts the system. The device uses GPS to obtain its current location (Kyoto Station) and sends it to the server. The server then searches a database for information about tourist spots around Kyoto Station and sends it to the device. The device converts the received tourist spot information into an audio guide, providing instructions such as, "This is Kyoto Station, located in the center of Kyoto." As the user approaches important points on their way to Kyoto Tower, the device vibrates to notify them. If the emotion engine detects the user's excitement, it provides additional information such as, "You can see a wonderful night view from the observation deck of Kyoto Tower."

[1883] Example 2: Sightseeing for foreign visitors (English users)

[1884] A foreign user arrives at Todaiji Temple in Nara and starts the system. The device obtains its current location and sends it to the server. The server retrieves information about Todaiji Temple from a database and generates a multilingual audio guide. The device plays the received English audio guide, stating, "This is the Great Buddha Hall of Todaiji Temple, built in the Nara period." If the emotion engine detects the user's stress level, it also provides information such as, "There is a quiet, relaxing park nearby."

[1885] Prompt Sentence Examples

[1886] As an example of an input prompt for a generative AI model, you could enter something like:

[1887] "Please explain in detail the process of the system that, when the user arrives at a tourist spot, acquires the user's current location using the smartphone, acquires tourist spot information from the server, and customizes the information according to the user's emotional state."

[1888]

[1889] "Describe the detailed process by which a real-time audio guide system uses a smartphone to acquire the current location, fetch tourist information from a server, and customize the information based on the user's emotional state."

[1890] This system can improve the travel experience by providing users with appropriate tourism information in real time via voice or text. Furthermore, the emotion engine enables information to be provided according to the user's emotional state, realizing a more personalized and satisfying tourism experience.

[1891] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1892] Step 1: Get current location

[1893] Input: The user launches the app and the GPS module obtains the current location.

[1894] Specific operation: A user arrives at a tourist destination and launches the app on their smartphone or smart glasses. The device immediately accesses the built-in GPS module to obtain the user's current location (latitude and longitude). The message "Retrieving user's current location..." is displayed.

[1895] Data processing or data calculation: Position information (e.g., 35.0116°N, 135.7681°E) is acquired and temporarily stored in the device's memory.

[1896] Output: The obtained location information is sent to the server.

[1897] Step 2: Search for tourist information

[1898] Input: The server receives the user's location information.

[1899] Specific operation: The server performs a database search based on the location information (latitude and longitude) received. The database stores tourist spot names, descriptions, audio guide URLs, etc.

[1900] Data processing or data calculation: The server analyzes the location information and searches for information on nearby tourist attractions in a database (e.g., MySQL).

[1901] Output: Tourist information (e.g., Kyoto Tower description text, audio URL, location information of important points) is obtained and sent to the user's device.

[1902] Step 3: Sending and receiving tourist destination information

[1903] Input: The server organizes tourist destination information from search results.

[1904] Specifi...

Claims

1. means for obtaining a user's current location; A means for searching for related tourist attraction information from a tourist attraction database based on the acquired location information; means for transmitting the searched tourist spot information to a user terminal; means for generating and playing back an audio guide based on the transmitted tourist attraction information; A system including:

2. The system according to claim 1, further comprising a vibration function for notifying the user of tourist information in real time based on the user's current location.

3. 2. The system according to claim 1, further comprising means for displaying tourist information in text format for hearing impaired people in order to provide universal service.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A