system

A system that uses real-time location-based audio guidance and feedback optimization enriches travel experiences by providing detailed geographical and historical information, addressing the limitations of traditional tour guides and vast information sources.

JP2026073338APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Modern travelers face challenges in deeply understanding the history and culture of their destinations due to the vast amount of information available, and traditional tour guides are costly, diminishing the educational value and depth of travel experiences.

Method used

A system that acquires user location information in real time, searches a database for relevant geographical information, generates audio content using a text-to-speech engine, and provides it to the user's terminal, allowing for immediate playback and continuous improvement based on user feedback.

Benefits of technology

Enhances travel experiences by providing detailed geographical and historical information in real time, enriching the educational value and personalizing the experience through user feedback optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026073338000001_ABST
    Figure 2026073338000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of obtaining the user's location information, A means for retrieving relevant geographical information from a database based on the aforementioned location information, A means for generating audio content based on the aforementioned searched geographical information, Means for transmitting the aforementioned audio content to a user terminal, means for playing the aforementioned audio content on the user terminal, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Even if modern travelers want to deeply understand the history and culture of their destinations, they are forced to search for and read the necessary data from a vast amount of information. In addition, hiring a traditional tour guide often costs a significant amount. Due to such circumstances, there is a problem that the educational value and depth of experience of travel are diminished.

Means for Solving the Problems

[0005] To address this challenge, the present invention proposes a system that acquires the user's location information in real time and searches a database for relevant geographical information based on that information. Furthermore, it includes a function to generate the retrieved information as audio content, send it to the user's terminal, and play it back. This audio content is generated using a text-to-speech engine and is immediately available to the user. In addition, by collecting user feedback and improving the accuracy of information provided in the future, the invention provides a means to further enrich individual travel experiences.

[0006] "User location information" refers to numerical data of latitude and longitude that indicates where the user is currently located.

[0007] "Geographical information" refers to data about history, culture, tourist attractions, etc., associated with a specific location.

[0008] A "database" is a collection of information systematically organized according to specific rules, which can be searched and updated.

[0009] "Audio content" refers to data that expresses textual information as sound and communicates it to users through their hearing.

[0010] A "text-to-speech engine" is a program or device that possesses the technology to convert text data into speech data.

[0011] A "user terminal" is an electronic device that a user directly operates, and includes smartphones, tablets, and other similar devices.

[0012] "Feedback" is the act of a user communicating their evaluation, opinion, or reaction to the information or service provided to a system. [Brief explanation of the drawing]

[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0014] An example of an embodiment of the system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), and the like.

[0017] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0018] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disk (e.g., hard disk), or magnetic tape, and the like.

[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0021] [First Embodiment]

[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0034] This invention is a system for providing relevant geographical information based on the user's location and guiding the user through it via voice. In this embodiment, the aim is to allow the user to obtain information in real time using their handheld device, thereby making the travel experience deeper and richer.

[0035] The user's device uses GPS functionality to determine its current location and sends this location information to the server. The server receives this information and searches its database for relevant geographical information. This allows the server to identify tourist information and historical background relevant to the user's current location.

[0036] Next, the server uses the acquired data to create a text guide for the user and generates audio content using a text-to-speech engine. This audio content is converted to an appropriate data format so that the user can listen to it through their device, and then transferred to the user's device.

[0037] The user's device receives this audio content and plays it back through speakers or headphones. This playback allows the user to hear detailed tourist information on the spot, deepening their understanding of the location. Furthermore, the server collects the user's feedback and uses it to improve future information provision. This feedback function allows the system to continuously optimize the user experience.

[0038] As a concrete example, if a user is in a tourist destination, the server can provide information about local historical events and important figures, and can also play quiz-style questions in audio format. For example, it could ask, "Guess what happened in this place in the 19th century and who was involved." In this way, users can gain high educational value even while on the go.

[0039] The following describes the processing flow.

[0040] Step 1:

[0041] The user's device uses its built-in GPS function to obtain its current location. Location information is recorded as latitude and longitude.

[0042] Step 2:

[0043] The user's device transmits the acquired location information to the server using a secure protocol. The location information is encrypted to protect the data.

[0044] Step 3:

[0045] The server analyzes location information received from the user's device. Based on this information, it searches the database for relevant geographical information, tourist attractions, and historical background.

[0046] Step 4:

[0047] The server generates text guides from the geographical information obtained through searches. These guides include tourist information and descriptions of specific historical events.

[0048] Step 5:

[0049] The server uses a text-to-speech engine to convert the text guide into audio data. This audio data is then formatted to be easily understood by the user.

[0050] Step 6:

[0051] The server sends the generated audio data to the user's terminal. During data transfer, error detection and correction are performed to maintain the reliability of the communication.

[0052] Step 7:

[0053] The user's device plays the received audio data. Audio guidance is provided through speakers or headphones, allowing the user to hear the instructions in real time.

[0054] Step 8:

[0055] Users provide feedback on the information provided through the application. This feedback is sent to the server to optimize future information provision.

[0056] (Example 1)

[0057] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0058] Modern travelers demand detailed and intuitive information about their destinations, but traditional voice guidance systems struggle to fully customize based on location data and have limitations in providing information that reflects user feedback, thus failing to enrich the user experience.

[0059] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0060] In this invention, the server includes an information device for acquiring location information, an information processing device for retrieving relevant geographic information from an information storage device based on the location information, and an information conversion device for generating voice information based on the retrieved geographic information. This makes it possible to provide voice guidance optimized for individual users and to continuously improve the quality of information provision based on feedback.

[0061] "Location information" refers to digital data that represents the longitude and latitude of a specific point on Earth.

[0062] An "information device" is an electronic device that enables the acquisition, processing, and transmission of data.

[0063] An "information storage device" is a recording medium that includes a database and is capable of storing and retrieving diverse types of data.

[0064] An "information processing device" is a computer system that analyzes received data and extracts and provides relevant information.

[0065] An "information conversion device" is a device that has the function of converting data in a specific format into another format.

[0066] A "communication device" is a combination of hardware and software used to send and receive data between different electronic devices.

[0067] An "information terminal" is a device that allows users to interactively receive and display information.

[0068] A "playback device" is a device that represents digital content such as audio and video as a physical output.

[0069] An "information gathering device" is a system that receives input data from users and transmits it to a server.

[0070] This invention is a system that provides users with detailed geographical and historical information about a local area, which they can then listen to as audio guidance. The user uses a mobile device equipped with GPS to determine their current location. The device then transmits this location information to a server.

[0071] The server searches its internal database for relevant geographical information based on the received location data. The database contains information on tourist attractions and historical events related to geographical locations. The retrieved information is extracted as organized text data.

[0072] The server then converts the text data into speech data using a text-to-speech (TTS) engine. While general speech synthesis techniques are used here, open-source TTS engines are available as specific examples.

[0073] The generated audio data is sent from the server to the user's mobile device. After receiving this audio data, the device plays it back to the user through its speaker or headphones.

[0074] As a concrete example, when a user visits a historical tourist site, the server will provide audio information, including historical events in that area, so that the user can perceive that information while on site. This system aims to enhance the user's travel experience through efficient information provision.

[0075] An example of a prompt message might be: "We want to design an audio guidance system to provide historical background and tourist information related to a geographically specified location for commercial purposes. Please list the main functions required and explain the overall flow of the system."

[0076] This invention also includes a function to improve the accuracy of information provided based on user feedback, thereby continuously improving user satisfaction.

[0077] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0078] Step 1:

[0079] The user enables location services on their mobile device. This allows the device to obtain latitude and longitude via its GPS module. The input data is the local latitude and longitude collected by the device. The output is the acquired location information temporarily stored on the device.

[0080] Step 2:

[0081] The device transmits the collected location information to the server. Here, the latitude and longitude stored on the device are used as input, and a location information packet is generated as output for the server to receive. This process is performed via the HTTPS protocol, ensuring security.

[0082] Step 3:

[0083] The server accesses a database based on the received location information and searches for relevant geographical information. The input data is the latitude and longitude sent to the server. The server converts this into a database query and retrieves information about tourist destinations and historical events as output.

[0084] Step 4:

[0085] The server analyzes acquired geographical information and generates a text guide organized for the user. Raw data retrieved from a database is used as input, and concise, easy-to-understand text information is generated as output. This text information is customized based on the user's location and interests.

[0086] Step 5:

[0087] The server passes the organized text information to the text-to-speech engine, which converts it into audio data. The input data is the generated text guide. Based on this, audio data is generated and output as an MP3 audio file.

[0088] Step 6:

[0089] The server sends the generated audio file to the user's terminal. The input is the generated audio file, and the output is the audio data received by the user's terminal. This data is transferred using standard data streaming technology.

[0090] Step 7:

[0091] The terminal plays the received audio data through a speaker or headphones. The input data is an audio file received from the server, and the output is the voice guidance that the user hears. This playback function allows the user to receive real-time voice guidance on the spot.

[0092] Step 8:

[0093] Users can provide feedback on voice guidance through their devices. The input data is the feedback information entered by the user, and the output is the feedback data being sent to the server and used to improve future information provision.

[0094] (Application Example 1)

[0095] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0096] There is a need for safer and more effective ways for users of autonomous vehicles to obtain information about their surroundings while operating them. Conventional systems mainly rely on visual information provision, which can sometimes lead to situations where users are unable to concentrate on driving. It is necessary to solve these problems and improve safety and convenience for users.

[0097] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0098] In this invention, the server includes means for recognizing the user's geographical location signal, means for retrieving relevant regional information from a data storage area based on the geographical location signal, and means for generating voice output based on the retrieved regional information. This allows users in autonomous mobile vehicles to receive regional information by voice while moving via a head-mounted display device, enabling them to acquire information without relying on vision, thereby improving user safety and the efficiency of information provision.

[0099] "User geographic location signal" refers to a digital signal that indicates the user's current location, obtained using GPS or other location-determining technologies.

[0100] "Relevant local information" refers to information such as tourist destinations, historical facts, and events within a specific geographical area, and is considered useful data for users.

[0101] A "data storage area" is a digital storage system for accumulating information and is a database used to enable location-based searches.

[0102] "Audio output" refers to electronic signals that convert text information into sound, providing information to the user through hearing.

[0103] A "user terminal" is a device that has the function of receiving and playing audio, and includes smart glasses and personal digital assistants.

[0104] An "autonomous mobile vehicle" is a vehicle that can move on its own without human intervention, and is typically operated using AI and various sensors.

[0105] A "head-mounted display device" is a device worn on the head and designed to provide information by combining visual and auditory elements.

[0106] This invention is a system for acquiring a user's geographical location signal and providing relevant regional information based on that information. The system uses a GPS module to recognize the user's location and transmits the corresponding location information to a server. After receiving this, the server retrieves relevant regional information from its data storage area and generates voice output based on that information. The voice output is generated in real time using a string-to-speech engine and provided to the user via their user terminal, specifically a head-mounted display device.

[0107] The server uses the Python library `requests` to retrieve necessary information from external services. This data is then converted into speech format using the `text_to_speech` library. The head-mounted display device, acting as the user terminal, has a speaker to play the received audio data, providing information to the user without relying on their vision.

[0108] As a concrete example, if a user is traveling through a tourist destination in an autonomous vehicle, the server can provide audio information about the area's historical background and landmarks. For instance, information such as, "This region has long been known as a hot spring resort and is a popular destination visited by many tourists," could be provided. In this way, users can safely receive rich information while minimizing their reliance on visual cues. An example of a prompt to input into the generating AI model would be, "Please obtain information about the surrounding tourist attractions based on the current location and provide it as an audio guide."

[0109] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0110] Step 1:

[0111] The user obtains their current geographical location signal using a device. The input is location information, received from a GPS module. The output is latitude and longitude data. The device sends this data to the server.

[0112] Step 2:

[0113] The server retrieves relevant regional information from its data storage area based on the latitude and longitude data received from the terminal. The input is latitude and longitude, and the output is a dataset of regional information. The server filters the information by sending queries to the database.

[0114] Step 3:

[0115] The server creates descriptive text strings based on the acquired regional information dataset. The input is a regional information dataset, and the output is descriptive text. The server combines the data to generate structured sentences.

[0116] Step 4:

[0117] The server generates audio data using a string-to-speech engine to convert explanatory text into speech. The input is explanatory text, and the output is an audio file. The server converts the string into a predetermined format and generates an audio waveform.

[0118] Step 5:

[0119] The server sends the generated audio data to the terminal. The input is an audio file, and the output is the audio data that has been transferred to the user's terminal. The server sends the audio data to the terminal using the appropriate protocol.

[0120] Step 6:

[0121] The device plays the received audio data through the headset speaker. The input is the received audio data, and the output is the audio output to the user's hearing. The device decodes the audio file and plays the audio through the speaker.

[0122] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0123] This invention not only provides relevant geographical information based on the user's location, but also recognizes the user's emotions and dynamically adjusts the content provided accordingly. This system enables a more personalized travel experience for the user.

[0124] First, the user's device uses its GPS function to obtain its current location and sends this location information to the server. The server analyzes this location information and retrieves relevant geographical information from its database. At this time, a text guide is generated based on the retrieved geographical information.

[0125] Next, the user's device uses its built-in sensors, camera, and microphone to analyze the user's emotional state with an emotion engine. This is based on facial expression analysis, voice tone recognition, and physiological data obtained from sensors.

[0126] The server uses data obtained from the emotion engine to select or adjust text content according to the user's current emotions. For example, if the server detects that the user is excited, it might prioritize providing information about activities and adventures. This text content is then converted into audio data using a text-to-speech engine.

[0127] This audio data is sent from the server to the user's device and played back on the user's device. This allows the user to receive information appropriate to their emotional state at that time in real time. Furthermore, through a feedback function, the suitability of the information to the emotions is evaluated, and the system is continuously optimized.

[0128] As a concrete example, when a user visits a tourist destination, they can receive different information depending on their mood and reactions that day. If the server detects that the user is seeking relaxation, it will provide information about quiet parks and cafes in the area, along with related historical background. In this way, users can always obtain valuable information that matches their current emotional state.

[0129] The following describes the processing flow.

[0130] Step 1:

[0131] The user's device obtains its current location information using a GPS sensor. This location information is converted into a digital format as latitude and longitude.

[0132] Step 2:

[0133] The user's device sends the acquired location information to the server. The data is encrypted using a secure protocol before transmission.

[0134] Step 3:

[0135] The server uses the received location information to search its database for relevant geographical information. During this process, tourist attractions and historical information are extracted.

[0136] Step 4:

[0137] The microphone and camera on the user's device are used, and an emotion engine analyzes the user's emotional state. This analysis includes voice tone and facial recognition.

[0138] Step 5:

[0139] The server receives emotion data from the emotion engine and generates or adjusts text content according to the user's emotions. This content will be tailored to the user's interests and emotions.

[0140] Step 6:

[0141] The server uses a text-to-speech engine to convert the edited text content into audio data. This audio data is processed in a format that is comfortable for the user to listen to.

[0142] Step 7:

[0143] The server sends the generated audio data to the user's device. The transmitted data is delivered quickly and securely.

[0144] Step 8:

[0145] The user's device receives audio data and plays it back through the built-in speaker or connected earphones. The user can hear the guidance information directly.

[0146] Step 9:

[0147] Users provide feedback on the information they receive. This feedback includes evaluations of the information's content and its relevance to their emotional state.

[0148] Step 10:

[0149] The server analyzes user feedback and optimizes the entire system to improve the personalization accuracy of future information delivery. This process continuously improves the learning database.

[0150] (Example 2)

[0151] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0152] Conventional geographic information systems often provide certain information based on the user's location, but they do not take into account the user's individual emotional state. As a result, there is a lack of personalization of the user experience, and the content provided may not match the user's current emotions.

[0153] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0154] In this invention, the server includes means for acquiring the user's location information, means for retrieving relevant geographical information from a data storage device, and means for analyzing the user's emotional state using an emotion engine. This enables the provision of appropriate and personalized information according to the user's current emotional state.

[0155] "User location information" refers to coordinate data indicating the user's current location, and is typically location information obtained using the Global Positioning Satellite System (GPS).

[0156] "Geographical information" refers to various types of information related to a specific location, including data about the region such as tourist attractions, history, culture, and facilities.

[0157] A "data storage device" is a storage device or database that stores a wide variety of information, such as geographical information, and allows it to be retrieved as needed.

[0158] A "text guide" refers to information presented to users in the form of explanatory or instructive text, which is generated based on geographical information.

[0159] An "emotion engine" is a program or system for analyzing a user's emotional state, using facial expressions, voice, sensor data, etc., to infer emotions.

[0160] "Audio data" refers to digital audio data that is created by converting text content into speech and is intended to be played back on a user's device.

[0161] A "speech synthesis engine" is a program or system for converting text content into speech data, and is a technology for generating artificial speech.

[0162] "Responses" refer to feedback and behavioral data collected from users, which can be used to improve system performance and personalize the system.

[0163] This invention relates to a system that dynamically provides personalized content by utilizing the user's location information and emotions. This system is mainly composed of a combination of a server, a terminal, and various software engines.

[0164] First, the user obtains location information using their device. The device utilizes its built-in GPS sensor to measure latitude and longitude data. This location information is transmitted to the server in real time.

[0165] The server retrieves relevant geographical information from the data storage device based on the received location information. A database management system is used in this process to efficiently obtain geographical information.

[0166] Next, the device uses its built-in camera, microphone, and sensors to analyze the user's emotions in real time. The emotion engine processes data from various sensors and applies a dedicated analysis algorithm to infer the emotional state. This analysis includes facial recognition technology and voice analysis technology.

[0167] Based on the output from the emotion engine, the server dynamically generates or adjusts content according to the user's emotions. Appropriately selected text information is converted into audio data by the speech synthesis engine. The speech synthesis engine uses a specified speech model to convert the text into smooth speech.

[0168] Audio data is sent from the server to the user's device, which then plays the audio using its built-in media player. This allows the user to receive personalized information in real time, based on their location and emotions.

[0169] As a concrete example, consider a scenario where a user is visiting a tourist spot in a city. If facial analysis determines that the user is relaxed, the server will prioritize providing information about quiet places and cafes in that location. Related historical information will also be provided via audio, allowing the user to efficiently obtain information while walking.

[0170] An example of a prompt to a generative AI model is the instruction, "Generate tourist information suitable for when the user is in an excited state." In response to this prompt, the system will prioritize generating information about adventurous activities.

[0171] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0172] Step 1:

[0173] The user's device obtains its current location using GPS functionality. The device inputs latitude and longitude data from the Global Positioning Satellite System (GSO) and transmits this information to the server. The obtained location information serves as the basis for searching for geographical information about the user's surroundings.

[0174] Step 2:

[0175] The server uses location information received from the terminal as input to query the data storage device and retrieve relevant geographical information. Using a database management system, it searches for tourist attractions, facilities, and other regional information, and prepares the retrieved geographical information for use in the next step.

[0176] Step 3:

[0177] The user's device uses its built-in camera, microphone, and sensors to collect data for input into the emotion engine. This data includes facial image data, voice tone, and physiological data. The emotion engine processes this data to determine the user's current emotional state, using image processing algorithms and voice analysis models.

[0178] Step 4:

[0179] The server receives emotional state data output from the emotion engine as input and dynamically generates personalized content. This process uses a generative AI model to create text and information that matches the user's emotions. The generated text content is adjusted according to the user's interests and emotions.

[0180] Step 5:

[0181] The server inputs the selected or edited text content into a speech synthesis engine and converts it into speech data. The synthesized speech is used to generate smooth speech data from the text. Here, a speech synthesis algorithm is applied to output speech in an easy-to-understand format.

[0182] Step 6:

[0183] The server sends the generated audio data to the user's device. The user's device uses its built-in media player to play the received audio data. This allows the user to hear appropriate information in real time, tailored to their location and emotions.

[0184] (Application Example 2)

[0185] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0186] Conventional geographic information systems provide information based solely on the user's location, making it difficult to deliver content tailored to the user's psychological state, such as emotions, stress levels, and curiosity, in real time. Furthermore, a lack of continuous content optimization based on user feedback and individual needs for tourism information was a significant challenge.

[0187] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0188] In this invention, the server includes means for acquiring the user's location information, means for analyzing the acquired user's emotional state, and means for generating personalized audio content based on the emotional state and location information. This makes it possible to provide personalized geographical information and audio guidance that corresponds to the user's psychological state.

[0189] "Means for acquiring user location information" refers to a device or method that has the function of determining the user's current geographical location and acquiring it as digital information.

[0190] "Means for retrieving geographical information from a data set" refers to a device or method that has the function of retrieving and extracting relevant geographical information from a database or the like based on acquired location data.

[0191] "Means for analyzing a user's emotional state" refers to a device or method that has the function of determining a user's psychological emotional state using the user's facial expressions, voice, and other physiological data.

[0192] "Means for generating personalized audio content" refers to a device or method that has the function of creating audio information tailored to individual needs based on the user's emotional state and location information.

[0193] The "feature for collecting user feedback" is a function that provides a mechanism for users to proactively submit their own opinions and suggestions for improvement regarding the information provided.

[0194] This invention is a system for providing personalized geographical information to users. The system mainly consists of a server and a user terminal.

[0195] First, the user's device acquires location information using a GPS module. This location information is sent to a server, which then searches its database for relevant geographical information based on this data. During this process, the server analyzes facial expression and voice data sent from the device to obtain the user's emotional state. Here, a facial expression recognition library (e.g., OpenCV) and voice tone recognition software are used.

[0196] The server uses acquired emotional state and location information to generate personalized audio content. This audio content is converted from text to audio data using a text-to-speech conversion device (e.g., Amazon Polly). The generated audio content is sent to the user's device and played back on the device. This allows the user to receive information in audio format that is tailored to their emotional state in real time.

[0197] Furthermore, it includes a function to collect feedback from users, and the system continuously optimizes content quality based on user feedback.

[0198] For example, if the system determines that a user has arrived at a tourist destination and is relaxing, it can provide information about nearby quiet parks or cafes and offer audio guidance about the historical background.

[0199] An example of a prompt for a generative AI model is: "Analyze the user's emotional state in real time based on their facial expressions and voice tone, and provide location-related tourist information in an engaging narration."

[0200] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0201] Step 1:

[0202] The device uses a GPS module to obtain the user's current location. The obtained location information is sent to the server. The input is the location information, and the output is the location information sent to the server.

[0203] Step 2:

[0204] The server searches the database for relevant geographical information based on the received location information. The input is the transmitted location information, and the output is the retrieved geographical information. The server executes queries and extracts the geographical information.

[0205] Step 3:

[0206] The device uses its camera and microphone to collect the user's facial expressions and voice, and generates emotion data. The input is the user's facial expressions and voice data, and the output is the generated emotion data. The device performs facial recognition and voice analysis.

[0207] Step 4:

[0208] The server analyzes emotional data to identify the user's emotional state. The input is emotional data, and the output is the identified emotional state. The server utilizes a facial expression recognition library and voice tone analysis software.

[0209] Step 5:

[0210] The server generates personalized text content based on identified emotional states and retrieved geographical information. The input is emotional state and geographical information, and the output is text content. The server selects information that matches the user's preferences.

[0211] Step 6:

[0212] The server converts the generated text content into audio data using a text-to-speech conversion processing unit. The input is text content, and the output is audio data. The server performs the conversion using a text-to-speech engine.

[0213] Step 7:

[0214] The server sends the generated audio data to the terminal. The input is the audio data, and the output is the audio data sent to the terminal. The server transfers the data.

[0215] Step 8:

[0216] The device plays the received audio data. The input is the received audio data, and the output is the audio guide provided to the user. The device activates the audio playback function.

[0217] Step 9:

[0218] Users provide feedback on the information they receive. The input is the information provided, and the output is the user's feedback. Users send their feedback via text or voice.

[0219] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0220] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0221] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0222] [Second Embodiment]

[0223] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0224] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0225] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0226] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0227] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0228] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0229] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0230] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0231] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0232] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0233] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0234] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0235] This invention is a system for providing relevant geographical information based on the user's location and guiding the user through it via voice. In this embodiment, the aim is to allow the user to obtain information in real time using their handheld device, thereby making the travel experience deeper and richer.

[0236] The user's device uses GPS functionality to determine its current location and sends this location information to the server. The server receives this information and searches its database for relevant geographical information. This allows the server to identify tourist information and historical background relevant to the user's current location.

[0237] Next, the server uses the acquired data to create a text guide for the user and generates audio content using a text-to-speech engine. This audio content is converted to an appropriate data format so that the user can listen to it through their device, and then transferred to the user's device.

[0238] The user's device receives this audio content and plays it back through speakers or headphones. This playback allows the user to hear detailed tourist information on the spot, deepening their understanding of the location. Furthermore, the server collects the user's feedback and uses it to improve future information provision. This feedback function allows the system to continuously optimize the user experience.

[0239] As a concrete example, if a user is in a tourist destination, the server can provide information about local historical events and important figures, and can also play quiz-style questions in audio format. For example, it could ask, "Guess what happened in this place in the 19th century and who was involved." In this way, users can gain high educational value even while on the go.

[0240] The following describes the processing flow.

[0241] Step 1:

[0242] The user's device uses its built-in GPS function to obtain its current location. Location information is recorded as latitude and longitude.

[0243] Step 2:

[0244] The user's device transmits the acquired location information to the server using a secure protocol. The location information is encrypted to protect the data.

[0245] Step 3:

[0246] The server analyzes location information received from the user's device. Based on this information, it searches the database for relevant geographical information, tourist attractions, and historical background.

[0247] Step 4:

[0248] The server generates text guides from the geographical information obtained through searches. These guides include tourist information and descriptions of specific historical events.

[0249] Step 5:

[0250] The server uses a text-to-speech engine to convert the text guide into audio data. This audio data is then formatted to be easily understood by the user.

[0251] Step 6:

[0252] The server sends the generated audio data to the user's terminal. During data transfer, error detection and correction are performed to maintain the reliability of the communication.

[0253] Step 7:

[0254] The user's device plays the received audio data. Audio guidance is provided through speakers or headphones, allowing the user to hear the instructions in real time.

[0255] Step 8:

[0256] Users provide feedback on the information provided through the application. This feedback is sent to the server to optimize future information provision.

[0257] (Example 1)

[0258] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0259] Modern travelers demand detailed and intuitive information about their destinations, but traditional voice guidance systems struggle to fully customize based on location data and have limitations in providing information that reflects user feedback, thus failing to enrich the user experience.

[0260] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0261] In this invention, the server includes an information device for acquiring location information, an information processing device for retrieving relevant geographic information from an information storage device based on the location information, and an information conversion device for generating voice information based on the retrieved geographic information. This makes it possible to provide voice guidance optimized for individual users and to continuously improve the quality of information provision based on feedback.

[0262] "Location information" refers to digital data that represents the longitude and latitude of a specific point on Earth.

[0263] An "information device" is an electronic device that enables the acquisition, processing, and transmission of data.

[0264] An "information storage device" is a recording medium that includes a database and is capable of storing and retrieving diverse types of data.

[0265] An "information processing device" is a computer system that analyzes received data and extracts and provides relevant information.

[0266] An "information conversion device" is a device that has the function of converting data in a specific format into another format.

[0267] A "communication device" is a combination of hardware and software used to send and receive data between different electronic devices.

[0268] An "information terminal" is a device that allows users to interactively receive and display information.

[0269] A "playback device" is a device that represents digital content such as audio and video as a physical output.

[0270] An "information gathering device" is a system that receives input data from users and transmits it to a server.

[0271] This invention is a system that provides users with detailed geographical and historical information about a local area, which they can then listen to as audio guidance. The user uses a mobile device equipped with GPS to determine their current location. The device then transmits this location information to a server.

[0272] The server searches its internal database for relevant geographical information based on the received location data. The database contains information on tourist attractions and historical events related to geographical locations. The retrieved information is extracted as organized text data.

[0273] The server then converts the text data into speech data using a text-to-speech (TTS) engine. While general speech synthesis techniques are used here, open-source TTS engines are available as specific examples.

[0274] The generated audio data is sent from the server to the user's mobile device. After receiving this audio data, the device plays it back to the user through its speaker or headphones.

[0275] As a concrete example, when a user visits a historical tourist site, the server will provide audio information, including historical events in that area, so that the user can perceive that information while on site. This system aims to enhance the user's travel experience through efficient information provision.

[0276] An example of a prompt message might be: "We want to design an audio guidance system to provide historical background and tourist information related to a geographically specified location for commercial purposes. Please list the main functions required and explain the overall flow of the system."

[0277] This invention also includes a function to improve the accuracy of information provided based on user feedback, thereby continuously improving user satisfaction.

[0278] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0279] Step 1:

[0280] The user enables location services on their mobile device. This allows the device to obtain latitude and longitude via its GPS module. The input data is the local latitude and longitude collected by the device. The output is the acquired location information temporarily stored on the device.

[0281] Step 2:

[0282] The device transmits the collected location information to the server. Here, the latitude and longitude stored on the device are used as input, and a location information packet is generated as output for the server to receive. This process is performed via the HTTPS protocol, ensuring security.

[0283] Step 3:

[0284] The server accesses the database based on the received location information and searches for relevant geographical information. The input data is the latitude and longitude sent to the server. The server converts this into a database query and obtains information about tourist attractions and historical events as output.

[0285] Step 4:

[0286] The server analyzes the obtained geographical information and generates a text guide organized for the user. Raw data obtained from the database is used as input, and concise and easy-to-understand text information is generated as output. This text information is customized based on the user's location and interests.

[0287] Step 5:

[0288] The server passes the organized text information to a text-to-speech engine and converts it into audio data. The input data is the generated text guide. Based on this, audio data is generated and output as an MP3-formatted audio file.

[0289] Step 6:

[0290] The server sends the generated audio file to the user's terminal. The input is the generated audio file, and the output is the audio data received by the user's terminal. This data is transferred using normal data streaming technology.

[0291] Step 7:

[0292] The terminal plays the received audio data through a speaker or headphones. The input data is the audio file received from the server, and the output is the voice guidance that the user listens to. With this playback function, the user can receive real-time voice guidance on the spot.

[0293] Step 8:

[0294] Users can provide feedback on voice guidance through their devices. The input data is the feedback information entered by the user, and the output is the feedback data being sent to the server and used to improve future information provision.

[0295] (Application Example 1)

[0296] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0297] There is a need for safer and more effective ways for users of autonomous vehicles to obtain information about their surroundings while operating them. Conventional systems mainly rely on visual information provision, which can sometimes lead to situations where users are unable to concentrate on driving. It is necessary to solve these problems and improve safety and convenience for users.

[0298] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0299] In this invention, the server includes means for recognizing the user's geographical location signal, means for retrieving relevant regional information from a data storage area based on the geographical location signal, and means for generating voice output based on the retrieved regional information. This allows users in autonomous mobile vehicles to receive regional information by voice while moving via a head-mounted display device, enabling them to acquire information without relying on vision, thereby improving user safety and the efficiency of information provision.

[0300] "User geographic location signal" refers to a digital signal that indicates the user's current location, obtained using GPS or other location-determining technologies.

[0301] "Relevant local information" refers to information such as tourist destinations, historical facts, and events within a specific geographical area, and is considered useful data for users.

[0302] The "data storage area" is digital storage for accumulating information and is a database used to enable searches based on location information.

[0303] "Voice output" is an electronic signal that converts text information into sound and provides information to the user through hearing.

[0304] The "user terminal" is a device that has the function of receiving and playing back voices, and includes smart glasses, mobile information terminals, etc.

[0305] An "autonomous mobile vehicle" is a vehicle that can move under self-control without human intervention and usually operates using AI and various sensors.

[0306] A "head-mounted display device" is a device that is worn on the head of the body and is for providing information by combining visual and auditory elements.

[0307] This invention is a system for acquiring a user's geographical position signal and providing related regional information based on that information. The system uses a GPS module to recognize the user's position and transmits the corresponding position information to the server. After receiving this, the server searches for related regional information from the data storage area and generates voice output based on that information. The voice output is generated in real time using a text-to-speech engine and is provided to the user via the user's terminal, specifically a head-mounted display device.

[0308] The server uses a Python library and utilizes requests to obtain the necessary information from external services. This data is further converted into audio format using a text_to_speech library. The head-mounted display device as the user terminal has a speaker for playing back the received audio data and provides information to the user without relying on vision.

[0309] As a concrete example, if a user is traveling through a tourist destination in an autonomous vehicle, the server can provide audio information about the area's historical background and landmarks. For instance, information such as, "This region has long been known as a hot spring resort and is a popular destination visited by many tourists," could be provided. In this way, users can safely receive rich information while minimizing their reliance on visual cues. An example of a prompt to input into the generating AI model would be, "Please obtain information about the surrounding tourist attractions based on the current location and provide it as an audio guide."

[0310] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0311] Step 1:

[0312] The user obtains their current geographical location signal using a device. The input is location information, received from a GPS module. The output is latitude and longitude data. The device sends this data to the server.

[0313] Step 2:

[0314] The server retrieves relevant regional information from its data storage area based on the latitude and longitude data received from the terminal. The input is latitude and longitude, and the output is a dataset of regional information. The server filters the information by sending queries to the database.

[0315] Step 3:

[0316] The server creates descriptive text strings based on the acquired regional information dataset. The input is a regional information dataset, and the output is descriptive text. The server combines the data to generate structured sentences.

[0317] Step 4:

[0318] The server generates audio data using a string-to-speech engine to convert explanatory text into speech. The input is explanatory text, and the output is an audio file. The server converts the string into a predetermined format and generates an audio waveform.

[0319] Step 5:

[0320] The server sends the generated audio data to the terminal. The input is an audio file, and the output is the audio data that has been transferred to the user's terminal. The server sends the audio data to the terminal using the appropriate protocol.

[0321] Step 6:

[0322] The device plays the received audio data through the headset speaker. The input is the received audio data, and the output is the audio output to the user's hearing. The device decodes the audio file and plays the audio through the speaker.

[0323] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0324] This invention not only provides relevant geographical information based on the user's location, but also recognizes the user's emotions and dynamically adjusts the content provided accordingly. This system enables a more personalized travel experience for the user.

[0325] First, the user's device uses its GPS function to obtain its current location and sends this location information to the server. The server analyzes this location information and retrieves relevant geographical information from its database. At this time, a text guide is generated based on the retrieved geographical information.

[0326] Next, the user's device uses its built-in sensors, camera, and microphone to analyze the user's emotional state with an emotion engine. This is based on facial expression analysis, voice tone recognition, and physiological data obtained from sensors.

[0327] The server uses data obtained from the emotion engine to select or adjust text content according to the user's current emotions. For example, if the server detects that the user is excited, it might prioritize providing information about activities and adventures. This text content is then converted into audio data using a text-to-speech engine.

[0328] This audio data is sent from the server to the user's device and played back on the user's device. This allows the user to receive information appropriate to their emotional state at that time in real time. Furthermore, through a feedback function, the suitability of the information to the emotions is evaluated, and the system is continuously optimized.

[0329] As a concrete example, when a user visits a tourist destination, they can receive different information depending on their mood and reactions that day. If the server detects that the user is seeking relaxation, it will provide information about quiet parks and cafes in the area, along with related historical background. In this way, users can always obtain valuable information that matches their current emotional state.

[0330] The following describes the processing flow.

[0331] Step 1:

[0332] The user's device obtains its current location information using a GPS sensor. This location information is converted into a digital format as latitude and longitude.

[0333] Step 2:

[0334] The user's device sends the acquired location information to the server. The data is encrypted using a secure protocol before transmission.

[0335] Step 3:

[0336] The server uses the received location information to search its database for relevant geographical information. During this process, tourist attractions and historical information are extracted.

[0337] Step 4:

[0338] The microphone and camera on the user's device are used, and an emotion engine analyzes the user's emotional state. This analysis includes voice tone and facial recognition.

[0339] Step 5:

[0340] The server receives emotion data from the emotion engine and generates or adjusts text content according to the user's emotions. This content will be tailored to the user's interests and emotions.

[0341] Step 6:

[0342] The server uses a text-to-speech engine to convert the edited text content into audio data. This audio data is processed in a format that is comfortable for the user to listen to.

[0343] Step 7:

[0344] The server sends the generated audio data to the user's device. The transmitted data is delivered quickly and securely.

[0345] Step 8:

[0346] The user's device receives audio data and plays it back through the built-in speaker or connected earphones. The user can hear the guidance information directly.

[0347] Step 9:

[0348] Users provide feedback on the information they receive. This feedback includes evaluations of the information's content and its relevance to their emotional state.

[0349] Step 10:

[0350] The server analyzes user feedback and optimizes the entire system to improve the personalization accuracy of future information delivery. This process continuously improves the learning database.

[0351] (Example 2)

[0352] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0353] Conventional geographic information systems often provide certain information based on the user's location, but they do not take into account the user's individual emotional state. As a result, there is a lack of personalization of the user experience, and the content provided may not match the user's current emotions.

[0354] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0355] In this invention, the server includes means for acquiring the user's location information, means for retrieving relevant geographical information from a data storage device, and means for analyzing the user's emotional state using an emotion engine. This enables the provision of appropriate and personalized information according to the user's current emotional state.

[0356] "User location information" refers to coordinate data indicating the user's current location, and is typically location information obtained using the Global Positioning Satellite System (GPS).

[0357] "Geographical information" refers to various types of information related to a specific location, including data about the region such as tourist attractions, history, culture, and facilities.

[0358] A "data storage device" is a storage device or database that stores a wide variety of information, such as geographical information, and allows it to be retrieved as needed.

[0359] A "text guide" refers to information presented to users in the form of explanatory or instructive text, which is generated based on geographical information.

[0360] An "emotion engine" is a program or system for analyzing a user's emotional state, using facial expressions, voice, sensor data, etc., to infer emotions.

[0361] "Audio data" refers to digital audio data that is created by converting text content into speech and is intended to be played back on a user's device.

[0362] A "speech synthesis engine" is a program or system for converting text content into speech data, and is a technology for generating artificial speech.

[0363] "Responses" refer to feedback and behavioral data collected from users, which can be used to improve system performance and personalize the system.

[0364] This invention relates to a system that dynamically provides personalized content by utilizing the user's location information and emotions. This system is mainly composed of a combination of a server, a terminal, and various software engines.

[0365] First, the user obtains location information using their device. The device utilizes its built-in GPS sensor to measure latitude and longitude data. This location information is transmitted to the server in real time.

[0366] The server retrieves relevant geographical information from the data storage device based on the received location information. A database management system is used in this process to efficiently obtain geographical information.

[0367] Next, the device uses its built-in camera, microphone, and sensors to analyze the user's emotions in real time. The emotion engine processes data from various sensors and applies a dedicated analysis algorithm to infer the emotional state. This analysis includes facial recognition technology and voice analysis technology.

[0368] Based on the output from the emotion engine, the server dynamically generates or adjusts content according to the user's emotions. Appropriately selected text information is converted into audio data by the speech synthesis engine. The speech synthesis engine uses a specified speech model to convert the text into smooth speech.

[0369] Audio data is sent from the server to the user's device, which then plays the audio using its built-in media player. This allows the user to receive personalized information in real time, based on their location and emotions.

[0370] As a concrete example, consider a scenario where a user is visiting a tourist spot in a city. If facial analysis determines that the user is relaxed, the server will prioritize providing information about quiet places and cafes in that location. Related historical information will also be provided via audio, allowing the user to efficiently obtain information while walking.

[0371] An example of a prompt to a generative AI model is the instruction, "Generate tourist information suitable for when the user is in an excited state." In response to this prompt, the system will prioritize generating information about adventurous activities.

[0372] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0373] Step 1:

[0374] The user's device obtains its current location using GPS functionality. The device inputs latitude and longitude data from the Global Positioning Satellite System (GSO) and transmits this information to the server. The obtained location information serves as the basis for searching for geographical information about the user's surroundings.

[0375] Step 2:

[0376] The server uses location information received from the terminal as input to query the data storage device and retrieve relevant geographical information. Using a database management system, it searches for tourist attractions, facilities, and other regional information, and prepares the retrieved geographical information for use in the next step.

[0377] Step 3:

[0378] The user's device uses its built-in camera, microphone, and sensors to collect data for input into the emotion engine. This data includes facial image data, voice tone, and physiological data. The emotion engine processes this data to determine the user's current emotional state, using image processing algorithms and voice analysis models.

[0379] Step 4:

[0380] The server receives emotional state data output from the emotion engine as input and dynamically generates personalized content. This process uses a generative AI model to create text and information that matches the user's emotions. The generated text content is adjusted according to the user's interests and emotions.

[0381] Step 5:

[0382] The server inputs the selected or edited text content into a speech synthesis engine and converts it into speech data. The synthesized speech is used to generate smooth speech data from the text. Here, a speech synthesis algorithm is applied to output speech in an easy-to-understand format.

[0383] Step 6:

[0384] The server sends the generated audio data to the user's device. The user's device uses its built-in media player to play the received audio data. This allows the user to hear appropriate information in real time, tailored to their location and emotions.

[0385] (Application Example 2)

[0386] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0387] Conventional geographic information systems provide information based solely on the user's location, making it difficult to deliver content tailored to the user's psychological state, such as emotions, stress levels, and curiosity, in real time. Furthermore, a lack of continuous content optimization based on user feedback and individual needs for tourism information was a significant challenge.

[0388] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0389] In this invention, the server includes means for acquiring the user's location information, means for analyzing the acquired user's emotional state, and means for generating personalized audio content based on the emotional state and location information. This makes it possible to provide personalized geographical information and audio guidance that corresponds to the user's psychological state.

[0390] "Means for acquiring user location information" refers to a device or method that has the function of determining the user's current geographical location and acquiring it as digital information.

[0391] "Means for retrieving geographical information from a data set" refers to a device or method that has the function of retrieving and extracting relevant geographical information from a database or the like based on acquired location data.

[0392] "Means for analyzing a user's emotional state" refers to a device or method that has the function of determining a user's psychological emotional state using the user's facial expressions, voice, and other physiological data.

[0393] "Means for generating personalized audio content" refers to a device or method that has the function of creating audio information tailored to individual needs based on the user's emotional state and location information.

[0394] The "feature for collecting user feedback" is a function that provides a mechanism for users to proactively submit their own opinions and suggestions for improvement regarding the information provided.

[0395] This invention is a system for providing personalized geographical information to users. The system mainly consists of a server and a user terminal.

[0396] First, the user's device acquires location information using a GPS module. This location information is sent to a server, which then searches its database for relevant geographical information based on this data. During this process, the server analyzes facial expression and voice data sent from the device to obtain the user's emotional state. Here, a facial expression recognition library (e.g., OpenCV) and voice tone recognition software are used.

[0397] The server uses acquired emotional state and location information to generate personalized audio content. This audio content is converted from text to audio data using a text-to-speech conversion device (e.g., Amazon Polly). The generated audio content is sent to the user's device and played back on the device. This allows the user to receive information in audio format that is tailored to their emotional state in real time.

[0398] Furthermore, it includes a function to collect feedback from users, and the system continuously optimizes content quality based on user feedback.

[0399] For example, if the system determines that a user has arrived at a tourist destination and is relaxing, it can provide information about nearby quiet parks or cafes and offer audio guidance about the historical background.

[0400] An example of a prompt for a generative AI model is: "Analyze the user's emotional state in real time based on their facial expressions and voice tone, and provide location-related tourist information in an engaging narration."

[0401] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0402] Step 1:

[0403] The device uses a GPS module to obtain the user's current location. The obtained location information is sent to the server. The input is the location information, and the output is the location information sent to the server.

[0404] Step 2:

[0405] The server searches the database for relevant geographical information based on the received location information. The input is the transmitted location information, and the output is the retrieved geographical information. The server executes queries and extracts the geographical information.

[0406] Step 3:

[0407] The device uses its camera and microphone to collect the user's facial expressions and voice, and generates emotion data. The input is the user's facial expressions and voice data, and the output is the generated emotion data. The device performs facial recognition and voice analysis.

[0408] Step 4:

[0409] The server analyzes emotional data to identify the user's emotional state. The input is emotional data, and the output is the identified emotional state. The server utilizes a facial expression recognition library and voice tone analysis software.

[0410] Step 5:

[0411] The server generates personalized text content based on identified emotional states and retrieved geographical information. The input is emotional state and geographical information, and the output is text content. The server selects information that matches the user's preferences.

[0412] Step 6:

[0413] The server converts the generated text content into audio data using a text-to-speech conversion processing unit. The input is text content, and the output is audio data. The server performs the conversion using a text-to-speech engine.

[0414] Step 7:

[0415] The server sends the generated audio data to the terminal. The input is the audio data, and the output is the audio data sent to the terminal. The server transfers the data.

[0416] Step 8:

[0417] The device plays the received audio data. The input is the received audio data, and the output is the audio guide provided to the user. The device activates the audio playback function.

[0418] Step 9:

[0419] Users provide feedback on the information they receive. The input is the information provided, and the output is the user's feedback. Users send their feedback via text or voice.

[0420] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0421] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0422] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0423] [Third Embodiment]

[0424] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0425] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0426] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0427] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0428] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0429] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0430] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0431] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0432] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0433] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0434] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0435] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0436] This invention is a system for providing relevant geographical information based on the user's location and guiding the user through it via voice. In this embodiment, the aim is to allow the user to obtain information in real time using their handheld device, thereby making the travel experience deeper and richer.

[0437] The user's device uses GPS functionality to determine its current location and sends this location information to the server. The server receives this information and searches its database for relevant geographical information. This allows the server to identify tourist information and historical background relevant to the user's current location.

[0438] Next, the server uses the acquired data to create a text guide for the user and generates audio content using a text-to-speech engine. This audio content is converted to an appropriate data format so that the user can listen to it through their device, and then transferred to the user's device.

[0439] The user's device receives this audio content and plays it back through speakers or headphones. This playback allows the user to hear detailed tourist information on the spot, deepening their understanding of the location. Furthermore, the server collects the user's feedback and uses it to improve future information provision. This feedback function allows the system to continuously optimize the user experience.

[0440] As a concrete example, if a user is in a tourist destination, the server can provide information about local historical events and important figures, and can also play quiz-style questions in audio format. For example, it could ask, "Guess what happened in this place in the 19th century and who was involved." In this way, users can gain high educational value even while on the go.

[0441] The following describes the processing flow.

[0442] Step 1:

[0443] The user's device uses its built-in GPS function to obtain its current location. Location information is recorded as latitude and longitude.

[0444] Step 2:

[0445] The user's device transmits the acquired location information to the server using a secure protocol. The location information is encrypted to protect the data.

[0446] Step 3:

[0447] The server analyzes location information received from the user's device. Based on this information, it searches the database for relevant geographical information, tourist attractions, and historical background.

[0448] Step 4:

[0449] The server generates text guides from the geographical information obtained through searches. These guides include tourist information and descriptions of specific historical events.

[0450] Step 5:

[0451] The server uses a text-to-speech engine to convert the text guide into audio data. This audio data is then formatted to be easily understood by the user.

[0452] Step 6:

[0453] The server sends the generated audio data to the user's terminal. During data transfer, error detection and correction are performed to maintain the reliability of the communication.

[0454] Step 7:

[0455] The user's device plays the received audio data. Audio guidance is provided through speakers or headphones, allowing the user to hear the instructions in real time.

[0456] Step 8:

[0457] Users provide feedback on the information provided through the application. This feedback is sent to the server to optimize future information provision.

[0458] (Example 1)

[0459] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0460] Modern travelers demand detailed and intuitive information about their destinations, but traditional voice guidance systems struggle to fully customize based on location data and have limitations in providing information that reflects user feedback, thus failing to enrich the user experience.

[0461] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0462] In this invention, the server includes an information device for acquiring location information, an information processing device for retrieving relevant geographic information from an information storage device based on the location information, and an information conversion device for generating voice information based on the retrieved geographic information. This makes it possible to provide voice guidance optimized for individual users and to continuously improve the quality of information provision based on feedback.

[0463] "Location information" refers to digital data that represents the longitude and latitude of a specific point on Earth.

[0464] An "information device" is an electronic device that enables the acquisition, processing, and transmission of data.

[0465] An "information storage device" is a recording medium that includes a database and is capable of storing and retrieving diverse types of data.

[0466] An "information processing device" is a computer system that analyzes received data and extracts and provides relevant information.

[0467] An "information conversion device" is a device that has the function of converting data in a specific format into another format.

[0468] A "communication device" is a combination of hardware and software used to send and receive data between different electronic devices.

[0469] An "information terminal" is a device that allows users to interactively receive and display information.

[0470] A "playback device" is a device that represents digital content such as audio and video as a physical output.

[0471] An "information gathering device" is a system that receives input data from users and transmits it to a server.

[0472] This invention is a system that provides users with detailed geographical and historical information about a local area, which they can then listen to as audio guidance. The user uses a mobile device equipped with GPS to determine their current location. The device then transmits this location information to a server.

[0473] The server searches its internal database for relevant geographical information based on the received location data. The database contains information on tourist attractions and historical events related to geographical locations. The retrieved information is extracted as organized text data.

[0474] The server then converts the text data into speech data using a text-to-speech (TTS) engine. While general speech synthesis techniques are used here, open-source TTS engines are available as specific examples.

[0475] The generated audio data is sent from the server to the user's mobile device. After receiving this audio data, the device plays it back to the user through its speaker or headphones.

[0476] As a concrete example, when a user visits a historical tourist site, the server will provide audio information, including historical events in that area, so that the user can perceive that information while on site. This system aims to enhance the user's travel experience through efficient information provision.

[0477] An example of a prompt message might be: "We want to design an audio guidance system to provide historical background and tourist information related to a geographically specified location for commercial purposes. Please list the main functions required and explain the overall flow of the system."

[0478] This invention also includes a function to improve the accuracy of information provided based on user feedback, thereby continuously improving user satisfaction.

[0479] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0480] Step 1:

[0481] The user enables location services on their mobile device. This allows the device to obtain latitude and longitude via its GPS module. The input data is the local latitude and longitude collected by the device. The output is the acquired location information temporarily stored on the device.

[0482] Step 2:

[0483] The device transmits the collected location information to the server. Here, the latitude and longitude stored on the device are used as input, and a location information packet is generated as output for the server to receive. This process is performed via the HTTPS protocol, ensuring security.

[0484] Step 3:

[0485] The server accesses a database based on the received location information and searches for relevant geographical information. The input data is the latitude and longitude sent to the server. The server converts this into a database query and retrieves information about tourist destinations and historical events as output.

[0486] Step 4:

[0487] The server analyzes acquired geographical information and generates a text guide organized for the user. Raw data retrieved from a database is used as input, and concise, easy-to-understand text information is generated as output. This text information is customized based on the user's location and interests.

[0488] Step 5:

[0489] The server passes the organized text information to the text-to-speech engine, which converts it into audio data. The input data is the generated text guide. Based on this, audio data is generated and output as an MP3 audio file.

[0490] Step 6:

[0491] The server sends the generated audio file to the user's terminal. The input is the generated audio file, and the output is the audio data received by the user's terminal. This data is transferred using standard data streaming technology.

[0492] Step 7:

[0493] The terminal plays the received audio data through a speaker or headphones. The input data is an audio file received from the server, and the output is the voice guidance that the user hears. This playback function allows the user to receive real-time voice guidance on the spot.

[0494] Step 8:

[0495] Users can provide feedback on voice guidance through their devices. The input data is the feedback information entered by the user, and the output is the feedback data being sent to the server and used to improve future information provision.

[0496] (Application Example 1)

[0497] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0498] There is a need for safer and more effective ways for users of autonomous vehicles to obtain information about their surroundings while operating them. Conventional systems mainly rely on visual information provision, which can sometimes lead to situations where users are unable to concentrate on driving. It is necessary to solve these problems and improve safety and convenience for users.

[0499] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0500] In this invention, the server includes means for recognizing the user's geographical location signal, means for retrieving relevant regional information from a data storage area based on the geographical location signal, and means for generating voice output based on the retrieved regional information. This allows users in autonomous mobile vehicles to receive regional information by voice while moving via a head-mounted display device, enabling them to acquire information without relying on vision, thereby improving user safety and the efficiency of information provision.

[0501] "User geographic location signal" refers to a digital signal that indicates the user's current location, obtained using GPS or other location-determining technologies.

[0502] "Relevant local information" refers to information such as tourist destinations, historical facts, and events within a specific geographical area, and is considered useful data for users.

[0503] A "data storage area" is a digital storage system for accumulating information and is a database used to enable location-based searches.

[0504] "Audio output" refers to electronic signals that convert text information into sound, providing information to the user through hearing.

[0505] A "user terminal" is a device that has the function of receiving and playing audio, and includes smart glasses and personal digital assistants.

[0506] An "autonomous mobile vehicle" is a vehicle that can move on its own without human intervention, and is typically operated using AI and various sensors.

[0507] A "head-mounted display device" is a device worn on the head and designed to provide information by combining visual and auditory elements.

[0508] This invention is a system for acquiring a user's geographical location signal and providing relevant regional information based on that information. The system uses a GPS module to recognize the user's location and transmits the corresponding location information to a server. After receiving this, the server retrieves relevant regional information from its data storage area and generates voice output based on that information. The voice output is generated in real time using a string-to-speech engine and provided to the user via their user terminal, specifically a head-mounted display device.

[0509] The server uses the Python library `requests` to retrieve necessary information from external services. This data is then converted into speech format using the `text_to_speech` library. The head-mounted display device, acting as the user terminal, has a speaker to play the received audio data, providing information to the user without relying on their vision.

[0510] As a concrete example, if a user is traveling through a tourist destination in an autonomous vehicle, the server can provide audio information about the area's historical background and landmarks. For instance, information such as, "This region has long been known as a hot spring resort and is a popular destination visited by many tourists," could be provided. In this way, users can safely receive rich information while minimizing their reliance on visual cues. An example of a prompt to input into the generating AI model would be, "Please obtain information about the surrounding tourist attractions based on the current location and provide it as an audio guide."

[0511] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0512] Step 1:

[0513] The user obtains their current geographical location signal using a device. The input is location information, received from a GPS module. The output is latitude and longitude data. The device sends this data to the server.

[0514] Step 2:

[0515] The server retrieves relevant regional information from its data storage area based on the latitude and longitude data received from the terminal. The input is latitude and longitude, and the output is a dataset of regional information. The server filters the information by sending queries to the database.

[0516] Step 3:

[0517] The server creates descriptive text strings based on the acquired regional information dataset. The input is a regional information dataset, and the output is descriptive text. The server combines the data to generate structured sentences.

[0518] Step 4:

[0519] The server generates audio data using a string-to-speech engine to convert explanatory text into speech. The input is explanatory text, and the output is an audio file. The server converts the string into a predetermined format and generates an audio waveform.

[0520] Step 5:

[0521] The server sends the generated audio data to the terminal. The input is an audio file, and the output is the audio data that has been transferred to the user's terminal. The server sends the audio data to the terminal using the appropriate protocol.

[0522] Step 6:

[0523] The device plays the received audio data through the headset speaker. The input is the received audio data, and the output is the audio output to the user's hearing. The device decodes the audio file and plays the audio through the speaker.

[0524] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0525] This invention not only provides relevant geographical information based on the user's location, but also recognizes the user's emotions and dynamically adjusts the content provided accordingly. This system enables a more personalized travel experience for the user.

[0526] First, the user's device uses its GPS function to obtain its current location and sends this location information to the server. The server analyzes this location information and retrieves relevant geographical information from its database. At this time, a text guide is generated based on the retrieved geographical information.

[0527] Next, the user's device uses its built-in sensors, camera, and microphone to analyze the user's emotional state with an emotion engine. This is based on facial expression analysis, voice tone recognition, and physiological data obtained from sensors.

[0528] The server uses data obtained from the emotion engine to select or adjust text content according to the user's current emotions. For example, if the server detects that the user is excited, it might prioritize providing information about activities and adventures. This text content is then converted into audio data using a text-to-speech engine.

[0529] This audio data is sent from the server to the user's device and played back on the user's device. This allows the user to receive information appropriate to their emotional state at that time in real time. Furthermore, through a feedback function, the suitability of the information to the emotions is evaluated, and the system is continuously optimized.

[0530] As a concrete example, when a user visits a tourist destination, they can receive different information depending on their mood and reactions that day. If the server detects that the user is seeking relaxation, it will provide information about quiet parks and cafes in the area, along with related historical background. In this way, users can always obtain valuable information that matches their current emotional state.

[0531] The following describes the processing flow.

[0532] Step 1:

[0533] The user's device obtains its current location information using a GPS sensor. This location information is converted into a digital format as latitude and longitude.

[0534] Step 2:

[0535] The user's device sends the acquired location information to the server. The data is encrypted using a secure protocol before transmission.

[0536] Step 3:

[0537] The server uses the received location information to search its database for relevant geographical information. During this process, tourist attractions and historical information are extracted.

[0538] Step 4:

[0539] The microphone and camera on the user's device are used, and an emotion engine analyzes the user's emotional state. This analysis includes voice tone and facial recognition.

[0540] Step 5:

[0541] The server receives emotion data from the emotion engine and generates or adjusts text content according to the user's emotions. This content will be tailored to the user's interests and emotions.

[0542] Step 6:

[0543] The server uses a text-to-speech engine to convert the edited text content into audio data. This audio data is processed in a format that is comfortable for the user to listen to.

[0544] Step 7:

[0545] The server sends the generated audio data to the user's device. The transmitted data is delivered quickly and securely.

[0546] Step 8:

[0547] The user's device receives audio data and plays it back through the built-in speaker or connected earphones. The user can hear the guidance information directly.

[0548] Step 9:

[0549] Users provide feedback on the information they receive. This feedback includes evaluations of the information's content and its relevance to their emotional state.

[0550] Step 10:

[0551] The server analyzes user feedback and optimizes the entire system to improve the personalization accuracy of future information delivery. This process continuously improves the learning database.

[0552] (Example 2)

[0553] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0554] Conventional geographic information systems often provide certain information based on the user's location, but they do not take into account the user's individual emotional state. As a result, there is a lack of personalization of the user experience, and the content provided may not match the user's current emotions.

[0555] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0556] In this invention, the server includes means for acquiring the user's location information, means for retrieving relevant geographical information from a data storage device, and means for analyzing the user's emotional state using an emotion engine. This enables the provision of appropriate and personalized information according to the user's current emotional state.

[0557] "User location information" refers to coordinate data indicating the user's current location, and is typically location information obtained using the Global Positioning Satellite System (GPS).

[0558] "Geographical information" refers to various types of information related to a specific location, including data about the region such as tourist attractions, history, culture, and facilities.

[0559] A "data storage device" is a storage device or database that stores a wide variety of information, such as geographical information, and allows it to be retrieved as needed.

[0560] A "text guide" refers to information presented to users in the form of explanatory or instructive text, which is generated based on geographical information.

[0561] An "emotion engine" is a program or system for analyzing a user's emotional state, using facial expressions, voice, sensor data, etc., to infer emotions.

[0562] "Audio data" refers to digital audio data that is created by converting text content into speech and is intended to be played back on a user's device.

[0563] A "speech synthesis engine" is a program or system for converting text content into speech data, and is a technology for generating artificial speech.

[0564] "Responses" refer to feedback and behavioral data collected from users, which can be used to improve system performance and personalize the system.

[0565] This invention relates to a system that dynamically provides personalized content by utilizing the user's location information and emotions. This system is mainly composed of a combination of a server, a terminal, and various software engines.

[0566] First, the user obtains location information using their device. The device utilizes its built-in GPS sensor to measure latitude and longitude data. This location information is transmitted to the server in real time.

[0567] The server retrieves relevant geographical information from the data storage device based on the received location information. A database management system is used in this process to efficiently obtain geographical information.

[0568] Next, the device uses its built-in camera, microphone, and sensors to analyze the user's emotions in real time. The emotion engine processes data from various sensors and applies a dedicated analysis algorithm to infer the emotional state. This analysis includes facial recognition technology and voice analysis technology.

[0569] Based on the output from the emotion engine, the server dynamically generates or adjusts content according to the user's emotions. Appropriately selected text information is converted into audio data by the speech synthesis engine. The speech synthesis engine uses a specified speech model to convert the text into smooth speech.

[0570] Audio data is sent from the server to the user's device, which then plays the audio using its built-in media player. This allows the user to receive personalized information in real time, based on their location and emotions.

[0571] As a concrete example, consider a scenario where a user is visiting a tourist spot in a city. If facial analysis determines that the user is relaxed, the server will prioritize providing information about quiet places and cafes in that location. Related historical information will also be provided via audio, allowing the user to efficiently obtain information while walking.

[0572] An example of a prompt to a generative AI model is the instruction, "Generate tourist information suitable for when the user is in an excited state." In response to this prompt, the system will prioritize generating information about adventurous activities.

[0573] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0574] Step 1:

[0575] The user's device obtains its current location using GPS functionality. The device inputs latitude and longitude data from the Global Positioning Satellite System (GSO) and transmits this information to the server. The obtained location information serves as the basis for searching for geographical information about the user's surroundings.

[0576] Step 2:

[0577] The server uses location information received from the terminal as input to query the data storage device and retrieve relevant geographical information. Using a database management system, it searches for tourist attractions, facilities, and other regional information, and prepares the retrieved geographical information for use in the next step.

[0578] Step 3:

[0579] The user's device uses its built-in camera, microphone, and sensors to collect data for input into the emotion engine. This data includes facial image data, voice tone, and physiological data. The emotion engine processes this data to determine the user's current emotional state, using image processing algorithms and voice analysis models.

[0580] Step 4:

[0581] The server receives emotional state data output from the emotion engine as input and dynamically generates personalized content. This process uses a generative AI model to create text and information that matches the user's emotions. The generated text content is adjusted according to the user's interests and emotions.

[0582] Step 5:

[0583] The server inputs the selected or edited text content into a speech synthesis engine and converts it into speech data. The synthesized speech is used to generate smooth speech data from the text. Here, a speech synthesis algorithm is applied to output speech in an easy-to-understand format.

[0584] Step 6:

[0585] The server sends the generated audio data to the user's device. The user's device uses its built-in media player to play the received audio data. This allows the user to hear appropriate information in real time, tailored to their location and emotions.

[0586] (Application Example 2)

[0587] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0588] Conventional geographic information systems provide information based solely on the user's location, making it difficult to deliver content tailored to the user's psychological state, such as emotions, stress levels, and curiosity, in real time. Furthermore, a lack of continuous content optimization based on user feedback and individual needs for tourism information was a significant challenge.

[0589] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0590] In this invention, the server includes means for acquiring the user's location information, means for analyzing the acquired user's emotional state, and means for generating personalized audio content based on the emotional state and location information. This makes it possible to provide personalized geographical information and audio guidance that corresponds to the user's psychological state.

[0591] "Means for acquiring user location information" refers to a device or method that has the function of determining the user's current geographical location and acquiring it as digital information.

[0592] "Means for retrieving geographical information from a data set" refers to a device or method that has the function of retrieving and extracting relevant geographical information from a database or the like based on acquired location data.

[0593] "Means for analyzing a user's emotional state" refers to a device or method that has the function of determining a user's psychological emotional state using the user's facial expressions, voice, and other physiological data.

[0594] "Means for generating personalized audio content" refers to a device or method that has the function of creating audio information tailored to individual needs based on the user's emotional state and location information.

[0595] The "feature for collecting user feedback" is a function that provides a mechanism for users to proactively submit their own opinions and suggestions for improvement regarding the information provided.

[0596] This invention is a system for providing personalized geographical information to users. The system mainly consists of a server and a user terminal.

[0597] First, the user's device acquires location information using a GPS module. This location information is sent to a server, which then searches its database for relevant geographical information based on this data. During this process, the server analyzes facial expression and voice data sent from the device to obtain the user's emotional state. Here, a facial expression recognition library (e.g., OpenCV) and voice tone recognition software are used.

[0598] The server uses acquired emotional state and location information to generate personalized audio content. This audio content is converted from text to audio data using a text-to-speech conversion device (e.g., Amazon Polly). The generated audio content is sent to the user's device and played back on the device. This allows the user to receive information in audio format that is tailored to their emotional state in real time.

[0599] Furthermore, it includes a function to collect feedback from users, and the system continuously optimizes content quality based on user feedback.

[0600] For example, if the system determines that a user has arrived at a tourist destination and is relaxing, it can provide information about nearby quiet parks or cafes and offer audio guidance about the historical background.

[0601] An example of a prompt for a generative AI model is: "Analyze the user's emotional state in real time based on their facial expressions and voice tone, and provide location-related tourist information in an engaging narration."

[0602] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0603] Step 1:

[0604] The device uses a GPS module to obtain the user's current location. The obtained location information is sent to the server. The input is the location information, and the output is the location information sent to the server.

[0605] Step 2:

[0606] The server searches the database for relevant geographical information based on the received location information. The input is the transmitted location information, and the output is the retrieved geographical information. The server executes queries and extracts the geographical information.

[0607] Step 3:

[0608] The device uses its camera and microphone to collect the user's facial expressions and voice, and generates emotion data. The input is the user's facial expressions and voice data, and the output is the generated emotion data. The device performs facial recognition and voice analysis.

[0609] Step 4:

[0610] The server analyzes emotional data to identify the user's emotional state. The input is emotional data, and the output is the identified emotional state. The server utilizes a facial expression recognition library and voice tone analysis software.

[0611] Step 5:

[0612] The server generates personalized text content based on identified emotional states and retrieved geographical information. The input is emotional state and geographical information, and the output is text content. The server selects information that matches the user's preferences.

[0613] Step 6:

[0614] The server converts the generated text content into audio data using a text-to-speech conversion processing unit. The input is text content, and the output is audio data. The server performs the conversion using a text-to-speech engine.

[0615] Step 7:

[0616] The server sends the generated audio data to the terminal. The input is the audio data, and the output is the audio data sent to the terminal. The server transfers the data.

[0617] Step 8:

[0618] The device plays the received audio data. The input is the received audio data, and the output is the audio guide provided to the user. The device activates the audio playback function.

[0619] Step 9:

[0620] Users provide feedback on the information they receive. The input is the information provided, and the output is the user's feedback. Users send their feedback via text or voice.

[0621] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0622] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0623] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0624] [Fourth Embodiment]

[0625] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0626] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0627] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0628] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0629] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0630] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0631] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0632] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0633] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0634] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0635] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0636] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0637] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0638] This invention is a system for providing relevant geographical information based on the user's location and guiding the user through it via voice. In this embodiment, the aim is to allow the user to obtain information in real time using their handheld device, thereby making the travel experience deeper and richer.

[0639] The user's device uses GPS functionality to determine its current location and sends this location information to the server. The server receives this information and searches its database for relevant geographical information. This allows the server to identify tourist information and historical background relevant to the user's current location.

[0640] Next, the server uses the acquired data to create a text guide for the user and generates audio content using a text-to-speech engine. This audio content is converted to an appropriate data format so that the user can listen to it through their device, and then transferred to the user's device.

[0641] The user's device receives this audio content and plays it back through speakers or headphones. This playback allows the user to hear detailed tourist information on the spot, deepening their understanding of the location. Furthermore, the server collects the user's feedback and uses it to improve future information provision. This feedback function allows the system to continuously optimize the user experience.

[0642] As a concrete example, if a user is in a tourist destination, the server can provide information about local historical events and important figures, and can also play quiz-style questions in audio format. For example, it could ask, "Guess what happened in this place in the 19th century and who was involved." In this way, users can gain high educational value even while on the go.

[0643] The following describes the processing flow.

[0644] Step 1:

[0645] The user's device uses its built-in GPS function to obtain its current location. Location information is recorded as latitude and longitude.

[0646] Step 2:

[0647] The user's device transmits the acquired location information to the server using a secure protocol. The location information is encrypted to protect the data.

[0648] Step 3:

[0649] The server analyzes location information received from the user's device. Based on this information, it searches the database for relevant geographical information, tourist attractions, and historical background.

[0650] Step 4:

[0651] The server generates text guides from the geographical information obtained through searches. These guides include tourist information and descriptions of specific historical events.

[0652] Step 5:

[0653] The server uses a text-to-speech engine to convert the text guide into audio data. This audio data is then formatted to be easily understood by the user.

[0654] Step 6:

[0655] The server sends the generated audio data to the user's terminal. During data transfer, error detection and correction are performed to maintain the reliability of the communication.

[0656] Step 7:

[0657] The user's device plays the received audio data. Audio guidance is provided through speakers or headphones, allowing the user to hear the instructions in real time.

[0658] Step 8:

[0659] Users provide feedback on the information provided through the application. This feedback is sent to the server to optimize future information provision.

[0660] (Example 1)

[0661] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0662] Modern travelers demand detailed and intuitive information about their destinations, but traditional voice guidance systems struggle to fully customize based on location data and have limitations in providing information that reflects user feedback, thus failing to enrich the user experience.

[0663] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0664] In this invention, the server includes an information device for acquiring location information, an information processing device for retrieving relevant geographic information from an information storage device based on the location information, and an information conversion device for generating voice information based on the retrieved geographic information. This makes it possible to provide voice guidance optimized for individual users and to continuously improve the quality of information provision based on feedback.

[0665] "Location information" refers to digital data that represents the longitude and latitude of a specific point on Earth.

[0666] An "information device" is an electronic device that enables the acquisition, processing, and transmission of data.

[0667] An "information storage device" is a recording medium that includes a database and is capable of storing and retrieving diverse types of data.

[0668] An "information processing device" is a computer system that analyzes received data and extracts and provides relevant information.

[0669] An "information conversion device" is a device that has the function of converting data in a specific format into another format.

[0670] A "communication device" is a combination of hardware and software used to send and receive data between different electronic devices.

[0671] An "information terminal" is a device that allows users to interactively receive and display information.

[0672] A "playback device" is a device that represents digital content such as audio and video as a physical output.

[0673] An "information gathering device" is a system that receives input data from users and transmits it to a server.

[0674] This invention is a system that provides users with detailed geographical and historical information about a local area, which they can then listen to as audio guidance. The user uses a mobile device equipped with GPS to determine their current location. The device then transmits this location information to a server.

[0675] The server searches its internal database for relevant geographical information based on the received location data. The database contains information on tourist attractions and historical events related to geographical locations. The retrieved information is extracted as organized text data.

[0676] The server then converts the text data into speech data using a text-to-speech (TTS) engine. While general speech synthesis techniques are used here, open-source TTS engines are available as specific examples.

[0677] The generated audio data is sent from the server to the user's mobile device. After receiving this audio data, the device plays it back to the user through its speaker or headphones.

[0678] As a concrete example, when a user visits a historical tourist site, the server will provide audio information, including historical events in that area, so that the user can perceive that information while on site. This system aims to enhance the user's travel experience through efficient information provision.

[0679] An example of a prompt message might be: "We want to design an audio guidance system to provide historical background and tourist information related to a geographically specified location for commercial purposes. Please list the main functions required and explain the overall flow of the system."

[0680] This invention also includes a function to improve the accuracy of information provided based on user feedback, thereby continuously improving user satisfaction.

[0681] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0682] Step 1:

[0683] The user enables location services on their mobile device. This allows the device to obtain latitude and longitude via its GPS module. The input data is the local latitude and longitude collected by the device. The output is the acquired location information temporarily stored on the device.

[0684] Step 2:

[0685] The device transmits the collected location information to the server. Here, the latitude and longitude stored on the device are used as input, and a location information packet is generated as output for the server to receive. This process is performed via the HTTPS protocol, ensuring security.

[0686] Step 3:

[0687] The server accesses a database based on the received location information and searches for relevant geographical information. The input data is the latitude and longitude sent to the server. The server converts this into a database query and retrieves information about tourist destinations and historical events as output.

[0688] Step 4:

[0689] The server analyzes acquired geographical information and generates a text guide organized for the user. Raw data retrieved from a database is used as input, and concise, easy-to-understand text information is generated as output. This text information is customized based on the user's location and interests.

[0690] Step 5:

[0691] The server passes the organized text information to the text-to-speech engine, which converts it into audio data. The input data is the generated text guide. Based on this, audio data is generated and output as an MP3 audio file.

[0692] Step 6:

[0693] The server sends the generated audio file to the user's terminal. The input is the generated audio file, and the output is the audio data received by the user's terminal. This data is transferred using standard data streaming technology.

[0694] Step 7:

[0695] The terminal plays the received audio data through a speaker or headphones. The input data is an audio file received from the server, and the output is the voice guidance that the user hears. This playback function allows the user to receive real-time voice guidance on the spot.

[0696] Step 8:

[0697] Users can provide feedback on voice guidance through their devices. The input data is the feedback information entered by the user, and the output is the feedback data being sent to the server and used to improve future information provision.

[0698] (Application Example 1)

[0699] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0700] There is a need for safer and more effective ways for users of autonomous vehicles to obtain information about their surroundings while operating them. Conventional systems mainly rely on visual information provision, which can sometimes lead to situations where users are unable to concentrate on driving. It is necessary to solve these problems and improve safety and convenience for users.

[0701] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0702] In this invention, the server includes means for recognizing the user's geographical location signal, means for retrieving relevant regional information from a data storage area based on the geographical location signal, and means for generating voice output based on the retrieved regional information. This allows users in autonomous mobile vehicles to receive regional information by voice while moving via a head-mounted display device, enabling them to acquire information without relying on vision, thereby improving user safety and the efficiency of information provision.

[0703] "User geographic location signal" refers to a digital signal that indicates the user's current location, obtained using GPS or other location-determining technologies.

[0704] "Relevant local information" refers to information such as tourist destinations, historical facts, and events within a specific geographical area, and is considered useful data for users.

[0705] A "data storage area" is a digital storage system for accumulating information and is a database used to enable location-based searches.

[0706] "Audio output" refers to electronic signals that convert text information into sound, providing information to the user through hearing.

[0707] A "user terminal" is a device that has the function of receiving and playing audio, and includes smart glasses and personal digital assistants.

[0708] An "autonomous mobile vehicle" is a vehicle that can move on its own without human intervention, and is typically operated using AI and various sensors.

[0709] A "head-mounted display device" is a device worn on the head and designed to provide information by combining visual and auditory elements.

[0710] This invention is a system for acquiring a user's geographical location signal and providing relevant regional information based on that information. The system uses a GPS module to recognize the user's location and transmits the corresponding location information to a server. After receiving this, the server retrieves relevant regional information from its data storage area and generates voice output based on that information. The voice output is generated in real time using a string-to-speech engine and provided to the user via their user terminal, specifically a head-mounted display device.

[0711] The server uses the Python library `requests` to retrieve necessary information from external services. This data is then converted into speech format using the `text_to_speech` library. The head-mounted display device, acting as the user terminal, has a speaker to play the received audio data, providing information to the user without relying on their vision.

[0712] As a concrete example, if a user is traveling through a tourist destination in an autonomous vehicle, the server can provide audio information about the area's historical background and landmarks. For instance, information such as, "This region has long been known as a hot spring resort and is a popular destination visited by many tourists," could be provided. In this way, users can safely receive rich information while minimizing their reliance on visual cues. An example of a prompt to input into the generating AI model would be, "Please obtain information about the surrounding tourist attractions based on the current location and provide it as an audio guide."

[0713] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0714] Step 1:

[0715] The user obtains their current geographical location signal using a device. The input is location information, received from a GPS module. The output is latitude and longitude data. The device sends this data to the server.

[0716] Step 2:

[0717] The server retrieves relevant regional information from its data storage area based on the latitude and longitude data received from the terminal. The input is latitude and longitude, and the output is a dataset of regional information. The server filters the information by sending queries to the database.

[0718] Step 3:

[0719] The server creates descriptive text strings based on the acquired regional information dataset. The input is a regional information dataset, and the output is descriptive text. The server combines the data to generate structured sentences.

[0720] Step 4:

[0721] The server generates audio data using a string-to-speech engine to convert explanatory text into speech. The input is explanatory text, and the output is an audio file. The server converts the string into a predetermined format and generates an audio waveform.

[0722] Step 5:

[0723] The server sends the generated audio data to the terminal. The input is an audio file, and the output is the audio data that has been transferred to the user's terminal. The server sends the audio data to the terminal using the appropriate protocol.

[0724] Step 6:

[0725] The device plays the received audio data through the headset speaker. The input is the received audio data, and the output is the audio output to the user's hearing. The device decodes the audio file and plays the audio through the speaker.

[0726] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0727] This invention not only provides relevant geographical information based on the user's location, but also recognizes the user's emotions and dynamically adjusts the content provided accordingly. This system enables a more personalized travel experience for the user.

[0728] First, the user's device uses its GPS function to obtain its current location and sends this location information to the server. The server analyzes this location information and retrieves relevant geographical information from its database. At this time, a text guide is generated based on the retrieved geographical information.

[0729] Next, the user's device uses its built-in sensors, camera, and microphone to analyze the user's emotional state with an emotion engine. This is based on facial expression analysis, voice tone recognition, and physiological data obtained from sensors.

[0730] The server uses data obtained from the emotion engine to select or adjust text content according to the user's current emotions. For example, if the server detects that the user is excited, it might prioritize providing information about activities and adventures. This text content is then converted into audio data using a text-to-speech engine.

[0731] This audio data is sent from the server to the user's device and played back on the user's device. This allows the user to receive information appropriate to their emotional state at that time in real time. Furthermore, through a feedback function, the suitability of the information to the emotions is evaluated, and the system is continuously optimized.

[0732] As a concrete example, when a user visits a tourist destination, they can receive different information depending on their mood and reactions that day. If the server detects that the user is seeking relaxation, it will provide information about quiet parks and cafes in the area, along with related historical background. In this way, users can always obtain valuable information that matches their current emotional state.

[0733] The following describes the processing flow.

[0734] Step 1:

[0735] The user's device obtains its current location information using a GPS sensor. This location information is converted into a digital format as latitude and longitude.

[0736] Step 2:

[0737] The user's device sends the acquired location information to the server. The data is encrypted using a secure protocol before transmission.

[0738] Step 3:

[0739] The server uses the received location information to search its database for relevant geographical information. During this process, tourist attractions and historical information are extracted.

[0740] Step 4:

[0741] The microphone and camera on the user's device are used, and an emotion engine analyzes the user's emotional state. This analysis includes voice tone and facial recognition.

[0742] Step 5:

[0743] The server receives emotion data from the emotion engine and generates or adjusts text content according to the user's emotions. This content will be tailored to the user's interests and emotions.

[0744] Step 6:

[0745] The server uses a text-to-speech engine to convert the edited text content into audio data. This audio data is processed in a format that is comfortable for the user to listen to.

[0746] Step 7:

[0747] The server sends the generated audio data to the user's device. The transmitted data is delivered quickly and securely.

[0748] Step 8:

[0749] The user's device receives audio data and plays it back through the built-in speaker or connected earphones. The user can hear the guidance information directly.

[0750] Step 9:

[0751] Users provide feedback on the information they receive. This feedback includes evaluations of the information's content and its relevance to their emotional state.

[0752] Step 10:

[0753] The server analyzes user feedback and optimizes the entire system to improve the personalization accuracy of future information delivery. This process continuously improves the learning database.

[0754] (Example 2)

[0755] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0756] Conventional geographic information systems often provide certain information based on the user's location, but they do not take into account the user's individual emotional state. As a result, there is a lack of personalization of the user experience, and the content provided may not match the user's current emotions.

[0757] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0758] In this invention, the server includes means for acquiring the user's location information, means for retrieving relevant geographical information from a data storage device, and means for analyzing the user's emotional state using an emotion engine. This enables the provision of appropriate and personalized information according to the user's current emotional state.

[0759] "User location information" refers to coordinate data indicating the user's current location, and is typically location information obtained using the Global Positioning Satellite System (GPS).

[0760] "Geographical information" refers to various types of information related to a specific location, including data about the region such as tourist attractions, history, culture, and facilities.

[0761] A "data storage device" is a storage device or database that stores a wide variety of information, such as geographical information, and allows it to be retrieved as needed.

[0762] A "text guide" refers to information presented to users in the form of explanatory or instructive text, which is generated based on geographical information.

[0763] An "emotion engine" is a program or system for analyzing a user's emotional state, using facial expressions, voice, sensor data, etc., to infer emotions.

[0764] "Audio data" refers to digital audio data that is created by converting text content into speech and is intended to be played back on a user's device.

[0765] A "speech synthesis engine" is a program or system for converting text content into speech data, and is a technology for generating artificial speech.

[0766] "Responses" refer to feedback and behavioral data collected from users, which can be used to improve system performance and personalize the system.

[0767] This invention relates to a system that dynamically provides personalized content by utilizing the user's location information and emotions. This system is mainly composed of a combination of a server, a terminal, and various software engines.

[0768] First, the user obtains location information using their device. The device utilizes its built-in GPS sensor to measure latitude and longitude data. This location information is transmitted to the server in real time.

[0769] The server retrieves relevant geographical information from the data storage device based on the received location information. A database management system is used in this process to efficiently obtain geographical information.

[0770] Next, the device uses its built-in camera, microphone, and sensors to analyze the user's emotions in real time. The emotion engine processes data from various sensors and applies a dedicated analysis algorithm to infer the emotional state. This analysis includes facial recognition technology and voice analysis technology.

[0771] Based on the output from the emotion engine, the server dynamically generates or adjusts content according to the user's emotions. Appropriately selected text information is converted into audio data by the speech synthesis engine. The speech synthesis engine uses a specified speech model to convert the text into smooth speech.

[0772] Audio data is sent from the server to the user's device, which then plays the audio using its built-in media player. This allows the user to receive personalized information in real time, based on their location and emotions.

[0773] As a concrete example, consider a scenario where a user is visiting a tourist spot in a city. If facial analysis determines that the user is relaxed, the server will prioritize providing information about quiet places and cafes in that location. Related historical information will also be provided via audio, allowing the user to efficiently obtain information while walking.

[0774] An example of a prompt to a generative AI model is the instruction, "Generate tourist information suitable for when the user is in an excited state." In response to this prompt, the system will prioritize generating information about adventurous activities.

[0775] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0776] Step 1:

[0777] The user's device obtains its current location using GPS functionality. The device inputs latitude and longitude data from the Global Positioning Satellite System (GSO) and transmits this information to the server. The obtained location information serves as the basis for searching for geographical information about the user's surroundings.

[0778] Step 2:

[0779] The server uses location information received from the terminal as input to query the data storage device and retrieve relevant geographical information. Using a database management system, it searches for tourist attractions, facilities, and other regional information, and prepares the retrieved geographical information for use in the next step.

[0780] Step 3:

[0781] The user's device uses its built-in camera, microphone, and sensors to collect data for input into the emotion engine. This data includes facial image data, voice tone, and physiological data. The emotion engine processes this data to determine the user's current emotional state, using image processing algorithms and voice analysis models.

[0782] Step 4:

[0783] The server receives emotional state data output from the emotion engine as input and dynamically generates personalized content. This process uses a generative AI model to create text and information that matches the user's emotions. The generated text content is adjusted according to the user's interests and emotions.

[0784] Step 5:

[0785] The server inputs the selected or edited text content into a speech synthesis engine and converts it into speech data. The synthesized speech is used to generate smooth speech data from the text. Here, a speech synthesis algorithm is applied to output speech in an easy-to-understand format.

[0786] Step 6:

[0787] The server sends the generated audio data to the user's device. The user's device uses its built-in media player to play the received audio data. This allows the user to hear appropriate information in real time, tailored to their location and emotions.

[0788] (Application Example 2)

[0789] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0790] Conventional geographic information systems provide information based solely on the user's location, making it difficult to deliver content tailored to the user's psychological state, such as emotions, stress levels, and curiosity, in real time. Furthermore, a lack of continuous content optimization based on user feedback and individual needs for tourism information was a significant challenge.

[0791] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0792] In this invention, the server includes means for acquiring the user's location information, means for analyzing the acquired user's emotional state, and means for generating personalized audio content based on the emotional state and location information. This makes it possible to provide personalized geographical information and audio guidance that corresponds to the user's psychological state.

[0793] "Means for acquiring user location information" refers to a device or method that has the function of determining the user's current geographical location and acquiring it as digital information.

[0794] "Means for retrieving geographical information from a data set" refers to a device or method that has the function of retrieving and extracting relevant geographical information from a database or the like based on acquired location data.

[0795] "Means for analyzing a user's emotional state" refers to a device or method that has the function of determining a user's psychological emotional state using the user's facial expressions, voice, and other physiological data.

[0796] "Means for generating personalized audio content" refers to a device or method that has the function of creating audio information tailored to individual needs based on the user's emotional state and location information.

[0797] The "feature for collecting user feedback" is a function that provides a mechanism for users to proactively submit their own opinions and suggestions for improvement regarding the information provided.

[0798] This invention is a system for providing personalized geographical information to users. The system mainly consists of a server and a user terminal.

[0799] First, the user's device acquires location information using a GPS module. This location information is sent to a server, which then searches its database for relevant geographical information based on this data. During this process, the server analyzes facial expression and voice data sent from the device to obtain the user's emotional state. Here, a facial expression recognition library (e.g., OpenCV) and voice tone recognition software are used.

[0800] The server uses acquired emotional state and location information to generate personalized audio content. This audio content is converted from text to audio data using a text-to-speech conversion device (e.g., Amazon Polly). The generated audio content is sent to the user's device and played back on the device. This allows the user to receive information in audio format that is tailored to their emotional state in real time.

[0801] Furthermore, it includes a function to collect feedback from users, and the system continuously optimizes content quality based on user feedback.

[0802] For example, if the system determines that a user has arrived at a tourist destination and is relaxing, it can provide information about nearby quiet parks or cafes and offer audio guidance about the historical background.

[0803] An example of a prompt for a generative AI model is: "Analyze the user's emotional state in real time based on their facial expressions and voice tone, and provide location-related tourist information in an engaging narration."

[0804] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0805] Step 1:

[0806] The device uses a GPS module to obtain the user's current location. The obtained location information is sent to the server. The input is the location information, and the output is the location information sent to the server.

[0807] Step 2:

[0808] The server searches the database for relevant geographical information based on the received location information. The input is the transmitted location information, and the output is the retrieved geographical information. The server executes queries and extracts the geographical information.

[0809] Step 3:

[0810] The device uses its camera and microphone to collect the user's facial expressions and voice, and generates emotion data. The input is the user's facial expressions and voice data, and the output is the generated emotion data. The device performs facial recognition and voice analysis.

[0811] Step 4:

[0812] The server analyzes emotional data to identify the user's emotional state. The input is emotional data, and the output is the identified emotional state. The server utilizes a facial expression recognition library and voice tone analysis software.

[0813] Step 5:

[0814] The server generates personalized text content based on identified emotional states and retrieved geographical information. The input is emotional state and geographical information, and the output is text content. The server selects information that matches the user's preferences.

[0815] Step 6:

[0816] The server converts the generated text content into audio data using a text-to-speech conversion processing unit. The input is text content, and the output is audio data. The server performs the conversion using a text-to-speech engine.

[0817] Step 7:

[0818] The server sends the generated audio data to the terminal. The input is the audio data, and the output is the audio data sent to the terminal. The server transfers the data.

[0819] Step 8:

[0820] The device plays the received audio data. The input is the received audio data, and the output is the audio guide provided to the user. The device activates the audio playback function.

[0821] Step 9:

[0822] Users provide feedback on the information they receive. The input is the information provided, and the output is the user's feedback. Users send their feedback via text or voice.

[0823] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0824] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0825] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0826] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0827] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0828] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0829] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0830] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0831] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0832] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0833] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0834] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0835] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0836] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0837] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0838] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0839] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0840] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0841] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0842] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0843] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0844] The following is further disclosed regarding the embodiments described above.

[0845] (Claim 1)

[0846] A means of obtaining the user's location information,

[0847] A means for retrieving relevant geographical information from a database based on the aforementioned location information,

[0848] A means for generating audio content based on the aforementioned searched geographical information,

[0849] Means for transmitting the aforementioned audio content to a user terminal,

[0850] means for playing the aforementioned audio content on the user terminal,

[0851] A system that includes this.

[0852] (Claim 2)

[0853] The system according to claim 1, comprising a text-to-speech engine for converting text content into audio data.

[0854] (Claim 3)

[0855] The system according to claim 1, further comprising a function to collect user feedback and improve the accuracy of future information provision.

[0856] "Example 1"

[0857] (Claim 1)

[0858] An information device that acquires location information,

[0859] An information processing device that retrieves relevant geographic information from an information storage device based on location information,

[0860] An information conversion device that generates audio information based on retrieved geographic information,

[0861] A communication device that transmits voice information to the user's information terminal,

[0862] An information terminal equipped with a playback device for playing audio information,

[0863] An information gathering device that collects user feedback and uses it to improve future information provision,

[0864] A system that includes this.

[0865] (Claim 2)

[0866] The system according to claim 1, comprising a speech conversion device for generating speech information.

[0867] (Claim 3)

[0868] The system according to claim 1, comprising interactive educational elements based on the content of audio information.

[0869] "Application Example 1"

[0870] (Claim 1)

[0871] A means of recognizing the user's geographical location signal,

[0872] A means for retrieving relevant regional information from a data storage area based on the aforementioned geographic location signal,

[0873] A means for generating audio output based on the aforementioned searched regional information,

[0874] Means for transmitting the aforementioned audio output to the user terminal,

[0875] means for playing the aforementioned audio output on the user terminal,

[0876] A means to be mounted on a head-mounted display device for providing information about the surrounding area during the operation of an autonomous mobile vehicle,

[0877] A system that includes this.

[0878] (Claim 2)

[0879] The system according to claim 1, comprising a string-to-speech engine for converting text information into speech signals.

[0880] (Claim 3)

[0881] The system according to claim 1, further comprising a function to collect feedback from users and improve the accuracy of future information provision.

[0882] "Example 2 of combining an emotion engine"

[0883] (Claim 1)

[0884] A means of obtaining the user's location information,

[0885] A means for retrieving relevant geographical information from a data storage device based on the aforementioned location information,

[0886] A means for generating a text guide based on the aforementioned geographical information,

[0887] A means of using an emotion engine to analyze the emotional state of users,

[0888] Means for selecting or adjusting text content according to the aforementioned emotional state,

[0889] Means for converting the selected or adjusted text content into audio data,

[0890] Means for transmitting the aforementioned audio data to the user terminal,

[0891] means for playing the aforementioned audio data on the user terminal,

[0892] A system that includes this.

[0893] (Claim 2)

[0894] The system according to claim 1, comprising a speech synthesis engine for converting text content into speech data.

[0895] (Claim 3)

[0896] The system according to claim 1, further comprising a function to collect user feedback and improve the accuracy of future information provision.

[0897] "Application example 2 when combining with an emotional engine"

[0898] (Claim 1)

[0899] A means of obtaining the user's location information,

[0900] A means for retrieving relevant geographical information from a data set based on the aforementioned location information,

[0901] A means of analyzing the emotional state of acquired users,

[0902] A means for generating personalized audio content based on the aforementioned emotional state and retrieved geographical information,

[0903] Means for transmitting the aforementioned audio content to a user terminal,

[0904] means for playing the aforementioned audio content on the user terminal,

[0905] A system that includes this.

[0906] (Claim 2)

[0907] The system according to claim 1, comprising a character information / speech conversion processing device for converting text content into speech data.

[0908] (Claim 3)

[0909] The system according to claim 1, which includes a function to collect user feedback and improve the accuracy of future information provision. [Explanation of symbols]

[0910] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of obtaining the user's location information, A means for retrieving relevant geographical information from a database based on the aforementioned location information, A means for generating audio content based on the aforementioned searched geographical information, Means for transmitting the aforementioned audio content to a user terminal, means for playing the aforementioned audio content on the user terminal, A system that includes this.

2. The system according to claim 1, comprising a text-to-speech engine for converting text content into audio data.

3. The system according to claim 1, further comprising a function to collect user feedback and improve the accuracy of future information provision.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A