system
The system addresses the limitations of conventional tourism guidance by using high-precision location and emotion analysis with generative AI to deliver personalized and real-time information, improving the visitor experience.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
Conventional tourism guidance methods are time-consuming, lack accuracy in location information, and struggle to provide personalized information to a large number of visitors, making it difficult to offer real-time and individualized guidance.
A system utilizing high-precision location information acquisition, voice guidance data, and generative AI to provide location-specific and personalized information in response to user requests, with real-time tracking and emotion analysis for tailored guidance.
Enriches the visitor experience by providing accurate, real-time, and personalized information based on user location and emotional state, enhancing the understanding and satisfaction of tourists and facility visitors.
Smart Images

Figure 2026073412000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the conventional tourism guidance method, visitors had to read physical explanation panels within a wide tourist area, which was time-consuming and laborious. Also, since the accuracy of location information in general tourism applications was not sufficient, it was difficult to immediately provide detailed information about specific attractions. Furthermore, there was also a problem that it was difficult to provide personalized information to a large number of visitors. There is a demand for providing a method that solves such problems and easily provides real-time and individualized tourism guidance to visitors.
Means for Solving the Problems
[0005] The present invention is a system comprising means for acquiring highly accurate location information, means for providing voice guidance data related to a specific location based on said location information, means for generating additional information in response to user requests using a generation AI, and means for outputting said additional information as voice data. With this system, when a visitor reaches a specific location, appropriate voice guidance is automatically provided based on highly accurate location information, and they can request additional information according to their interests. Furthermore, by tracking the visitor's movements in real time and personalizing the information, the problems of effort and accuracy that have been challenges in the past can be effectively resolved.
[0006] "High-precision location information acquisition means" refers to devices and systems that use technology to accurately determine a user's location with an error margin of a few centimeters.
[0007] "Voice guidance data" refers to data that expresses information related to a specific place or event in audio form, providing information to the user through hearing.
[0008] "Generative AI" is a system that uses artificial intelligence technology to analyze data and generate new, appropriate information in response to user requests and needs.
[0009] "Additional information" refers to further details and related information that, in addition to basic information, are tailored to the user's interests and questions.
[0010] "Audio data" refers to audio signals stored or transmitted in digital format, which are the result of converting textual information into auditory information. [Brief explanation of the drawing]
[0011] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3]This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0012] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0013] First, let's explain the terminology used in the following explanation.
[0014] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0015] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0016] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0017] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0018] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0019] [First Embodiment]
[0020] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0021] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0022] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0023] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0024] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0025] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0026] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0027] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0028] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0029] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0030] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0031] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0032] This invention is a system that uses a high-precision location information acquisition method to provide users with voice guidance on information about specific tourist destinations. The following describes the system's program-based processing in natural language.
[0033] Server Role
[0034] The server retrieves relevant information from a database based on location data and uses generative AI to generate voice guidance tailored to the user. Specifically, the server receives the user's location information transmitted from the device and retrieves data on tourist spots related to that location from the database. This data includes historical background, cultural significance, and recommended points of interest. Furthermore, the server also utilizes generative AI to generate additional information in response to specific user requests. The generated information is converted into voice data and sent to the device.
[0035] Terminal role
[0036] The device acquires the user's location information in real time and sends it to the server. By playing the received audio data back to the user, it provides tourist information. The device also has the function to accept user input and send additional requests to the server. This allows for flexible responses when the user asks additional questions or shows new interests.
[0037] User experience
[0038] While visiting tourist attractions, users can deepen their understanding of the sites by receiving location-specific audio guidance. By carrying a smartphone, users can listen to explanations based on their current location, obtaining diverse information in real time without using their hands. For example, if a user is in front of a historical building, they can hear about the building's history, architect, and related anecdotes. Furthermore, if the user wants to know more details about the building, they can request additional information by operating their device. In response to this request, the server generates more detailed information and related stories, delivering them to the user as audio data.
[0039] Thus, the system of the present invention makes it possible to enrich the visitor experience at tourist destinations and efficiently provide personalized information.
[0040] The following describes the processing flow.
[0041] Step 1:
[0042] The device acquires location information. The device uses its built-in GPS function and high-precision location measurement device to determine the user's current location. This location information is updated at regular intervals and immediately transmitted to the server.
[0043] Step 2:
[0044] The server receives location information and retrieves corresponding tourist destination data. The server accesses the database and searches for information on the relevant tourist spot based on the received location information. The retrieved information includes historical background, points of interest, and related anecdotes.
[0045] Step 3:
[0046] The server uses a generation AI to create voice guidance data. Based on the user's location and specific request, the server uses the generation AI to generate appropriate voice guidance. The generated text data is then converted into voice data, preparing it for delivery to the user.
[0047] Step 4:
[0048] The server sends audio data to the terminal. The server sends the generated audio data to the terminal via the network and prepares it for voice guidance to the user.
[0049] Step 5:
[0050] The device plays an audio guide. The device plays the received audio data to the user, providing information about a specific tourist spot. Users can deepen their understanding of the tourist spot by listening to the audio guide through their smartphone.
[0051] Step 6:
[0052] Users can request additional information. After listening to the voice guidance, users can send requests via their device if they want more detailed information or further questions.
[0053] Step 7:
[0054] The terminal sends a request to the server. The terminal sends a request for additional user information and questions to the server in text format, asking for the necessary information to be provided.
[0055] Step 8:
[0056] The server generates and sends information based on the request. The server analyzes the user's request and retrieves new relevant information from the database or generates it using a generative AI. The newly generated audio data is sent to the terminal and provided to the user.
[0057] Step 9:
[0058] The device plays additional information. The device plays additional audio data received from the server, providing the user with further sightseeing information. By repeating this process, the user can obtain a detailed and personalized sightseeing experience.
[0059] (Example 1)
[0060] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0061] Current tourist information systems have limitations in providing real-time information and generating additional information based on the user's location, and they struggle to effectively deliver personalized experiences. In particular, their inability to provide dynamic information tailored to the user's interests and preferences is a major problem.
[0062] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0063] In this invention, the server includes a high-precision location information acquisition device, a device that provides voice guidance information related to a specific geographical area based on the location information, and a device that generates additional information in response to user requests using a generation AI model. This enables the real-time provision of personalized guidance information based on the user's current location information.
[0064] A "high-precision location information acquisition device" is a device that can accurately and precisely identify a user's geographical location and provide that information in real time.
[0065] A "geographical area" is a concept that refers to the physical extent of a place, including specific places of interest or tourist attractions.
[0066] "Audio guidance information" refers to data generated to provide users with information such as tourist attractions, geographical features, and historical background in audio format.
[0067] A "generative AI model" is an artificial intelligence program that uses natural language processing and machine learning to automatically generate information in response to user requests.
[0068] A "prompt statement" is an input statement that contains instructions or requests necessary for a generative AI model to provide meaningful information.
[0069] "Personalized guidance information" refers to information that is tailored to specific needs and interests based on the user's current location, past behavior, and preferences.
[0070] This invention is a system that uses a high-precision location information acquisition device to provide users with voice guidance on information about a specific geographical area. This makes it possible to provide users with real-time information about tourist spots they are visiting.
[0071] The server retrieves relevant geographical information from a database based on the user's location information transmitted from the terminal. At this time, the server organizes the information to provide clear and easy-to-understand guidance based on highly accurate location data. Furthermore, the server uses a generative AI model to generate additional information in response to the user's request. The generated information is then converted into audio data using speech synthesis technology and transmitted to the terminal.
[0072] The device transmits real-time location information obtained via its built-in GPS function to a server, and plays back audio data received from the server to the user. This allows the user to obtain information in a natural way.
[0073] Users can use their smartphones or tablets to listen to audio guides played from their devices, deepening their understanding of the background and highlights of tourist destinations. For example, when a user visits a particular historical building, they can receive detailed audio information about its history and background. Furthermore, if a user has specific interests or questions, they can make additional requests through their device, which can generate even more detailed information.
[0074] An example of a prompt message would be a request such as, "Please tell me about the founder of this historic building." In response, the server can use a generative AI model to retrieve appropriate information and provide it to the user as voice guidance.
[0075] In this way, the system can efficiently provide personalized tourist information to users.
[0076] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0077] Step 1:
[0078] The device uses GPS functionality to obtain the user's current location information. This location information is represented as latitude and longitude coordinates and sent to the server. The input is the user's real-time location information, and the output is the transmission of location data to the server.
[0079] Step 2:
[0080] The server retrieves information about the corresponding geographical area from its database based on the location information received from the device. This process uses the location information as a search query to extract details about relevant tourist attractions. The input is the location information from the device, and the output is the retrieved information about tourist attractions.
[0081] Step 3:
[0082] The server uses a generative AI model to construct the content of the voice guidance based on the acquired tourist spot information. It creates prompt sentences and inputs them into the generative AI model to generate appropriate voice guidance content. The input is tourist spot information and user prompt sentences, and the output is the generated voice guidance content.
[0083] Step 4:
[0084] The server uses a speech synthesis API to convert the generated voice guidance content into audio data. This process converts text data into an audio data format. The input is the voice guidance content generated by the AI model, and the output is playable audio data.
[0085] Step 5:
[0086] The server sends the converted audio data to the terminal. This communication is conducted securely over the internet. The input is audio data, and the output is the transmission of data to the terminal.
[0087] Step 6:
[0088] The terminal plays audio data received from the server to the user. Audio guidance is provided through the terminal's speaker or earphones, allowing the user to hear information about tourist destinations. The input is audio data from the server, and the output is the playback of audio guidance to the user.
[0089] Step 7:
[0090] If the user needs further information, they can operate the terminal to request additional information. This request is sent from the terminal to the server. Input is a new question or request from the user, and output is the request sent to the server.
[0091] Step 8:
[0092] The server receives the user's additional request and generates new information again using the generation AI model. This new information is then converted back into audio data and sent to the terminal. The input is the user's request and prompt text, and the output is the audio data of the additional information.
[0093] Step 9:
[0094] The device plays newly received audio data to provide the user with further instructions. The input is audio data containing additional information, and the output is the playback of the audio guidance to the user.
[0095] (Application Example 1)
[0096] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0097] In modern commercial facilities, the amount of information provided to visitors is limited, making it difficult to provide efficient guidance tailored to individual preferences. As a result, visitors are unable to effectively obtain information about products and stores, hindering their satisfaction. Furthermore, the lack of real-time location-based guidance prevents the provision of a dynamic shopping experience.
[0098] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0099] In this invention, the server includes means for acquiring high-precision location information, means for providing voice guidance data related to a specific location based on the location information, means for generating additional information in response to user requests using a generating AI, and means for providing user-specific guidance based on information within the commercial facility. This enables the provision of real-time, dynamic, and personalized information to individual visitors.
[0100] "High-precision location information acquisition method" refers to a technology that allows a user's precise current location to be determined in real time through a digital device.
[0101] "Voice guidance data" refers to digital data used to provide users with useful information in audio format.
[0102] "Generative AI" is an artificial intelligence technology that creates and outputs appropriate information based on user requests.
[0103] "Information within a commercial facility" refers to detailed information about products, stores, campaigns, etc., related to the commercial facility.
[0104] "Personalized guidance for users" refers to a service that provides customized information based on the individual user's preferences and behavioral history.
[0105] This system provides real-time location-based voice guidance services to visitors within commercial facilities. The implementation of this invention requires high-precision location information acquisition means, voice guidance data provision means, and information generation means using AI. Specific hardware includes portable digital devices such as smartphones and tablets, and built-in GPS modules. The software used includes a real-time location tracking API, an information provision server, a speech synthesis engine, and a generative AI model.
[0106] The server receives highly accurate location information from the user's device and uses that data to collect relevant information within the commercial facility. This information includes store, product, campaign, and event information. The generating AI considers the user's current location and request to create personalized guidance. The generated guidance is then converted into voice data using a speech synthesis engine.
[0107] The device has the function of playing guidance to the user via audio data provided by the server. The audio guidance can be customized based on the user's preferences and past conversation history. As a result, users can receive information hands-free, enriching their shopping experience.
[0108] For example, let's say a user is standing in front of a clothing store in a shopping mall. The user's smartphone sends its location information to a server, which then provides relevant promotions and information on popular products. The generated voice message might say, "This store is currently having a sale on its latest collection. Please request additional information if you need more details." An example of a prompt might be, "Please provide information on stores around latitude=35.6895, longitude=139.6917."
[0109] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0110] Step 1:
[0111] The device uses its built-in GPS module to obtain the user's current location in real time. The input is the device's GPS data, and the output is numerical information of latitude and longitude. This numerical information is sent to the server via the location API on the device.
[0112] Step 2:
[0113] The server receives location information transmitted from the terminal. Based on this, it retrieves relevant information within the commercial facility from the database. The input is latitude and longitude information sent from the terminal, and the output is detailed information about the corresponding store or product. This data is searched and retrieved using database queries.
[0114] Step 3:
[0115] The server uses the acquired information to input prompts into the generating AI model, which then generates user-optimized guidance information. The input is information about the commercial facility obtained from the database, and the output is user-specific guidance content. Based on these inputs, the generating AI performs natural language processing to construct guidance text appropriate to the context.
[0116] Step 4:
[0117] The server inputs the generated guidance information into a speech synthesis engine and converts it into audio data. The input is the generated guidance content (in text format), and the output is playable audio data. Speech synthesis is performed using text-to-speech technology to generate audio that is easy for the user to understand.
[0118] Step 5:
[0119] The terminal plays audio data received from the server and provides guidance to the user. The input is audio data transmitted from the server, and the output is the user's auditory reception of the information. The terminal outputs the audio through a speaker or earphones via its playback function.
[0120] Step 6:
[0121] If the user needs further information after hearing the instructions, they can send additional requests to the server via the terminal. Input is the user's voice or text request, and output is an information request corresponding to the request's instructions. The terminal collects new requests through the user interface and sends them to the server.
[0122] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0123] This invention provides a system that combines a high-precision location information acquisition method, an emotion engine, and a generative AI to provide voice guidance that takes into account the user's emotional state. This system collects information about tourist spots based on the user's location information, recognizes the user's emotions, and then generates personalized voice guidance accordingly.
[0124] Server Role
[0125] The server first receives the user's location information transmitted from the terminal. It analyzes the location information and retrieves data on relevant tourist spots from a database. Furthermore, the emotion engine receives the user's emotional state and provides this to the generating AI. The generating AI uses this information to generate voice guidance with appropriate tone and content based on the user's emotions. This voice guidance is adjusted to a calm tone if the user wants to relax, and to include details if the AI determines the user is interested. The generated voice data is then transmitted to the terminal and provided to the user.
[0126] Terminal role
[0127] The device acquires the user's location information with high accuracy and transmits it to the server in real time. Through an emotion engine, it estimates the user's emotions from their voice, facial expressions, and actions, and transmits this information to the server. The device also receives voice guidance data from the server and plays it back to the user. The content of the voice guidance is applied based on emotion recognition, serving to further personalize the user's experience.
[0128] User experience
[0129] Users can access the system via their smartphones to receive detailed, personalized audio guidance about tourist attractions. For example, if the emotion engine detects that a user is slightly tired while standing in front of a historical building, the guidance will be shorter and focus on key points. Conversely, if it detects high levels of excitement or interest, additional anecdotes and trivia will be provided. This allows users to effortlessly receive information tailored to their emotional state, making the sightseeing experience richer and more convenient.
[0130] Thus, the present invention integrates location information, emotion recognition, and AI technology to realize a system that provides a high level of satisfaction to visitors.
[0131] The following describes the processing flow.
[0132] Step 1:
[0133] The device acquires location information and emotion data. The device uses its built-in high-precision GPS function to determine the user's current location and activates an emotion engine that analyzes the user's emotions from their facial expressions and voice using the camera and microphone.
[0134] Step 2:
[0135] The device sends location information and emotional data acquired by the device to the server. The device packages the analyzed emotional state and location information and sends the data to the server in real time.
[0136] Step 3:
[0137] The server retrieves tourist spot information based on location data. The server accesses a database and extracts detailed information about the tourist spot corresponding to the transmitted location data.
[0138] Step 4:
[0139] The server adjusts the voice guidance content based on the user's emotions. The server provides the received emotion data to the generating AI, which then edits the tourist spot information guidance to match the user's mood. For excited users, it provides detailed and abundant information, while for users seeking relaxation, it provides concise and calming content.
[0140] Step 5:
[0141] The server generates voice guidance data and sends it to the terminal. The voice data, generated according to the user's emotional state, is sent to the terminal via the network, and the guidance is ready for the user.
[0142] Step 6:
[0143] The device plays voice guidance to the user. The device plays the received voice data and provides the user with customized guidance information. This process allows the user to receive the best possible guidance experience tailored to their emotions.
[0144] Step 7:
[0145] If needed, users can request further information or additional questions via the terminal. If a user wants to learn more about information they are interested in, they can operate the terminal to send additional requests to the server.
[0146] Step 8:
[0147] The device sends the user's additional request to the server. The device organizes the user's new request and sends it to the server along with sentiment data.
[0148] Step 9:
[0149] The server generates new information and sends additional voice guidance, adjusted according to the user's emotions, to the device. The server uses a generation AI to generate new relevant information and sends the readjusted voice data back to the device.
[0150] Step 10:
[0151] The device plays additional audio to provide users with deeper guidance. The device plays new audio guidance to further enrich the user experience. This ensures that users always receive information tailored to their interests and emotions.
[0152] (Example 2)
[0153] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0154] Conventional voice guidance systems struggled to provide personalized information that took into account the user's emotional state, and merely offered general directions based on location information. Furthermore, they lacked sufficient real-time information updates and optimization of voice guidance according to the user's emotional state, resulting in a limited user experience.
[0155] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0156] In this invention, the server includes a high-precision positioning means, a means for providing voice guidance information related to a specific area based on the positioning, and a means for generating additional information based on the user's emotional state using a generation AI model. This makes it possible to provide real-time optimized voice guidance that takes into account the user's emotional state.
[0157] "High-precision positioning means" refers to technology that acquires the user's location as latitude and longitude data with high accuracy.
[0158] "Means for providing audio guidance information related to a specific region" refers to technology that provides cultural, historical, and tourist information about a location in audio format based on acquired location information.
[0159] A "generative AI model" refers to the structure of artificial intelligence that uses machine learning algorithms to generate optimal additional information based on input data.
[0160] "Means for analyzing a user's emotional state" refers to technology that uses audio and video data acquired from microphones and cameras to identify the emotions a user is currently experiencing.
[0161] "Means for adjusting the tone and level of detail of voice guidance information" refers to technologies that change the atmosphere and depth of content of the voice information provided according to the user's emotional state.
[0162] "A means of tracking in real time and instantly updating location-specific guidance information" refers to technology that continuously monitors the user's movements and provides the latest, appropriate information whenever their location changes.
[0163] "Adaptive means" refers to technologies that personalize the information provided based on past data and user preferences.
[0164] This invention is a system that provides personalized voice guidance to users by combining multiple technological elements. Specifically, it utilizes high-precision positioning means, emotion analysis functions, generative AI models, and voice output functions.
[0165] Hardware and software to use
[0166] The device is equipped with a GPS module, microphone, and camera, which are used to acquire the user's location and emotional state. The location information is transmitted to the server in real time, and the audio and video data from the microphone and camera are processed by emotion analysis software. The emotion analysis software infers the user's emotional state based on their voice tone and facial expressions.
[0167] The server generates voice guidance using a generative AI model based on information sent from the terminal. The generative AI model receives regional information based on the user's location and emotional state as input, and generates prompt sentences. Based on this, it creates guidance with appropriate content and tone. This generated voice guidance data is then sent back to the terminal.
[0168] Specific example
[0169] For example, suppose a user is standing in front of a historical building. The device identifies its location via GPS and informs the server that it is a tourist spot. Simultaneously, the microphone captures the user's voice, and emotion analysis software recognizes that the user is slightly tired. Once this information is sent to the server, the server provides a prompt message to the AI model, such as "The user is in front of a historical building. Their emotional state is fatigued. Please create a concise guide," and sends the generated guide to the device.
[0170] In this way, users can obtain information that matches their emotions at the time, making their sightseeing experience more fulfilling. An example of a prompt might be: "The user is in front of Kinkaku-ji Temple. Their emotional state is excited. Please create an audio guide that includes detailed history and anecdotes."
[0171] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0172] Step 1:
[0173] The device acquires the user's current location using a high-precision GPS module. The GPS receiver analyzes satellite data to calculate latitude and longitude information. This location information is sent to the server as output data from the device.
[0174] Step 2:
[0175] The device captures the user's voice and facial expressions using a microphone and camera. The digital data obtained from the microphone and camera is input into emotion analysis software. The software analyzes the phoneme and visual features to identify the user's emotional state (e.g., excitement, fatigue). This emotional information is sent to the server as output data from the device.
[0176] Step 3:
[0177] The server searches its database for relevant tourist spot information based on location information received from the terminal. Once the location information is entered, data for tourist spots in the corresponding area is retrieved as server output. This includes the spot's name, description, and any special events.
[0178] Step 4:
[0179] The server inputs emotional information received from the terminal and acquired tourist spot information into a generating AI model. The generating AI model generates prompt sentences and creates personalized voice guidance based on them. The input data consists of emotional information and tourist spot information, and the output is voice guidance data tailored to the user's emotions.
[0180] Step 5:
[0181] The server sends the generated audio guidance data to the terminal. The terminal receives this audio data and plays it for the user through its built-in speaker. The audio guidance is tailored to enhance the user experience. The terminal's final output is for the user to listen to the audio guidance and gain a deeper understanding of the tourist destination.
[0182] (Application Example 2)
[0183] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0184] In recent years, many physical facilities have been required to improve the visitor experience. However, providing personalized guidance information that takes into account visitors' emotions and preferences in real time is technically complex, and achieving satisfactory results is often difficult. Therefore, there is a need to develop a system that provides a personalized experience based on the visitor's location and emotions.
[0185] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0186] In this invention, the server includes a high-precision location information acquisition device, a device that provides voice guidance information related to a specific facility based on the location information, a device that generates additional information customized based on customer preferences and past interactions using generative AI technology, a device that includes a process of analyzing the customer's emotional state using an emotion engine and personalizing the additional information, and a device that outputs the personalized information as voice data. As a result, visitors can receive optimal guidance information that matches their emotions and preferences at that moment.
[0187] A "high-precision location information acquisition device" is a device that accurately grasps the user's current location in real time and provides necessary information.
[0188] "Audio guidance information" refers to content that presents useful information about the facility or place the user is visiting in audio format.
[0189] "Generative AI technology" is an artificial intelligence technology that generates personalized additional information based on the user's past behavior and preferences.
[0190] An "emotion engine" is software or hardware that analyzes data such as a user's voice and facial expressions to estimate their emotional state.
[0191] "Personalized information" refers to guidance information and content that is optimized based on the individual user's preferences and emotional state.
[0192] "Audio data is digital data used to convey information generated by a device to the user as sound."
[0193] This invention provides a personalized guide system to enhance the customer experience in tourist destinations and commercial facilities. This system acquires customer location information with high accuracy and analyzes their emotional state to provide more personalized guidance. The main functions of the system and its embodiments are described below.
[0194] The server uses a high-precision location acquisition device (e.g., a GPS module) to determine the user's current location. Next, it retrieves data on relevant facilities and tourist spots from a database based on the location information and provides it as voice guidance information.
[0195] The server also uses an emotion engine (e.g., Affectiva's SDK) to estimate the user's emotional state in real time from their voice and facial expression data. This emotional state is then input into generative AI technology (e.g., OpenAI's GPT-3®). The generative AI considers the user's preferences and past interactions to generate personalized additional information. This generated information is output as personalized voice data.
[0196] For example, if a user smiles when visiting a specific facility in a tourist area, the emotion engine will determine that the user is feeling excited. This information is then provided to the generative AI model, which, using a prompt such as, "Please explain in detail the local history and manufacturing process of the local specialty that the tourist is interested in," provides detailed location-related information via voice.
[0197] In this way, the present invention can significantly improve the experience at commercial facilities and tourist destinations by utilizing location information and sentiment analysis to provide each customer with the most suitable information in real time.
[0198] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0199] Step 1:
[0200] The terminal acquires the user's location information using a high-precision location information acquisition device and transmits that data to the server. The input is the location information acquired by the terminal, and the output is the transmission of location information data to the server. Through this process, the server can recognize the user's current location and acquire appropriate facility data.
[0201] Step 2:
[0202] The server analyzes the received location information and retrieves information about facilities and tourist attractions associated with that location from a database. The input is the user's location information, and the output is data about the relevant facilities. Based on the location information, the server searches for relevant information and prepares data according to the user's actions.
[0203] Step 3:
[0204] The device acquires the user's voice and facial expression data, estimates the emotional state using an emotion engine, and sends that data to the server. The input is the user's voice and facial expression data, and the output is the estimated emotional state. Here, the device collects data using voice and video sensors.
[0205] Step 4:
[0206] The server inputs emotional state and facility data into a generating AI model, which then uses prompts to generate personalized additional information. The input is the user's emotional state and facility data, and the output is customized guidance information. The generating AI model uses this information to create appropriate guidance content.
[0207] Step 5:
[0208] The server outputs the generated additional information as audio data and sends it to the terminal. The input is personalized information from the generating AI, and the output is audio guidance data played on the terminal. In this process, the server uses speech synthesis technology to convert the digital data into audio.
[0209] Step 6:
[0210] The user receives audio data through their device and listens to personalized guidance in real time. The input is audio data from the server, and the output is the audio information received by the user. This allows the user to receive optimal guidance tailored to their emotions and interests.
[0211] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0212] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0213] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0214] [Second Embodiment]
[0215] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0216] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0217] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0218] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0219] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0220] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0221] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0222] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0223] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0224] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0225] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0226] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0227] This invention is a system that uses a high-precision location information acquisition method to provide users with voice guidance on information about specific tourist destinations. The following describes the system's program-based processing in natural language.
[0228] Server Role
[0229] The server retrieves relevant information from a database based on location data and uses generative AI to generate voice guidance tailored to the user. Specifically, the server receives the user's location information transmitted from the device and retrieves data on tourist spots related to that location from the database. This data includes historical background, cultural significance, and recommended points of interest. Furthermore, the server also utilizes generative AI to generate additional information in response to specific user requests. The generated information is converted into voice data and sent to the device.
[0230] Terminal role
[0231] The device acquires the user's location information in real time and sends it to the server. By playing the received audio data back to the user, it provides tourist information. The device also has the function to accept user input and send additional requests to the server. This allows for flexible responses when the user asks additional questions or shows new interests.
[0232] User experience
[0233] While visiting tourist attractions, users can deepen their understanding of the sites by receiving location-specific audio guidance. By carrying a smartphone, users can listen to explanations based on their current location, obtaining diverse information in real time without using their hands. For example, if a user is in front of a historical building, they can hear about the building's history, architect, and related anecdotes. Furthermore, if the user wants to know more details about the building, they can request additional information by operating their device. In response to this request, the server generates more detailed information and related stories, delivering them to the user as audio data.
[0234] Thus, the system of the present invention makes it possible to enrich the visitor experience at tourist destinations and efficiently provide personalized information.
[0235] The following describes the processing flow.
[0236] Step 1:
[0237] The device acquires location information. The device uses its built-in GPS function and high-precision location measurement device to determine the user's current location. This location information is updated at regular intervals and immediately transmitted to the server.
[0238] Step 2:
[0239] The server receives location information and retrieves corresponding tourist destination data. The server accesses the database and searches for information on the relevant tourist spot based on the received location information. The retrieved information includes historical background, points of interest, and related anecdotes.
[0240] Step 3:
[0241] The server uses a generation AI to create voice guidance data. Based on the user's location and specific request, the server uses the generation AI to generate appropriate voice guidance. The generated text data is then converted into voice data, preparing it for delivery to the user.
[0242] Step 4:
[0243] The server sends audio data to the terminal. The server sends the generated audio data to the terminal via the network and prepares it for voice guidance to the user.
[0244] Step 5:
[0245] The device plays an audio guide. The device plays the received audio data to the user, providing information about a specific tourist spot. Users can deepen their understanding of the tourist spot by listening to the audio guide through their smartphone.
[0246] Step 6:
[0247] Users can request additional information. After listening to the voice guidance, users can send requests via their device if they want more detailed information or further questions.
[0248] Step 7:
[0249] The terminal sends a request to the server. The terminal sends a request for additional user information and questions to the server in text format, asking for the necessary information to be provided.
[0250] Step 8:
[0251] The server generates and sends information based on the request. The server analyzes the user's request and retrieves new relevant information from the database or generates it using a generative AI. The newly generated audio data is sent to the terminal and provided to the user.
[0252] Step 9:
[0253] The device plays additional information. The device plays additional audio data received from the server, providing the user with further sightseeing information. By repeating this process, the user can obtain a detailed and personalized sightseeing experience.
[0254] (Example 1)
[0255] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0256] Current tourist information systems have limitations in providing real-time information and generating additional information based on the user's location, and they struggle to effectively deliver personalized experiences. In particular, their inability to provide dynamic information tailored to the user's interests and preferences is a major problem.
[0257] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0258] In this invention, the server includes a high-precision location information acquisition device, a device that provides voice guidance information related to a specific geographical area based on the location information, and a device that generates additional information in response to user requests using a generation AI model. This enables the real-time provision of personalized guidance information based on the user's current location information.
[0259] A "high-precision location information acquisition device" is a device that can accurately and precisely identify a user's geographical location and provide that information in real time.
[0260] A "geographical area" is a concept that refers to the physical extent of a place, including specific places of interest or tourist attractions.
[0261] "Audio guidance information" refers to data generated to provide users with information such as tourist attractions, geographical features, and historical background in audio format.
[0262] A "generative AI model" is an artificial intelligence program that uses natural language processing and machine learning to automatically generate information in response to user requests.
[0263] A "prompt statement" is an input statement that contains instructions or requests necessary for a generative AI model to provide meaningful information.
[0264] "Personalized guidance information" refers to information that is tailored to specific needs and interests based on the user's current location, past behavior, and preferences.
[0265] This invention is a system that uses a high-precision location information acquisition device to provide users with voice guidance on information about a specific geographical area. This makes it possible to provide users with real-time information about tourist spots they are visiting.
[0266] The server retrieves relevant geographical information from a database based on the user's location information transmitted from the terminal. At this time, the server organizes the information to provide clear and easy-to-understand guidance based on highly accurate location data. Furthermore, the server uses a generative AI model to generate additional information in response to the user's request. The generated information is then converted into audio data using speech synthesis technology and transmitted to the terminal.
[0267] The device transmits real-time location information obtained via its built-in GPS function to a server, and plays back audio data received from the server to the user. This allows the user to obtain information in a natural way.
[0268] Users can use their smartphones or tablets to listen to audio guides played from their devices, deepening their understanding of the background and highlights of tourist destinations. For example, when a user visits a particular historical building, they can receive detailed audio information about its history and background. Furthermore, if a user has specific interests or questions, they can make additional requests through their device, which can generate even more detailed information.
[0269] An example of a prompt message would be a request such as, "Please tell me about the founder of this historic building." In response, the server can use a generative AI model to retrieve appropriate information and provide it to the user as voice guidance.
[0270] In this way, the system can efficiently provide personalized tourist information to users.
[0271] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0272] Step 1:
[0273] The device uses GPS functionality to obtain the user's current location information. This location information is represented as latitude and longitude coordinates and sent to the server. The input is the user's real-time location information, and the output is the transmission of location data to the server.
[0274] Step 2:
[0275] The server retrieves information about the corresponding geographical area from its database based on the location information received from the device. This process uses the location information as a search query to extract details about relevant tourist attractions. The input is the location information from the device, and the output is the retrieved information about tourist attractions.
[0276] Step 3:
[0277] The server uses a generative AI model to construct the content of the voice guidance based on the acquired tourist spot information. It creates prompt sentences and inputs them into the generative AI model to generate appropriate voice guidance content. The input is tourist spot information and user prompt sentences, and the output is the generated voice guidance content.
[0278] Step 4:
[0279] The server uses a speech synthesis API to convert the generated voice guidance content into audio data. This process converts text data into an audio data format. The input is the voice guidance content generated by the AI model, and the output is playable audio data.
[0280] Step 5:
[0281] The server sends the converted audio data to the terminal. This communication is conducted securely over the internet. The input is audio data, and the output is the transmission of data to the terminal.
[0282] Step 6:
[0283] The terminal plays the voice data received from the server for the user. Voice guidance is provided through the terminal's speaker or earphone, and the user can listen to information about the tourist destination. The input is the voice data from the server, and the output is the playback of the voice guidance to the user.
[0284] Step 7:
[0285] If the user needs further information, the user can operate the terminal to request additional information. This request is sent from the terminal to the server. The input is the user's new question or request, and the output is the transmission of the request to the server.
[0286] Step 8:
[0287] The server receives the user's additional request and uses the regenerative AI model to generate new information again. This new information is converted into voice data again and sent to the terminal. The input is the user's request and the prompt text, and the output is the voice data of the additional information.
[0288] Step 9:
[0289] The terminal plays the newly received voice data and provides further guidance to the user. The input is the voice data of the additional information, and the output is the playback of the voice guidance to the user.
[0290] (Application Example 1)
[0291] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0292] In modern commercial facilities, the amount of information provided to visitors is limited, making it difficult to provide efficient guidance tailored to individual preferences. As a result, visitors are unable to effectively obtain information about products and stores, hindering their satisfaction. Furthermore, the lack of real-time location-based guidance prevents the provision of a dynamic shopping experience.
[0293] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0294] In this invention, the server includes means for acquiring high-precision location information, means for providing voice guidance data related to a specific location based on the location information, means for generating additional information in response to user requests using a generating AI, and means for providing user-specific guidance based on information within the commercial facility. This enables the provision of real-time, dynamic, and personalized information to individual visitors.
[0295] "High-precision location information acquisition method" refers to a technology that allows a user's precise current location to be determined in real time through a digital device.
[0296] "Voice guidance data" refers to digital data used to provide users with useful information in audio format.
[0297] "Generative AI" is an artificial intelligence technology that creates and outputs appropriate information based on user requests.
[0298] "Information within a commercial facility" refers to detailed information about products, stores, campaigns, etc., related to the commercial facility.
[0299] "Personalized guidance for users" refers to a service that provides customized information based on the individual user's preferences and behavioral history.
[0300] This system provides real-time location-based voice guidance services to visitors within commercial facilities. The implementation of this invention requires high-precision location information acquisition means, voice guidance data provision means, and information generation means using AI. Specific hardware includes portable digital devices such as smartphones and tablets, and built-in GPS modules. The software used includes a real-time location tracking API, an information provision server, a speech synthesis engine, and a generative AI model.
[0301] The server receives highly accurate location information from the user's device and uses that data to collect relevant information within the commercial facility. This information includes store, product, campaign, and event information. The generating AI considers the user's current location and request to create personalized guidance. The generated guidance is then converted into voice data using a speech synthesis engine.
[0302] The device has the function of playing guidance to the user via audio data provided by the server. The audio guidance can be customized based on the user's preferences and past conversation history. As a result, users can receive information hands-free, enriching their shopping experience.
[0303] For example, let's say a user is standing in front of a clothing store in a shopping mall. The user's smartphone sends its location information to a server, which then provides relevant promotions and information on popular products. The generated voice message might say, "This store is currently having a sale on its latest collection. Please request additional information if you need more details." An example of a prompt might be, "Please provide information on stores around latitude=35.6895, longitude=139.6917."
[0304] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0305] Step 1:
[0306] The terminal uses the built-in GPS module to obtain the user's current location in real time. The input is the GPS data of the device, and the output is the numerical information of latitude and longitude. This numerical information is sent to the server through the location information API on the device.
[0307] Step 2:
[0308] The server receives the location information sent from the terminal. Based on this, it obtains the relevant information within the commercial facility from the database. The input is the latitude and longitude information sent from the terminal, and the output is the detailed information of the corresponding store and products. These data are retrieved through database queries.
[0309] Step 3:
[0310] The server inputs the prompt text into the generated AI model using the obtained information to generate the optimized guidance information for the user. The input is the information within the commercial facility obtained from the database, and the output is the guidance content specialized for the user. The generative AI performs natural language processing based on these inputs to construct a guidance text suitable for the context.
[0311] Step 4:
[0312] The server inputs the generated guidance information into the speech synthesis engine and converts it into audio data. The input is the generated guidance content (in text format), and the output is the playable audio data. The speech synthesis is performed using text-to-speech technology to generate speech that can be easily understood by the user.
[0313] Step 5:
[0314] The terminal plays audio data received from the server and provides guidance to the user. The input is audio data transmitted from the server, and the output is the user's auditory reception of the information. The terminal outputs the audio through a speaker or earphones via its playback function.
[0315] Step 6:
[0316] If the user needs further information after hearing the instructions, they can send additional requests to the server via the terminal. Input is the user's voice or text request, and output is an information request corresponding to the request's instructions. The terminal collects new requests through the user interface and sends them to the server.
[0317] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0318] This invention provides a system that combines a high-precision location information acquisition method, an emotion engine, and a generative AI to provide voice guidance that takes into account the user's emotional state. This system collects information about tourist spots based on the user's location information, recognizes the user's emotions, and then generates personalized voice guidance accordingly.
[0319] Server Role
[0320] The server first receives the user's location information transmitted from the terminal. It analyzes the location information and retrieves data on relevant tourist spots from a database. Furthermore, the emotion engine receives the user's emotional state and provides this to the generating AI. The generating AI uses this information to generate voice guidance with appropriate tone and content based on the user's emotions. This voice guidance is adjusted to a calm tone if the user wants to relax, and to include details if the AI determines the user is interested. The generated voice data is then transmitted to the terminal and provided to the user.
[0321] Terminal role
[0322] The device acquires the user's location information with high accuracy and transmits it to the server in real time. Through an emotion engine, it estimates the user's emotions from their voice, facial expressions, and actions, and transmits this information to the server. The device also receives voice guidance data from the server and plays it back to the user. The content of the voice guidance is applied based on emotion recognition, serving to further personalize the user's experience.
[0323] User experience
[0324] Users can access the system via their smartphones to receive detailed, personalized audio guidance about tourist attractions. For example, if the emotion engine detects that a user is slightly tired while standing in front of a historical building, the guidance will be shorter and focus on key points. Conversely, if it detects high levels of excitement or interest, additional anecdotes and trivia will be provided. This allows users to effortlessly receive information tailored to their emotional state, making the sightseeing experience richer and more convenient.
[0325] Thus, the present invention integrates location information, emotion recognition, and AI technology to realize a system that provides a high level of satisfaction to visitors.
[0326] The following describes the processing flow.
[0327] Step 1:
[0328] The device acquires location information and emotion data. The device uses its built-in high-precision GPS function to determine the user's current location and activates an emotion engine that analyzes the user's emotions from their facial expressions and voice using the camera and microphone.
[0329] Step 2:
[0330] The device sends location information and emotional data acquired by the device to the server. The device packages the analyzed emotional state and location information and sends the data to the server in real time.
[0331] Step 3:
[0332] The server retrieves tourist spot information based on location data. The server accesses a database and extracts detailed information about the tourist spot corresponding to the transmitted location data.
[0333] Step 4:
[0334] The server adjusts the voice guidance content based on the user's emotions. The server provides the received emotion data to the generating AI, which then edits the tourist spot information guidance to match the user's mood. For excited users, it provides detailed and abundant information, while for users seeking relaxation, it provides concise and calming content.
[0335] Step 5:
[0336] The server generates voice guidance data and sends it to the terminal. The voice data, generated according to the user's emotional state, is sent to the terminal via the network, and the guidance is ready for the user.
[0337] Step 6:
[0338] The device plays voice guidance to the user. The device plays the received voice data and provides the user with customized guidance information. This process allows the user to receive the best possible guidance experience tailored to their emotions.
[0339] Step 7:
[0340] If needed, users can request further information or additional questions via the terminal. If a user wants to learn more about information they are interested in, they can operate the terminal to send additional requests to the server.
[0341] Step 8:
[0342] The device sends the user's additional request to the server. The device organizes the user's new request and sends it to the server along with sentiment data.
[0343] Step 9:
[0344] The server generates new information and sends additional voice guidance, adjusted according to the user's emotions, to the device. The server uses a generation AI to generate new relevant information and sends the readjusted voice data back to the device.
[0345] Step 10:
[0346] The device plays additional audio to provide users with deeper guidance. The device plays new audio guidance to further enrich the user experience. This ensures that users always receive information tailored to their interests and emotions.
[0347] (Example 2)
[0348] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0349] Conventional voice guidance systems struggled to provide personalized information that took into account the user's emotional state, and merely offered general directions based on location information. Furthermore, they lacked sufficient real-time information updates and optimization of voice guidance according to the user's emotional state, resulting in a limited user experience.
[0350] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0351] In this invention, the server includes a high-precision positioning means, a means for providing voice guidance information related to a specific area based on the positioning, and a means for generating additional information based on the user's emotional state using a generation AI model. This makes it possible to provide real-time optimized voice guidance that takes into account the user's emotional state.
[0352] "High-precision positioning means" refers to technology that acquires the user's location as latitude and longitude data with high accuracy.
[0353] "Means for providing audio guidance information related to a specific region" refers to technology that provides cultural, historical, and tourist information about a location in audio format based on acquired location information.
[0354] A "generative AI model" refers to the structure of artificial intelligence that uses machine learning algorithms to generate optimal additional information based on input data.
[0355] "Means for analyzing a user's emotional state" refers to technology that uses audio and video data acquired from microphones and cameras to identify the emotions a user is currently experiencing.
[0356] "Means for adjusting the tone and level of detail of voice guidance information" refers to technologies that change the atmosphere and depth of content of the voice information provided according to the user's emotional state.
[0357] "A means of tracking in real time and instantly updating location-specific guidance information" refers to technology that continuously monitors the user's movements and provides the latest, appropriate information whenever their location changes.
[0358] "Adaptive means" refers to technologies that personalize the information provided based on past data and user preferences.
[0359] This invention is a system that provides personalized voice guidance to users by combining multiple technological elements. Specifically, it utilizes high-precision positioning means, emotion analysis functions, generative AI models, and voice output functions.
[0360] Hardware and software to use
[0361] The device is equipped with a GPS module, microphone, and camera, which are used to acquire the user's location and emotional state. The location information is transmitted to the server in real time, and the audio and video data from the microphone and camera are processed by emotion analysis software. The emotion analysis software infers the user's emotional state based on their voice tone and facial expressions.
[0362] The server generates voice guidance using a generative AI model based on information sent from the terminal. The generative AI model receives regional information based on the user's location and emotional state as input, and generates prompt sentences. Based on this, it creates guidance with appropriate content and tone. This generated voice guidance data is then sent back to the terminal.
[0363] Specific example
[0364] For example, suppose a user is standing in front of a historical building. The device identifies its location via GPS and informs the server that it is a tourist spot. Simultaneously, the microphone captures the user's voice, and emotion analysis software recognizes that the user is slightly tired. Once this information is sent to the server, the server provides a prompt message to the AI model, such as "The user is in front of a historical building. Their emotional state is fatigued. Please create a concise guide," and sends the generated guide to the device.
[0365] In this way, users can obtain information that matches their emotions at the time, making their sightseeing experience more fulfilling. An example of a prompt might be: "The user is in front of Kinkaku-ji Temple. Their emotional state is excited. Please create an audio guide that includes detailed history and anecdotes."
[0366] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0367] Step 1:
[0368] The device acquires the user's current location using a high-precision GPS module. The GPS receiver analyzes satellite data to calculate latitude and longitude information. This location information is sent to the server as output data from the device.
[0369] Step 2:
[0370] The device captures the user's voice and facial expressions using a microphone and camera. The digital data obtained from the microphone and camera is input into emotion analysis software. The software analyzes the phoneme and visual features to identify the user's emotional state (e.g., excitement, fatigue). This emotional information is sent to the server as output data from the device.
[0371] Step 3:
[0372] The server searches its database for relevant tourist spot information based on location information received from the terminal. Once the location information is entered, data for tourist spots in the corresponding area is retrieved as server output. This includes the spot's name, description, and any special events.
[0373] Step 4:
[0374] The server inputs emotional information received from the terminal and acquired tourist spot information into a generating AI model. The generating AI model generates prompt sentences and creates personalized voice guidance based on them. The input data consists of emotional information and tourist spot information, and the output is voice guidance data tailored to the user's emotions.
[0375] Step 5:
[0376] The server sends the generated audio guidance data to the terminal. The terminal receives this audio data and plays it for the user through its built-in speaker. The audio guidance is tailored to enhance the user experience. The terminal's final output is for the user to listen to the audio guidance and gain a deeper understanding of the tourist destination.
[0377] (Application Example 2)
[0378] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0379] In recent years, many physical facilities have been required to improve the visitor experience. However, providing personalized guidance information that takes into account visitors' emotions and preferences in real time is technically complex, and achieving satisfactory results is often difficult. Therefore, there is a need to develop a system that provides a personalized experience based on the visitor's location and emotions.
[0380] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0381] In this invention, the server includes a high-precision location information acquisition device, a device that provides voice guidance information related to a specific facility based on the location information, a device that generates additional information customized based on customer preferences and past interactions using generative AI technology, a device that includes a process of analyzing the customer's emotional state using an emotion engine and personalizing the additional information, and a device that outputs the personalized information as voice data. As a result, visitors can receive optimal guidance information that matches their emotions and preferences at that moment.
[0382] A "high-precision location information acquisition device" is a device that accurately grasps the user's current location in real time and provides necessary information.
[0383] "Audio guidance information" refers to content that presents useful information about the facility or place the user is visiting in audio format.
[0384] "Generative AI technology" is an artificial intelligence technology that generates personalized additional information based on the user's past behavior and preferences.
[0385] An "emotion engine" is software or hardware that analyzes data such as a user's voice and facial expressions to estimate their emotional state.
[0386] "Personalized information" refers to guidance information and content that is optimized based on the individual user's preferences and emotional state.
[0387] "Audio data is digital data used to convey information generated by a device to the user as sound."
[0388] This invention provides a personalized guide system to enhance the customer experience in tourist destinations and commercial facilities. This system acquires customer location information with high accuracy and analyzes their emotional state to provide more personalized guidance. The main functions of the system and its embodiments are described below.
[0389] The server uses a high-precision location acquisition device (e.g., a GPS module) to determine the user's current location. Next, it retrieves data on relevant facilities and tourist spots from a database based on the location information and provides it as voice guidance information.
[0390] The server also uses an emotion engine (e.g., Affectiva's SDK) to estimate the user's emotional state in real time from their voice and facial expression data. This emotional state is then input into generative AI technology (e.g., OpenAI's GPT-3). The generative AI considers the user's preferences and past interactions to generate personalized additional information. This generated information is output as personalized voice data.
[0391] For example, if a user smiles when visiting a specific facility in a tourist area, the emotion engine will determine that the user is feeling excited. This information is then provided to the generative AI model, which, using a prompt such as, "Please explain in detail the local history and manufacturing process of the local specialty that the tourist is interested in," provides detailed location-related information via voice.
[0392] In this way, the present invention can significantly improve the experience at commercial facilities and tourist destinations by utilizing location information and sentiment analysis to provide each customer with the most suitable information in real time.
[0393] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0394] Step 1:
[0395] The terminal acquires the user's location information using a high-precision location information acquisition device and transmits that data to the server. The input is the location information acquired by the terminal, and the output is the transmission of location information data to the server. Through this process, the server can recognize the user's current location and acquire appropriate facility data.
[0396] Step 2:
[0397] The server analyzes the received location information and retrieves information about facilities and tourist attractions associated with that location from a database. The input is the user's location information, and the output is data about the relevant facilities. Based on the location information, the server searches for relevant information and prepares data according to the user's actions.
[0398] Step 3:
[0399] The device acquires the user's voice and facial expression data, estimates the emotional state using an emotion engine, and sends that data to the server. The input is the user's voice and facial expression data, and the output is the estimated emotional state. Here, the device collects data using voice and video sensors.
[0400] Step 4:
[0401] The server inputs emotional state and facility data into a generating AI model, which then uses prompts to generate personalized additional information. The input is the user's emotional state and facility data, and the output is customized guidance information. The generating AI model uses this information to create appropriate guidance content.
[0402] Step 5:
[0403] The server outputs the generated additional information as audio data and sends it to the terminal. The input is personalized information from the generating AI, and the output is audio guidance data played on the terminal. In this process, the server uses speech synthesis technology to convert the digital data into audio.
[0404] Step 6:
[0405] The user receives audio data through their device and listens to personalized guidance in real time. The input is audio data from the server, and the output is the audio information received by the user. This allows the user to receive optimal guidance tailored to their emotions and interests.
[0406] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0407] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0408] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0409] [Third Embodiment]
[0410] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0411] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0412] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0413] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0414] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0415] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0416] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0417] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0418] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0419] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0420] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0421] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0422] This invention is a system that uses a high-precision location information acquisition method to provide users with voice guidance on information about specific tourist destinations. The following describes the system's program-based processing in natural language.
[0423] Server Role
[0424] The server retrieves relevant information from a database based on location data and uses generative AI to generate voice guidance tailored to the user. Specifically, the server receives the user's location information transmitted from the device and retrieves data on tourist spots related to that location from the database. This data includes historical background, cultural significance, and recommended points of interest. Furthermore, the server also utilizes generative AI to generate additional information in response to specific user requests. The generated information is converted into voice data and sent to the device.
[0425] Terminal role
[0426] The device acquires the user's location information in real time and sends it to the server. By playing the received audio data back to the user, it provides tourist information. The device also has the function to accept user input and send additional requests to the server. This allows for flexible responses when the user asks additional questions or shows new interests.
[0427] User experience
[0428] While visiting tourist attractions, users can deepen their understanding of the sites by receiving location-specific audio guidance. By carrying a smartphone, users can listen to explanations based on their current location, obtaining diverse information in real time without using their hands. For example, if a user is in front of a historical building, they can hear about the building's history, architect, and related anecdotes. Furthermore, if the user wants to know more details about the building, they can request additional information by operating their device. In response to this request, the server generates more detailed information and related stories, delivering them to the user as audio data.
[0429] Thus, the system of the present invention makes it possible to enrich the visitor experience at tourist destinations and efficiently provide personalized information.
[0430] The following describes the processing flow.
[0431] Step 1:
[0432] The device acquires location information. The device uses its built-in GPS function and high-precision location measurement device to determine the user's current location. This location information is updated at regular intervals and immediately transmitted to the server.
[0433] Step 2:
[0434] The server receives location information and retrieves corresponding tourist destination data. The server accesses the database and searches for information on the relevant tourist spot based on the received location information. The retrieved information includes historical background, points of interest, and related anecdotes.
[0435] Step 3:
[0436] The server uses a generation AI to create voice guidance data. Based on the user's location and specific request, the server uses the generation AI to generate appropriate voice guidance. The generated text data is then converted into voice data, preparing it for delivery to the user.
[0437] Step 4:
[0438] The server sends audio data to the terminal. The server sends the generated audio data to the terminal via the network and prepares it for voice guidance to the user.
[0439] Step 5:
[0440] The device plays an audio guide. The device plays the received audio data to the user, providing information about a specific tourist spot. Users can deepen their understanding of the tourist spot by listening to the audio guide through their smartphone.
[0441] Step 6:
[0442] Users can request additional information. After listening to the voice guidance, users can send requests via their device if they want more detailed information or further questions.
[0443] Step 7:
[0444] The terminal sends a request to the server. The terminal sends a request for additional user information and questions to the server in text format, asking for the necessary information to be provided.
[0445] Step 8:
[0446] The server generates and sends information based on the request. The server analyzes the user's request and retrieves new relevant information from the database or generates it using a generative AI. The newly generated audio data is sent to the terminal and provided to the user.
[0447] Step 9:
[0448] The device plays additional information. The device plays additional audio data received from the server, providing the user with further sightseeing information. By repeating this process, the user can obtain a detailed and personalized sightseeing experience.
[0449] (Example 1)
[0450] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0451] Current tourist information systems have limitations in providing real-time information and generating additional information based on the user's location, and they struggle to effectively deliver personalized experiences. In particular, their inability to provide dynamic information tailored to the user's interests and preferences is a major problem.
[0452] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0453] In this invention, the server includes a high-precision location information acquisition device, a device that provides voice guidance information related to a specific geographical area based on the location information, and a device that generates additional information in response to user requests using a generation AI model. This enables the real-time provision of personalized guidance information based on the user's current location information.
[0454] A "high-precision location information acquisition device" is a device that can accurately and precisely identify a user's geographical location and provide that information in real time.
[0455] A "geographical area" is a concept that refers to the physical extent of a place, including specific places of interest or tourist attractions.
[0456] "Audio guidance information" refers to data generated to provide users with information such as tourist attractions, geographical features, and historical background in audio format.
[0457] A "generative AI model" is an artificial intelligence program that uses natural language processing and machine learning to automatically generate information in response to user requests.
[0458] A "prompt statement" is an input statement that contains instructions or requests necessary for a generative AI model to provide meaningful information.
[0459] "Personalized guidance information" refers to information that is tailored to specific needs and interests based on the user's current location, past behavior, and preferences.
[0460] This invention is a system that uses a high-precision location information acquisition device to provide users with voice guidance on information about a specific geographical area. This makes it possible to provide users with real-time information about tourist spots they are visiting.
[0461] The server retrieves relevant geographical information from a database based on the user's location information transmitted from the terminal. At this time, the server organizes the information to provide clear and easy-to-understand guidance based on highly accurate location data. Furthermore, the server uses a generative AI model to generate additional information in response to the user's request. The generated information is then converted into audio data using speech synthesis technology and transmitted to the terminal.
[0462] The device transmits real-time location information obtained via its built-in GPS function to a server, and plays back audio data received from the server to the user. This allows the user to obtain information in a natural way.
[0463] Users can use their smartphones or tablets to listen to audio guides played from their devices, deepening their understanding of the background and highlights of tourist destinations. For example, when a user visits a particular historical building, they can receive detailed audio information about its history and background. Furthermore, if a user has specific interests or questions, they can make additional requests through their device, which can generate even more detailed information.
[0464] An example of a prompt message would be a request such as, "Please tell me about the founder of this historic building." In response, the server can use a generative AI model to retrieve appropriate information and provide it to the user as voice guidance.
[0465] In this way, the system can efficiently provide personalized tourist information to users.
[0466] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0467] Step 1:
[0468] The device uses GPS functionality to obtain the user's current location information. This location information is represented as latitude and longitude coordinates and sent to the server. The input is the user's real-time location information, and the output is the transmission of location data to the server.
[0469] Step 2:
[0470] The server retrieves information about the corresponding geographical area from its database based on the location information received from the device. This process uses the location information as a search query to extract details about relevant tourist attractions. The input is the location information from the device, and the output is the retrieved information about tourist attractions.
[0471] Step 3:
[0472] The server uses a generative AI model to construct the content of the voice guidance based on the acquired tourist spot information. It creates prompt sentences and inputs them into the generative AI model to generate appropriate voice guidance content. The input is tourist spot information and user prompt sentences, and the output is the generated voice guidance content.
[0473] Step 4:
[0474] The server uses a speech synthesis API to convert the generated voice guidance content into audio data. This process converts text data into an audio data format. The input is the voice guidance content generated by the AI model, and the output is playable audio data.
[0475] Step 5:
[0476] The server sends the converted audio data to the terminal. This communication is conducted securely over the internet. The input is audio data, and the output is the transmission of data to the terminal.
[0477] Step 6:
[0478] The terminal plays audio data received from the server to the user. Audio guidance is provided through the terminal's speaker or earphones, allowing the user to hear information about tourist destinations. The input is audio data from the server, and the output is the playback of audio guidance to the user.
[0479] Step 7:
[0480] If the user needs further information, they can operate the terminal to request additional information. This request is sent from the terminal to the server. Input is a new question or request from the user, and output is the request sent to the server.
[0481] Step 8:
[0482] The server receives the user's additional request and generates new information again using the generation AI model. This new information is then converted back into audio data and sent to the terminal. The input is the user's request and prompt text, and the output is the audio data of the additional information.
[0483] Step 9:
[0484] The device plays newly received audio data to provide the user with further instructions. The input is audio data containing additional information, and the output is the playback of the audio guidance to the user.
[0485] (Application Example 1)
[0486] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0487] In modern commercial facilities, the amount of information provided to visitors is limited, making it difficult to provide efficient guidance tailored to individual preferences. As a result, visitors are unable to effectively obtain information about products and stores, hindering their satisfaction. Furthermore, the lack of real-time location-based guidance prevents the provision of a dynamic shopping experience.
[0488] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0489] In this invention, the server includes means for acquiring high-precision location information, means for providing voice guidance data related to a specific location based on the location information, means for generating additional information in response to user requests using a generating AI, and means for providing user-specific guidance based on information within the commercial facility. This enables the provision of real-time, dynamic, and personalized information to individual visitors.
[0490] "High-precision location information acquisition method" refers to a technology that allows a user's precise current location to be determined in real time through a digital device.
[0491] "Voice guidance data" refers to digital data used to provide users with useful information in audio format.
[0492] "Generative AI" is an artificial intelligence technology that creates and outputs appropriate information based on user requests.
[0493] "Information within a commercial facility" refers to detailed information about products, stores, campaigns, etc., related to the commercial facility.
[0494] "Personalized guidance for users" refers to a service that provides customized information based on the individual user's preferences and behavioral history.
[0495] This system provides real-time location-based voice guidance services to visitors within commercial facilities. The implementation of this invention requires high-precision location information acquisition means, voice guidance data provision means, and information generation means using AI. Specific hardware includes portable digital devices such as smartphones and tablets, and built-in GPS modules. The software used includes a real-time location tracking API, an information provision server, a speech synthesis engine, and a generative AI model.
[0496] The server receives highly accurate location information from the user's device and uses that data to collect relevant information within the commercial facility. This information includes store, product, campaign, and event information. The generating AI considers the user's current location and request to create personalized guidance. The generated guidance is then converted into voice data using a speech synthesis engine.
[0497] The device has the function of playing guidance to the user via audio data provided by the server. The audio guidance can be customized based on the user's preferences and past conversation history. As a result, users can receive information hands-free, enriching their shopping experience.
[0498] For example, let's say a user is standing in front of a clothing store in a shopping mall. The user's smartphone sends its location information to a server, which then provides relevant promotions and information on popular products. The generated voice message might say, "This store is currently having a sale on its latest collection. Please request additional information if you need more details." An example of a prompt might be, "Please provide information on stores around latitude=35.6895, longitude=139.6917."
[0499] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0500] Step 1:
[0501] The device uses its built-in GPS module to obtain the user's current location in real time. The input is the device's GPS data, and the output is numerical information of latitude and longitude. This numerical information is sent to the server via the location API on the device.
[0502] Step 2:
[0503] The server receives location information transmitted from the terminal. Based on this, it retrieves relevant information within the commercial facility from the database. The input is latitude and longitude information sent from the terminal, and the output is detailed information about the corresponding store or product. This data is searched and retrieved using database queries.
[0504] Step 3:
[0505] The server uses the acquired information to input prompts into the generating AI model, which then generates user-optimized guidance information. The input is information about the commercial facility obtained from the database, and the output is user-specific guidance content. Based on these inputs, the generating AI performs natural language processing to construct guidance text appropriate to the context.
[0506] Step 4:
[0507] The server inputs the generated guidance information into a speech synthesis engine and converts it into audio data. The input is the generated guidance content (in text format), and the output is playable audio data. Speech synthesis is performed using text-to-speech technology to generate audio that is easy for the user to understand.
[0508] Step 5:
[0509] The terminal plays audio data received from the server and provides guidance to the user. The input is audio data transmitted from the server, and the output is the user's auditory reception of the information. The terminal outputs the audio through a speaker or earphones via its playback function.
[0510] Step 6:
[0511] If the user needs further information after hearing the instructions, they can send additional requests to the server via the terminal. Input is the user's voice or text request, and output is an information request corresponding to the request's instructions. The terminal collects new requests through the user interface and sends them to the server.
[0512] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0513] This invention provides a system that combines a high-precision location information acquisition method, an emotion engine, and a generative AI to provide voice guidance that takes into account the user's emotional state. This system collects information about tourist spots based on the user's location information, recognizes the user's emotions, and then generates personalized voice guidance accordingly.
[0514] Server Role
[0515] The server first receives the user's location information transmitted from the terminal. It analyzes the location information and retrieves data on relevant tourist spots from a database. Furthermore, the emotion engine receives the user's emotional state and provides this to the generating AI. The generating AI uses this information to generate voice guidance with appropriate tone and content based on the user's emotions. This voice guidance is adjusted to a calm tone if the user wants to relax, and to include details if the AI determines the user is interested. The generated voice data is then transmitted to the terminal and provided to the user.
[0516] Terminal role
[0517] The device acquires the user's location information with high accuracy and transmits it to the server in real time. Through an emotion engine, it estimates the user's emotions from their voice, facial expressions, and actions, and transmits this information to the server. The device also receives voice guidance data from the server and plays it back to the user. The content of the voice guidance is applied based on emotion recognition, serving to further personalize the user's experience.
[0518] User experience
[0519] Users can access the system via their smartphones to receive detailed, personalized audio guidance about tourist attractions. For example, if the emotion engine detects that a user is slightly tired while standing in front of a historical building, the guidance will be shorter and focus on key points. Conversely, if it detects high levels of excitement or interest, additional anecdotes and trivia will be provided. This allows users to effortlessly receive information tailored to their emotional state, making the sightseeing experience richer and more convenient.
[0520] Thus, the present invention integrates location information, emotion recognition, and AI technology to realize a system that provides a high level of satisfaction to visitors.
[0521] The following describes the processing flow.
[0522] Step 1:
[0523] The device acquires location information and emotion data. The device uses its built-in high-precision GPS function to determine the user's current location and activates an emotion engine that analyzes the user's emotions from their facial expressions and voice using the camera and microphone.
[0524] Step 2:
[0525] The device sends location information and emotional data acquired by the device to the server. The device packages the analyzed emotional state and location information and sends the data to the server in real time.
[0526] Step 3:
[0527] The server retrieves tourist spot information based on location data. The server accesses a database and extracts detailed information about the tourist spot corresponding to the transmitted location data.
[0528] Step 4:
[0529] The server adjusts the voice guidance content based on the user's emotions. The server provides the received emotion data to the generating AI, which then edits the tourist spot information guidance to match the user's mood. For excited users, it provides detailed and abundant information, while for users seeking relaxation, it provides concise and calming content.
[0530] Step 5:
[0531] The server generates voice guidance data and sends it to the terminal. The voice data, generated according to the user's emotional state, is sent to the terminal via the network, and the guidance is ready for the user.
[0532] Step 6:
[0533] The device plays voice guidance to the user. The device plays the received voice data and provides the user with customized guidance information. This process allows the user to receive the best possible guidance experience tailored to their emotions.
[0534] Step 7:
[0535] If needed, users can request further information or additional questions via the terminal. If a user wants to learn more about information they are interested in, they can operate the terminal to send additional requests to the server.
[0536] Step 8:
[0537] The device sends the user's additional request to the server. The device organizes the user's new request and sends it to the server along with sentiment data.
[0538] Step 9:
[0539] The server generates new information and sends additional voice guidance, adjusted according to the user's emotions, to the device. The server uses a generation AI to generate new relevant information and sends the readjusted voice data back to the device.
[0540] Step 10:
[0541] The device plays additional audio to provide users with deeper guidance. The device plays new audio guidance to further enrich the user experience. This ensures that users always receive information tailored to their interests and emotions.
[0542] (Example 2)
[0543] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0544] Conventional voice guidance systems struggled to provide personalized information that took into account the user's emotional state, and merely offered general directions based on location information. Furthermore, they lacked sufficient real-time information updates and optimization of voice guidance according to the user's emotional state, resulting in a limited user experience.
[0545] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0546] In this invention, the server includes a high-precision positioning means, a means for providing voice guidance information related to a specific area based on the positioning, and a means for generating additional information based on the user's emotional state using a generation AI model. This makes it possible to provide real-time optimized voice guidance that takes into account the user's emotional state.
[0547] "High-precision positioning means" refers to technology that acquires the user's location as latitude and longitude data with high accuracy.
[0548] "Means for providing audio guidance information related to a specific region" refers to technology that provides cultural, historical, and tourist information about a location in audio format based on acquired location information.
[0549] A "generative AI model" refers to the structure of artificial intelligence that uses machine learning algorithms to generate optimal additional information based on input data.
[0550] "Means for analyzing a user's emotional state" refers to technology that uses audio and video data acquired from microphones and cameras to identify the emotions a user is currently experiencing.
[0551] "Means for adjusting the tone and level of detail of voice guidance information" refers to technologies that change the atmosphere and depth of content of the voice information provided according to the user's emotional state.
[0552] "A means of tracking in real time and instantly updating location-specific guidance information" refers to technology that continuously monitors the user's movements and provides the latest, appropriate information whenever their location changes.
[0553] "Adaptive means" refers to technologies that personalize the information provided based on past data and user preferences.
[0554] This invention is a system that provides personalized voice guidance to users by combining multiple technological elements. Specifically, it utilizes high-precision positioning means, emotion analysis functions, generative AI models, and voice output functions.
[0555] Hardware and software to use
[0556] The device is equipped with a GPS module, microphone, and camera, which are used to acquire the user's location and emotional state. The location information is transmitted to the server in real time, and the audio and video data from the microphone and camera are processed by emotion analysis software. The emotion analysis software infers the user's emotional state based on their voice tone and facial expressions.
[0557] The server generates voice guidance using a generative AI model based on information sent from the terminal. The generative AI model receives regional information based on the user's location and emotional state as input, and generates prompt sentences. Based on this, it creates guidance with appropriate content and tone. This generated voice guidance data is then sent back to the terminal.
[0558] Specific example
[0559] For example, suppose a user is standing in front of a historical building. The device identifies its location via GPS and informs the server that it is a tourist spot. Simultaneously, the microphone captures the user's voice, and emotion analysis software recognizes that the user is slightly tired. Once this information is sent to the server, the server provides a prompt message to the AI model, such as "The user is in front of a historical building. Their emotional state is fatigued. Please create a concise guide," and sends the generated guide to the device.
[0560] In this way, users can obtain information that matches their emotions at the time, making their sightseeing experience more fulfilling. An example of a prompt might be: "The user is in front of Kinkaku-ji Temple. Their emotional state is excited. Please create an audio guide that includes detailed history and anecdotes."
[0561] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0562] Step 1:
[0563] The device acquires the user's current location using a high-precision GPS module. The GPS receiver analyzes satellite data to calculate latitude and longitude information. This location information is sent to the server as output data from the device.
[0564] Step 2:
[0565] The device captures the user's voice and facial expressions using a microphone and camera. The digital data obtained from the microphone and camera is input into emotion analysis software. The software analyzes the phoneme and visual features to identify the user's emotional state (e.g., excitement, fatigue). This emotional information is sent to the server as output data from the device.
[0566] Step 3:
[0567] The server searches its database for relevant tourist spot information based on location information received from the terminal. Once the location information is entered, data for tourist spots in the corresponding area is retrieved as server output. This includes the spot's name, description, and any special events.
[0568] Step 4:
[0569] The server inputs emotional information received from the terminal and acquired tourist spot information into a generating AI model. The generating AI model generates prompt sentences and creates personalized voice guidance based on them. The input data consists of emotional information and tourist spot information, and the output is voice guidance data tailored to the user's emotions.
[0570] Step 5:
[0571] The server sends the generated audio guidance data to the terminal. The terminal receives this audio data and plays it for the user through its built-in speaker. The audio guidance is tailored to enhance the user experience. The terminal's final output is for the user to listen to the audio guidance and gain a deeper understanding of the tourist destination.
[0572] (Application Example 2)
[0573] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0574] In recent years, many physical facilities have been required to improve the visitor experience. However, providing personalized guidance information that takes into account visitors' emotions and preferences in real time is technically complex, and achieving satisfactory results is often difficult. Therefore, there is a need to develop a system that provides a personalized experience based on the visitor's location and emotions.
[0575] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0576] In this invention, the server includes a high-precision location information acquisition device, a device that provides voice guidance information related to a specific facility based on the location information, a device that generates additional information customized based on customer preferences and past interactions using generative AI technology, a device that includes a process of analyzing the customer's emotional state using an emotion engine and personalizing the additional information, and a device that outputs the personalized information as voice data. As a result, visitors can receive optimal guidance information that matches their emotions and preferences at that moment.
[0577] A "high-precision location information acquisition device" is a device that accurately grasps the user's current location in real time and provides necessary information.
[0578] "Audio guidance information" refers to content that presents useful information about the facility or place the user is visiting in audio format.
[0579] "Generative AI technology" is an artificial intelligence technology that generates personalized additional information based on the user's past behavior and preferences.
[0580] An "emotion engine" is software or hardware that analyzes data such as a user's voice and facial expressions to estimate their emotional state.
[0581] "Personalized information" refers to guidance information and content that is optimized based on the individual user's preferences and emotional state.
[0582] "Audio data is digital data used to convey information generated by a device to the user as sound."
[0583] This invention provides a personalized guide system to enhance the customer experience in tourist destinations and commercial facilities. This system acquires customer location information with high accuracy and analyzes their emotional state to provide more personalized guidance. The main functions of the system and its embodiments are described below.
[0584] The server uses a high-precision location acquisition device (e.g., a GPS module) to determine the user's current location. Next, it retrieves data on relevant facilities and tourist spots from a database based on the location information and provides it as voice guidance information.
[0585] The server also uses an emotion engine (e.g., Affectiva's SDK) to estimate the user's emotional state in real time from their voice and facial expression data. This emotional state is then input into generative AI technology (e.g., OpenAI's GPT-3). The generative AI considers the user's preferences and past interactions to generate personalized additional information. This generated information is output as personalized voice data.
[0586] For example, if a user smiles when visiting a specific facility in a tourist area, the emotion engine will determine that the user is feeling excited. This information is then provided to the generative AI model, which, using a prompt such as, "Please explain in detail the local history and manufacturing process of the local specialty that the tourist is interested in," provides detailed location-related information via voice.
[0587] In this way, the present invention can significantly improve the experience at commercial facilities and tourist destinations by utilizing location information and sentiment analysis to provide each customer with the most suitable information in real time.
[0588] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0589] Step 1:
[0590] The terminal acquires the user's location information using a high-precision location information acquisition device and transmits that data to the server. The input is the location information acquired by the terminal, and the output is the transmission of location information data to the server. Through this process, the server can recognize the user's current location and acquire appropriate facility data.
[0591] Step 2:
[0592] The server analyzes the received location information and retrieves information about facilities and tourist attractions associated with that location from a database. The input is the user's location information, and the output is data about the relevant facilities. Based on the location information, the server searches for relevant information and prepares data according to the user's actions.
[0593] Step 3:
[0594] The device acquires the user's voice and facial expression data, estimates the emotional state using an emotion engine, and sends that data to the server. The input is the user's voice and facial expression data, and the output is the estimated emotional state. Here, the device collects data using voice and video sensors.
[0595] Step 4:
[0596] The server inputs emotional state and facility data into a generating AI model, which then uses prompts to generate personalized additional information. The input is the user's emotional state and facility data, and the output is customized guidance information. The generating AI model uses this information to create appropriate guidance content.
[0597] Step 5:
[0598] The server outputs the generated additional information as audio data and sends it to the terminal. The input is personalized information from the generating AI, and the output is audio guidance data played on the terminal. In this process, the server uses speech synthesis technology to convert the digital data into audio.
[0599] Step 6:
[0600] The user receives audio data through their device and listens to personalized guidance in real time. The input is audio data from the server, and the output is the audio information received by the user. This allows the user to receive optimal guidance tailored to their emotions and interests.
[0601] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0602] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0603] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0604] [Fourth Embodiment]
[0605] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0606] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0607] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0608] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0609] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0610] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0611] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0612] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0613] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0614] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0615] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0616] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0617] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0618] This invention is a system that uses a high-precision location information acquisition method to provide users with voice guidance on information about specific tourist destinations. The following describes the system's program-based processing in natural language.
[0619] Server Role
[0620] The server retrieves relevant information from a database based on location data and uses generative AI to generate voice guidance tailored to the user. Specifically, the server receives the user's location information transmitted from the device and retrieves data on tourist spots related to that location from the database. This data includes historical background, cultural significance, and recommended points of interest. Furthermore, the server also utilizes generative AI to generate additional information in response to specific user requests. The generated information is converted into voice data and sent to the device.
[0621] Terminal role
[0622] The device acquires the user's location information in real time and sends it to the server. By playing the received audio data back to the user, it provides tourist information. The device also has the function to accept user input and send additional requests to the server. This allows for flexible responses when the user asks additional questions or shows new interests.
[0623] User experience
[0624] While visiting tourist attractions, users can deepen their understanding of the sites by receiving location-specific audio guidance. By carrying a smartphone, users can listen to explanations based on their current location, obtaining diverse information in real time without using their hands. For example, if a user is in front of a historical building, they can hear about the building's history, architect, and related anecdotes. Furthermore, if the user wants to know more details about the building, they can request additional information by operating their device. In response to this request, the server generates more detailed information and related stories, delivering them to the user as audio data.
[0625] Thus, the system of the present invention makes it possible to enrich the visitor experience at tourist destinations and efficiently provide personalized information.
[0626] The following describes the processing flow.
[0627] Step 1:
[0628] The device acquires location information. The device uses its built-in GPS function and high-precision location measurement device to determine the user's current location. This location information is updated at regular intervals and immediately transmitted to the server.
[0629] Step 2:
[0630] The server receives location information and retrieves corresponding tourist destination data. The server accesses the database and searches for information on the relevant tourist spot based on the received location information. The retrieved information includes historical background, points of interest, and related anecdotes.
[0631] Step 3:
[0632] The server uses a generation AI to create voice guidance data. Based on the user's location and specific request, the server uses the generation AI to generate appropriate voice guidance. The generated text data is then converted into voice data, preparing it for delivery to the user.
[0633] Step 4:
[0634] The server sends audio data to the terminal. The server sends the generated audio data to the terminal via the network and prepares it for voice guidance to the user.
[0635] Step 5:
[0636] The device plays an audio guide. The device plays the received audio data to the user, providing information about a specific tourist spot. Users can deepen their understanding of the tourist spot by listening to the audio guide through their smartphone.
[0637] Step 6:
[0638] Users can request additional information. After listening to the voice guidance, users can send requests via their device if they want more detailed information or further questions.
[0639] Step 7:
[0640] The terminal sends a request to the server. The terminal sends a request for additional user information and questions to the server in text format, asking for the necessary information to be provided.
[0641] Step 8:
[0642] The server generates and sends information based on the request. The server analyzes the user's request and retrieves new relevant information from the database or generates it using a generative AI. The newly generated audio data is sent to the terminal and provided to the user.
[0643] Step 9:
[0644] The device plays additional information. The device plays additional audio data received from the server, providing the user with further sightseeing information. By repeating this process, the user can obtain a detailed and personalized sightseeing experience.
[0645] (Example 1)
[0646] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0647] Current tourist information systems have limitations in providing real-time information and generating additional information based on the user's location, and they struggle to effectively deliver personalized experiences. In particular, their inability to provide dynamic information tailored to the user's interests and preferences is a major problem.
[0648] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0649] In this invention, the server includes a high-precision location information acquisition device, a device that provides voice guidance information related to a specific geographical area based on the location information, and a device that generates additional information in response to user requests using a generation AI model. This enables the real-time provision of personalized guidance information based on the user's current location information.
[0650] A "high-precision location information acquisition device" is a device that can accurately and precisely identify a user's geographical location and provide that information in real time.
[0651] A "geographical area" is a concept that refers to the physical extent of a place, including specific places of interest or tourist attractions.
[0652] "Audio guidance information" refers to data generated to provide users with information such as tourist attractions, geographical features, and historical background in audio format.
[0653] A "generative AI model" is an artificial intelligence program that uses natural language processing and machine learning to automatically generate information in response to user requests.
[0654] A "prompt statement" is an input statement that contains instructions or requests necessary for a generative AI model to provide meaningful information.
[0655] "Personalized guidance information" refers to information that is tailored to specific needs and interests based on the user's current location, past behavior, and preferences.
[0656] This invention is a system that uses a high-precision location information acquisition device to provide users with voice guidance on information about a specific geographical area. This makes it possible to provide users with real-time information about tourist spots they are visiting.
[0657] The server retrieves relevant geographical information from a database based on the user's location information transmitted from the terminal. At this time, the server organizes the information to provide clear and easy-to-understand guidance based on highly accurate location data. Furthermore, the server uses a generative AI model to generate additional information in response to the user's request. The generated information is then converted into audio data using speech synthesis technology and transmitted to the terminal.
[0658] The device transmits real-time location information obtained via its built-in GPS function to a server, and plays back audio data received from the server to the user. This allows the user to obtain information in a natural way.
[0659] Users can use their smartphones or tablets to listen to audio guides played from their devices, deepening their understanding of the background and highlights of tourist destinations. For example, when a user visits a particular historical building, they can receive detailed audio information about its history and background. Furthermore, if a user has specific interests or questions, they can make additional requests through their device, which can generate even more detailed information.
[0660] An example of a prompt message would be a request such as, "Please tell me about the founder of this historic building." In response, the server can use a generative AI model to retrieve appropriate information and provide it to the user as voice guidance.
[0661] In this way, the system can efficiently provide personalized tourist information to users.
[0662] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0663] Step 1:
[0664] The device uses GPS functionality to obtain the user's current location information. This location information is represented as latitude and longitude coordinates and sent to the server. The input is the user's real-time location information, and the output is the transmission of location data to the server.
[0665] Step 2:
[0666] The server retrieves information about the corresponding geographical area from its database based on the location information received from the device. This process uses the location information as a search query to extract details about relevant tourist attractions. The input is the location information from the device, and the output is the retrieved information about tourist attractions.
[0667] Step 3:
[0668] The server uses a generative AI model to construct the content of the voice guidance based on the acquired tourist spot information. It creates prompt sentences and inputs them into the generative AI model to generate appropriate voice guidance content. The input is tourist spot information and user prompt sentences, and the output is the generated voice guidance content.
[0669] Step 4:
[0670] The server uses a speech synthesis API to convert the generated voice guidance content into audio data. This process converts text data into an audio data format. The input is the voice guidance content generated by the AI model, and the output is playable audio data.
[0671] Step 5:
[0672] The server sends the converted audio data to the terminal. This communication is conducted securely over the internet. The input is audio data, and the output is the transmission of data to the terminal.
[0673] Step 6:
[0674] The terminal plays audio data received from the server to the user. Audio guidance is provided through the terminal's speaker or earphones, allowing the user to hear information about tourist destinations. The input is audio data from the server, and the output is the playback of audio guidance to the user.
[0675] Step 7:
[0676] If the user needs further information, they can operate the terminal to request additional information. This request is sent from the terminal to the server. Input is a new question or request from the user, and output is the request sent to the server.
[0677] Step 8:
[0678] The server receives the user's additional request and generates new information again using the generation AI model. This new information is then converted back into audio data and sent to the terminal. The input is the user's request and prompt text, and the output is the audio data of the additional information.
[0679] Step 9:
[0680] The device plays newly received audio data to provide the user with further instructions. The input is audio data containing additional information, and the output is the playback of the audio guidance to the user.
[0681] (Application Example 1)
[0682] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0683] In modern commercial facilities, the amount of information provided to visitors is limited, making it difficult to provide efficient guidance tailored to individual preferences. As a result, visitors are unable to effectively obtain information about products and stores, hindering their satisfaction. Furthermore, the lack of real-time location-based guidance prevents the provision of a dynamic shopping experience.
[0684] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0685] In this invention, the server includes means for acquiring high-precision location information, means for providing voice guidance data related to a specific location based on the location information, means for generating additional information in response to user requests using a generating AI, and means for providing user-specific guidance based on information within the commercial facility. This enables the provision of real-time, dynamic, and personalized information to individual visitors.
[0686] "High-precision location information acquisition method" refers to a technology that allows a user's precise current location to be determined in real time through a digital device.
[0687] "Voice guidance data" refers to digital data used to provide users with useful information in audio format.
[0688] "Generative AI" is an artificial intelligence technology that creates and outputs appropriate information based on user requests.
[0689] "Information within a commercial facility" refers to detailed information about products, stores, campaigns, etc., related to the commercial facility.
[0690] "Personalized guidance for users" refers to a service that provides customized information based on the individual user's preferences and behavioral history.
[0691] This system provides real-time location-based voice guidance services to visitors within commercial facilities. The implementation of this invention requires high-precision location information acquisition means, voice guidance data provision means, and information generation means using AI. Specific hardware includes portable digital devices such as smartphones and tablets, and built-in GPS modules. The software used includes a real-time location tracking API, an information provision server, a speech synthesis engine, and a generative AI model.
[0692] The server receives highly accurate location information from the user's device and uses that data to collect relevant information within the commercial facility. This information includes store, product, campaign, and event information. The generating AI considers the user's current location and request to create personalized guidance. The generated guidance is then converted into voice data using a speech synthesis engine.
[0693] The device has the function of playing guidance to the user via audio data provided by the server. The audio guidance can be customized based on the user's preferences and past conversation history. As a result, users can receive information hands-free, enriching their shopping experience.
[0694] For example, let's say a user is standing in front of a clothing store in a shopping mall. The user's smartphone sends its location information to a server, which then provides relevant promotions and information on popular products. The generated voice message might say, "This store is currently having a sale on its latest collection. Please request additional information if you need more details." An example of a prompt might be, "Please provide information on stores around latitude=35.6895, longitude=139.6917."
[0695] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0696] Step 1:
[0697] The device uses its built-in GPS module to obtain the user's current location in real time. The input is the device's GPS data, and the output is numerical information of latitude and longitude. This numerical information is sent to the server via the location API on the device.
[0698] Step 2:
[0699] The server receives location information transmitted from the terminal. Based on this, it retrieves relevant information within the commercial facility from the database. The input is latitude and longitude information sent from the terminal, and the output is detailed information about the corresponding store or product. This data is searched and retrieved using database queries.
[0700] Step 3:
[0701] The server uses the acquired information to input prompts into the generating AI model, which then generates user-optimized guidance information. The input is information about the commercial facility obtained from the database, and the output is user-specific guidance content. Based on these inputs, the generating AI performs natural language processing to construct guidance text appropriate to the context.
[0702] Step 4:
[0703] The server inputs the generated guidance information into a speech synthesis engine and converts it into audio data. The input is the generated guidance content (in text format), and the output is playable audio data. Speech synthesis is performed using text-to-speech technology to generate audio that is easy for the user to understand.
[0704] Step 5:
[0705] The terminal plays audio data received from the server and provides guidance to the user. The input is audio data transmitted from the server, and the output is the user's auditory reception of the information. The terminal outputs the audio through a speaker or earphones via its playback function.
[0706] Step 6:
[0707] If the user needs further information after hearing the instructions, they can send additional requests to the server via the terminal. Input is the user's voice or text request, and output is an information request corresponding to the request's instructions. The terminal collects new requests through the user interface and sends them to the server.
[0708] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0709] This invention provides a system that combines a high-precision location information acquisition method, an emotion engine, and a generative AI to provide voice guidance that takes into account the user's emotional state. This system collects information about tourist spots based on the user's location information, recognizes the user's emotions, and then generates personalized voice guidance accordingly.
[0710] Server Role
[0711] The server first receives the user's location information transmitted from the terminal. It analyzes the location information and retrieves data on relevant tourist spots from a database. Furthermore, the emotion engine receives the user's emotional state and provides this to the generating AI. The generating AI uses this information to generate voice guidance with appropriate tone and content based on the user's emotions. This voice guidance is adjusted to a calm tone if the user wants to relax, and to include details if the AI determines the user is interested. The generated voice data is then transmitted to the terminal and provided to the user.
[0712] Terminal role
[0713] The device acquires the user's location information with high accuracy and transmits it to the server in real time. Through an emotion engine, it estimates the user's emotions from their voice, facial expressions, and actions, and transmits this information to the server. The device also receives voice guidance data from the server and plays it back to the user. The content of the voice guidance is applied based on emotion recognition, serving to further personalize the user's experience.
[0714] User experience
[0715] Users can access the system via their smartphones to receive detailed, personalized audio guidance about tourist attractions. For example, if the emotion engine detects that a user is slightly tired while standing in front of a historical building, the guidance will be shorter and focus on key points. Conversely, if it detects high levels of excitement or interest, additional anecdotes and trivia will be provided. This allows users to effortlessly receive information tailored to their emotional state, making the sightseeing experience richer and more convenient.
[0716] Thus, the present invention integrates location information, emotion recognition, and AI technology to realize a system that provides a high level of satisfaction to visitors.
[0717] The following describes the processing flow.
[0718] Step 1:
[0719] The device acquires location information and emotion data. The device uses its built-in high-precision GPS function to determine the user's current location and activates an emotion engine that analyzes the user's emotions from their facial expressions and voice using the camera and microphone.
[0720] Step 2:
[0721] The device sends location information and emotional data acquired by the device to the server. The device packages the analyzed emotional state and location information and sends the data to the server in real time.
[0722] Step 3:
[0723] The server retrieves tourist spot information based on location data. The server accesses a database and extracts detailed information about the tourist spot corresponding to the transmitted location data.
[0724] Step 4:
[0725] The server adjusts the voice guidance content based on the user's emotions. The server provides the received emotion data to the generating AI, which then edits the tourist spot information guidance to match the user's mood. For excited users, it provides detailed and abundant information, while for users seeking relaxation, it provides concise and calming content.
[0726] Step 5:
[0727] The server generates voice guidance data and sends it to the terminal. The voice data, generated according to the user's emotional state, is sent to the terminal via the network, and the guidance is ready for the user.
[0728] Step 6:
[0729] The device plays voice guidance to the user. The device plays the received voice data and provides the user with customized guidance information. This process allows the user to receive the best possible guidance experience tailored to their emotions.
[0730] Step 7:
[0731] If needed, users can request further information or additional questions via the terminal. If a user wants to learn more about information they are interested in, they can operate the terminal to send additional requests to the server.
[0732] Step 8:
[0733] The device sends the user's additional request to the server. The device organizes the user's new request and sends it to the server along with sentiment data.
[0734] Step 9:
[0735] The server generates new information and sends additional voice guidance, adjusted according to the user's emotions, to the device. The server uses a generation AI to generate new relevant information and sends the readjusted voice data back to the device.
[0736] Step 10:
[0737] The device plays additional audio to provide users with deeper guidance. The device plays new audio guidance to further enrich the user experience. This ensures that users always receive information tailored to their interests and emotions.
[0738] (Example 2)
[0739] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0740] Conventional voice guidance systems struggled to provide personalized information that took into account the user's emotional state, and merely offered general directions based on location information. Furthermore, they lacked sufficient real-time information updates and optimization of voice guidance according to the user's emotional state, resulting in a limited user experience.
[0741] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0742] In this invention, the server includes a high-precision positioning means, a means for providing voice guidance information related to a specific area based on the positioning, and a means for generating additional information based on the user's emotional state using a generation AI model. This makes it possible to provide real-time optimized voice guidance that takes into account the user's emotional state.
[0743] "High-precision positioning means" refers to technology that acquires the user's location as latitude and longitude data with high accuracy.
[0744] "Means for providing audio guidance information related to a specific region" refers to technology that provides cultural, historical, and tourist information about a location in audio format based on acquired location information.
[0745] A "generative AI model" refers to the structure of artificial intelligence that uses machine learning algorithms to generate optimal additional information based on input data.
[0746] "Means for analyzing a user's emotional state" refers to technology that uses audio and video data acquired from microphones and cameras to identify the emotions a user is currently experiencing.
[0747] "Means for adjusting the tone and level of detail of voice guidance information" refers to technologies that change the atmosphere and depth of content of the voice information provided according to the user's emotional state.
[0748] "A means of tracking in real time and instantly updating location-specific guidance information" refers to technology that continuously monitors the user's movements and provides the latest, appropriate information whenever their location changes.
[0749] "Adaptive means" refers to technologies that personalize the information provided based on past data and user preferences.
[0750] This invention is a system that provides personalized voice guidance to users by combining multiple technological elements. Specifically, it utilizes high-precision positioning means, emotion analysis functions, generative AI models, and voice output functions.
[0751] Hardware and software to use
[0752] The device is equipped with a GPS module, microphone, and camera, which are used to acquire the user's location and emotional state. The location information is transmitted to the server in real time, and the audio and video data from the microphone and camera are processed by emotion analysis software. The emotion analysis software infers the user's emotional state based on their voice tone and facial expressions.
[0753] The server generates voice guidance using a generative AI model based on information sent from the terminal. The generative AI model receives regional information based on the user's location and emotional state as input, and generates prompt sentences. Based on this, it creates guidance with appropriate content and tone. This generated voice guidance data is then sent back to the terminal.
[0754] Specific example
[0755] For example, suppose a user is standing in front of a historical building. The device identifies its location via GPS and informs the server that it is a tourist spot. Simultaneously, the microphone captures the user's voice, and emotion analysis software recognizes that the user is slightly tired. Once this information is sent to the server, the server provides a prompt message to the AI model, such as "The user is in front of a historical building. Their emotional state is fatigued. Please create a concise guide," and sends the generated guide to the device.
[0756] In this way, users can obtain information that matches their emotions at the time, making their sightseeing experience more fulfilling. An example of a prompt might be: "The user is in front of Kinkaku-ji Temple. Their emotional state is excited. Please create an audio guide that includes detailed history and anecdotes."
[0757] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0758] Step 1:
[0759] The device acquires the user's current location using a high-precision GPS module. The GPS receiver analyzes satellite data to calculate latitude and longitude information. This location information is sent to the server as output data from the device.
[0760] Step 2:
[0761] The device captures the user's voice and facial expressions using a microphone and camera. The digital data obtained from the microphone and camera is input into emotion analysis software. The software analyzes the phoneme and visual features to identify the user's emotional state (e.g., excitement, fatigue). This emotional information is sent to the server as output data from the device.
[0762] Step 3:
[0763] The server searches its database for relevant tourist spot information based on location information received from the terminal. Once the location information is entered, data for tourist spots in the corresponding area is retrieved as server output. This includes the spot's name, description, and any special events.
[0764] Step 4:
[0765] The server inputs emotional information received from the terminal and acquired tourist spot information into a generating AI model. The generating AI model generates prompt sentences and creates personalized voice guidance based on them. The input data consists of emotional information and tourist spot information, and the output is voice guidance data tailored to the user's emotions.
[0766] Step 5:
[0767] The server sends the generated audio guidance data to the terminal. The terminal receives this audio data and plays it for the user through its built-in speaker. The audio guidance is tailored to enhance the user experience. The terminal's final output is for the user to listen to the audio guidance and gain a deeper understanding of the tourist destination.
[0768] (Application Example 2)
[0769] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0770] In recent years, many physical facilities have been required to improve the visitor experience. However, providing personalized guidance information that takes into account visitors' emotions and preferences in real time is technically complex, and achieving satisfactory results is often difficult. Therefore, there is a need to develop a system that provides a personalized experience based on the visitor's location and emotions.
[0771] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0772] In this invention, the server includes a high-precision location information acquisition device, a device that provides voice guidance information related to a specific facility based on the location information, a device that generates additional information customized based on customer preferences and past interactions using generative AI technology, a device that includes a process of analyzing the customer's emotional state using an emotion engine and personalizing the additional information, and a device that outputs the personalized information as voice data. As a result, visitors can receive optimal guidance information that matches their emotions and preferences at that moment.
[0773] A "high-precision location information acquisition device" is a device that accurately grasps the user's current location in real time and provides necessary information.
[0774] "Audio guidance information" refers to content that presents useful information about the facility or place the user is visiting in audio format.
[0775] "Generative AI technology" is an artificial intelligence technology that generates personalized additional information based on the user's past behavior and preferences.
[0776] An "emotion engine" is software or hardware that analyzes data such as a user's voice and facial expressions to estimate their emotional state.
[0777] "Personalized information" refers to guidance information and content that is optimized based on the individual user's preferences and emotional state.
[0778] "Audio data is digital data used to convey information generated by a device to the user as sound."
[0779] This invention provides a personalized guide system to enhance the customer experience in tourist destinations and commercial facilities. This system acquires customer location information with high accuracy and analyzes their emotional state to provide more personalized guidance. The main functions of the system and its embodiments are described below.
[0780] The server uses a high-precision location acquisition device (e.g., a GPS module) to determine the user's current location. Next, it retrieves data on relevant facilities and tourist spots from a database based on the location information and provides it as voice guidance information.
[0781] The server also uses an emotion engine (e.g., Affectiva's SDK) to estimate the user's emotional state in real time from their voice and facial expression data. This emotional state is then input into generative AI technology (e.g., OpenAI's GPT-3). The generative AI considers the user's preferences and past interactions to generate personalized additional information. This generated information is output as personalized voice data.
[0782] For example, if a user smiles when visiting a specific facility in a tourist area, the emotion engine will determine that the user is feeling excited. This information is then provided to the generative AI model, which, using a prompt such as, "Please explain in detail the local history and manufacturing process of the local specialty that the tourist is interested in," provides detailed location-related information via voice.
[0783] In this way, the present invention can significantly improve the experience at commercial facilities and tourist destinations by utilizing location information and sentiment analysis to provide each customer with the most suitable information in real time.
[0784] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0785] Step 1:
[0786] The terminal acquires the user's location information using a high-precision location information acquisition device and transmits that data to the server. The input is the location information acquired by the terminal, and the output is the transmission of location information data to the server. Through this process, the server can recognize the user's current location and acquire appropriate facility data.
[0787] Step 2:
[0788] The server analyzes the received location information and retrieves information about facilities and tourist attractions associated with that location from a database. The input is the user's location information, and the output is data about the relevant facilities. Based on the location information, the server searches for relevant information and prepares data according to the user's actions.
[0789] Step 3:
[0790] The device acquires the user's voice and facial expression data, estimates the emotional state using an emotion engine, and sends that data to the server. The input is the user's voice and facial expression data, and the output is the estimated emotional state. Here, the device collects data using voice and video sensors.
[0791] Step 4:
[0792] The server inputs emotional state and facility data into a generating AI model, which then uses prompts to generate personalized additional information. The input is the user's emotional state and facility data, and the output is customized guidance information. The generating AI model uses this information to create appropriate guidance content.
[0793] Step 5:
[0794] The server outputs the generated additional information as audio data and sends it to the terminal. The input is personalized information from the generating AI, and the output is audio guidance data played on the terminal. In this process, the server uses speech synthesis technology to convert the digital data into audio.
[0795] Step 6:
[0796] The user receives audio data through their device and listens to personalized guidance in real time. The input is audio data from the server, and the output is the audio information received by the user. This allows the user to receive optimal guidance tailored to their emotions and interests.
[0797] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0798] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0799] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0800] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0801] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0802] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0803] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0804] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0805] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0806] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0807] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0808] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0809] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0810] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0811] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0812] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0813] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0814] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0815] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0816] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0817] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0818] The following is further disclosed regarding the embodiments described above.
[0819] (Claim 1)
[0820] A means of acquiring highly accurate location information,
[0821] A means for providing voice guidance data related to a specific location based on the aforementioned location information,
[0822] A means of generating additional information in response to a user's request using a generative AI,
[0823] A means for outputting the aforementioned additional information as audio data,
[0824] A system that includes this.
[0825] (Claim 2)
[0826] The system according to claim 1, comprising means for tracking the user's movement in real time and updating location-based guidance data in a timely manner.
[0827] (Claim 3)
[0828] The system according to claim 1, further comprising means for customizing voice guidance data based on user preferences and past interactions.
[0829] "Example 1"
[0830] (Claim 1)
[0831] A high-precision location information acquisition device,
[0832] A device that provides voice guidance information related to a specific geographical area based on the aforementioned location information,
[0833] A device that generates additional information in response to user requests using a generative AI model,
[0834] A device that transmits the aforementioned additional information as audio information,
[0835] The process of receiving current location information from the device and obtaining voice guidance information,
[0836] A function that generates prompt sentences using a generative AI model and customizes the content of voice guidance,
[0837] A system that includes this.
[0838] (Claim 2)
[0839] The system according to claim 1, comprising a device that tracks the user's movements in real time and updates location-based guidance information in a timely manner.
[0840] (Claim 3)
[0841] The system according to claim 1, further comprising a device that personalizes voice guidance information based on the user's preferences and previous interactions.
[0842] "Application Example 1"
[0843] (Claim 1)
[0844] A means of acquiring highly accurate location information,
[0845] A means for providing voice guidance data related to a specific location based on the aforementioned location information,
[0846] A means of generating additional information in response to a user's request using a generative AI,
[0847] A means for outputting the aforementioned additional information as acoustic data,
[0848] A means of providing user-specific guidance based on information within commercial facilities,
[0849] A system that includes this.
[0850] (Claim 2)
[0851] The system according to claim 1, comprising means for tracking the user's movement in real time and updating location-based guidance data in a timely manner.
[0852] (Claim 3)
[0853] The system according to claim 1, further comprising means for customizing voice guidance data based on the user's preferences and past interactions.
[0854] "Example 2 of combining an emotion engine"
[0855] (Claim 1)
[0856] High-precision positioning means,
[0857] Means for providing voice guidance information related to a specific area based on the aforementioned positioning,
[0858] A means for generating additional information based on the user's emotional state using a generative AI model,
[0859] A means for outputting the aforementioned additional information as audio information,
[0860] A means of analyzing the emotional state of users,
[0861] A means of adjusting the tone and level of detail of voice guidance information according to the user's emotional state,
[0862] A system that includes this.
[0863] (Claim 2)
[0864] The system according to claim 1, comprising means for tracking user behavior in real time and immediately updating location-specific guidance information.
[0865] (Claim 3)
[0866] The system according to claim 1, further comprising means for adapting voice guidance information based on the user's preferences and past interactions, and means for further optimizing it according to the user's emotional state.
[0867] "Application example 2 when combining with an emotional engine"
[0868] (Claim 1)
[0869] A high-precision location information acquisition device,
[0870] A device that provides voice guidance information related to a specific facility based on the aforementioned location information,
[0871] A device that uses generative AI technology to generate additional information customized based on customer preferences and past interactions,
[0872] An apparatus that includes a process of analyzing a customer's emotional state using an emotion engine and personalizing the aforementioned additional information,
[0873] A device that outputs the personalized information as audio data,
[0874] A system that includes this.
[0875] (Claim 2)
[0876] The system according to claim 1, comprising a device that tracks customer movement in real time and dynamically updates voice guidance information according to location and emotional state.
[0877] (Claim 3)
[0878] The system according to claim 1, further comprising a device that generates detailed information about specific items or places using prompts generated by a generative AI and integrates it into voice guidance information based on customer interests. [Explanation of Symbols]
[0879] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of acquiring highly accurate location information, A means for providing voice guidance data related to a specific location based on the aforementioned location information, A means of generating additional information in response to a user's request using a generative AI, A means for outputting the aforementioned additional information as audio data, A system that includes this.
2. The system according to claim 1, comprising means for tracking the user's movement in real time and updating location-based guidance data in a timely manner.
3. The system according to claim 1, further comprising means for customizing voice guidance data based on user preferences and past interactions.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A