system
The system addresses the limitations of conventional tourist guides by analyzing user queries, searching local databases for specific information, and delivering responses in multiple formats to meet individual tourist needs and emotional states, enhancing the user experience.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-11
- Publication Date
- 2026-04-23
AI Technical Summary
Conventional tourist guides provide only general information and fail to meet the specific needs of tourists unfamiliar with an area, lacking a means to efficiently collect and deliver detailed, personalized information from local residents.
A system that analyzes user questions in natural language, searches a local information database for niche information, and provides responses in text, video, and audio formats tailored to individual needs and emotional states.
Enables the delivery of detailed, personalized tourist information that meets individual needs and emotional states, enhancing the user experience by providing visually and audibly rich content.
Smart Images

Figure 2026069133000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Conventional tourist guides are limited to general information and are insufficient for tourists who are not familiar with a specific area. Also, it has been difficult to provide detailed tourist information according to individual needs. Therefore, it is required to utilize the unique information possessed by local residents, but there is a problem that there is no means to efficiently collect and provide that information to tourists.
Means for Solving the Problems
[0005] This invention accurately grasps the individual needs of tourists by providing an analysis means that analyzes questions entered by users in natural language and understands their intent. Furthermore, by using a search means that searches a local information database containing niche information collected from local residents based on the analyzed intent, it provides detailed information that goes beyond general tourist information. In addition, since the generated response is provided as a combination of text information, video information, and audio information, it is possible to deliver information to tourists in a visually and aurally easy-to-understand format.
[0006] A "user" refers to an individual or group that accesses the system and asks questions using natural language.
[0007] "Input method" refers to the interface that users use to input questions in natural language.
[0008] "Analysis means" refers to a system component that has the function of analyzing questions entered by the user and understanding their intent.
[0009] "Search method" refers to a system component that has the function of retrieving related data from a database based on the analyzed information.
[0010] "Generation means" refers to a function that generates responses to the user in text, video, and audio formats based on search results.
[0011] "Transmission means" refers to a system component that has the function of delivering the generated response to the user's terminal.
[0012] "Local information" refers to detailed information that is not generally widely known, collected from residents who are familiar with a particular area.
[0013] "Response" refers to data provided in response to a user's question, which may include text, video, and audio formats. [Brief explanation of the drawing]
[0014] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Mode for Carrying Out the Invention
[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0018] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0019] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0022] [First Embodiment]
[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0035] This invention is a system in which a user inputs a question about tourism in natural language, the system analyzes the question, and generates and provides a detailed response based on local information. The processing of this system's program is described below in natural language.
[0036] The user begins by typing a question into the chat screen on their device. For example, they might type a question like, "What are some recommended hidden tourist spots in Kyoto?" The device then sends this question directly to the server.
[0037] The server analyzes the received question. Using natural language processing techniques, it identifies the intent of the question and information such as place names. Here, the intent "recommended hidden tourist spots" and the region "Kyoto" are extracted. Next, the server accesses the database and searches for the necessary local information. Utilizing information collected from local residents, it finds tourist spots that are not generally known.
[0038] When the server generates a response based on the information obtained from the search, it can provide information in a combination of text, video, and audio. For example, if it finds information about a small, popular local garden in Kyoto, it can provide photos and videos of the garden along with the text information. The generated response is then sent to the terminal.
[0039] The terminal provides the user with received information in a visually and audibly easy-to-understand format. The user can confirm the appeal of the tourist spots introduced on the screen and gain a deeper understanding through audio and video. In this form, the present invention provides detailed and valuable information that goes beyond general tourist information and is tailored to individual needs.
[0040] The following describes the processing flow.
[0041] Step 1:
[0042] The user accesses the chat interface on their device and enters a question about sightseeing in natural language. For example, they might enter the question, "What are some local recommendations for restaurants in Tokyo?"
[0043] Step 2:
[0044] The terminal sends the user's entered questions to the server. This transmission occurs in real time and is the first step in quickly processing the user's requests.
[0045] Step 3:
[0046] The server analyzes the received question using a natural language processing engine. This analysis identifies the main point and intent of the question and extracts important keywords and phrases. In this case, the keywords "recommended restaurants" and "Tokyo" are extracted.
[0047] Step 4:
[0048] The server searches a local information database based on the analyzed information. This database contains niche information collected from local residents, and the server narrows down the relevant information based on this. In this example, it retrieves information on a specific restaurant that is highly rated locally.
[0049] Step 5:
[0050] The server generates a response to provide to the user based on the search results. It combines text, image data, video links, and audio files to prepare information in a format that is easily understandable to the user. For example, it might prepare details and photos of a specific restaurant in Tokyo.
[0051] Step 6:
[0052] The server sends the generated response to the terminal. The transmitted data is displayed or played back in an appropriate format on the user interface.
[0053] Step 7:
[0054] The terminal displays responses received from the server, outputting text information to the screen as needed, while also playing video and audio. This allows users to obtain detailed local tourist information and use it to plan their trips.
[0055] (Example 1)
[0056] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0057] With the development of the information society, users are increasingly seeking detailed and personalized information specific to their region. However, traditional information systems face challenges in accurately understanding user intent and providing specific, region-specific information. Furthermore, there is a demand for information provision that appeals not only to text but also to visual and auditory senses, providing rich content. This necessitates improving the user experience.
[0058] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0059] In this invention, the server includes an interface means for the user to input a question in natural language; an analysis means for receiving the question input by the user and analyzing the question using generative AI technology to understand the intent and related information; an information retrieval means for searching for relevant geographic information from a data set based on the analyzed intent; and an information distribution means for providing the generated response to the user. This enables the user to obtain detailed and diverse content that is relevant to their intent and includes region-specific information.
[0060] "Interface means" refers to devices or software that allow users to input questions into a system using natural language.
[0061] "Generative AI technology" refers to a technology that uses artificial intelligence to analyze the intent of a user's question from natural language input and derive an appropriate response.
[0062] "Analysis means" refers to the process or method of analyzing received questions using artificial intelligence technology to identify intent and related information.
[0063] "Information retrieval methods" refer to the processes and technologies used to search for relevant geographic information from a data set based on analyzed intent.
[0064] "Information generation means" refers to functions and methods for generating responses in various formats (text, video, audio) based on retrieved geographic information.
[0065] "Information distribution means" refers to communication means and display devices used to deliver the generated response to the user.
[0066] This invention is a system in which users ask questions about tourism in natural language and obtain local information based on those questions. The specific techniques for implementing this system are described below.
[0067] The user first enters their question through an interface on their device. This interface is designed to allow users to easily enter questions in natural language. For example, they can enter a question such as, "What are some recommended hidden tourist spots in Kyoto?" An example of a prompt would be, "What are some recommended hidden tourist spots in Kyoto? I want to know about places that aren't well known to locals."
[0068] The terminal sends the entered question to the server. The server analyzes the question using generative AI technology, such as the GPT model. The analysis identifies the user's intent and relevant information such as place names. Based on this analysis, the server searches for relevant geographic information from the data set. The data set also includes unique information collected from local residents.
[0069] Based on the search results, the server uses information generation tools to create responses in text, video, and audio formats. These responses can combine diverse media to form rich content. For example, if the server finds information about a hidden garden in Kyoto, it will provide text descriptions of its history and highlights, along with photos and videos of the garden.
[0070] Finally, the generated response is provided to the user through the terminal's information distribution means. The user can visually and audibly understand the answer to the question through the displayed information. In this way, the present invention makes it possible to provide the user with detailed and specific tourist information they desire.
[0071] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0072] Step 1:
[0073] Users input travel-related questions in natural language using an interface on their device. For example, they can enter specific questions such as, "What are some recommended hidden tourist spots in Kyoto?" The device receives this input, formats the data, and sends it to the server via the network.
[0074] Step 2:
[0075] The server analyzes questions received from terminals using generation AI technology. In this analysis, the AI model processes text data to identify the intent of the question, place names, and areas of interest. It receives natural language question text as input and outputs the analysis results. These analysis results may include keywords such as "hidden tourist spots" and "Kyoto."
[0076] Step 3:
[0077] The server searches for geographical information within the data set based on the analyzed keywords. The data set contains unique information collected from local residents. The input is the keywords from the analysis results, and the output is information on related tourist spots. Information retrieval is performed through a high-speed database search algorithm.
[0078] Step 4:
[0079] The server generates a multimedia-based response based on the information obtained from the search. Search results are used as input, and the output is content combining text, images, and video. This generation process utilizes information generation methods to provide visually and aurally rich content.
[0080] Step 5:
[0081] The terminal receives multimedia responses sent from the server and provides them to the user. Its input is multimedia data from the server, and its output is the presentation of visual and auditory information to the user. Based on this information, the user can obtain detailed and personalized answers to their questions.
[0082] (Application Example 1)
[0083] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0084] Traditional tourist information systems only provide information based on user input, making it difficult to provide detailed, location-specific information or hidden gems in real time when a user visits a particular tourist destination. Therefore, there was a need for a method that could personalize the tourist experience and provide users with a deeper understanding.
[0085] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0086] In this invention, the server includes an interface means for the user to input a question in natural language, a natural language processing means for receiving the question from the user, analyzing the question, and understanding its intent, and a search means for retrieving relevant local information from a data store based on the analyzed intent and the user's current location information. This makes it possible to automatically provide detailed information and guidance on hidden attractions related to a place when the user visits a tourist destination.
[0087] An "interface means" is an input device or platform for users to input questions in natural language.
[0088] "Natural language processing means" refers to technologies and systems that analyze questions entered by users and understand their intent.
[0089] The "search method" is a function that searches for relevant regional information from the data store based on the analyzed intent and the user's current location information.
[0090] "Information processing means" refers to the process of generating responses in the form of text, images, and sound based on the searched regional information.
[0091] "Means of communication" refers to a method or system for transmitting a generated response to a user through a visual display or auditory device.
[0092] "Local information" refers to data and insights related to a specific geographical area, and in particular, information that includes hidden places and anecdotes collected from local residents.
[0093] "Visual information" refers to data related to videos and images presented to the user.
[0094] "Acoustic information" refers to audio and sound-related data presented to the user.
[0095] This invention provides a system for users to input tourism-related questions in natural language and obtain detailed information based on those inputs. The system has the following configuration:
[0096] Users input questions in natural language via devices such as smart glasses or mobile devices. For example, they might input a question like, "What are some recommended tourist spots nearby?" through the interface. In this process, visual displays and sound devices implemented on the device are used.
[0097] The server analyzes received questions using natural language processing to understand their intent. For example, common natural language processing libraries such as NLTK and Spacy can be used. The analyzed information is combined with the user's current location information to access a data store and search for relevant regional information. The search mechanism extracts useful data to present to the user.
[0098] The information processing system uses a generative AI model to generate responses to the user in text, image, and audio formats based on the explored local information. These generated responses are then communicated to the user through visual displays and auditory devices. For example, when a user visits a temple in Kyoto, the system can provide information about the temple's history, highlights, and nearby hidden gems.
[0099] The following prompt statements can be used with the generative AI model.
[0100] Example prompt: "What are some hidden gems in this area that everyone overlooks?"
[0101] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0102] Step 1:
[0103] The user enters the question in natural language through their device. The input is in text format and can be done using the voice input function or touch interface of smart glasses. The user confirms the question on the device and sends the entered information to the server by pressing the submit button.
[0104] Step 2:
[0105] The server analyzes received questions using natural language processing (NLTK) techniques. By analyzing the input question text, it extracts the intent and keywords of the question. For example, intents such as "tourist spots" and "recommendations" are extracted here. Natural language processing libraries such as NLTK and Spacy are used for the analysis.
[0106] Step 3:
[0107] The server searches for local information from the data store based on the analyzed intent and the user's current location. The inputs are the extracted intent and GPS information. The data store retrieves information on relevant tourist spots and region-specific information. A search algorithm is used to efficiently collect highly relevant information.
[0108] Step 4:
[0109] Using information processing tools, a generative AI model generates a response based on acquired regional information. The input is the explored regional information. A generative AI model such as GPT-3 (registered trademark) is used for text generation, creating a response that combines visual and auditory information. The generated data is then ready to be sent to the user's terminal.
[0110] Step 5:
[0111] The terminal presents the response received from the server to the user through a visual display and audio device. The input is the response data sent from the server. The user can obtain detailed information about tourist attractions through the information displayed on the screen and the audio guides provided.
[0112] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0113] This invention provides a sophisticated information system that can respond to the individual emotional needs of users by combining an emotion engine with an information provision system for tourist information. Below, the program processing of this system will be explained in natural language, along with specific examples.
[0114] The user first accesses the chat interface on their device and enters travel-related questions and requests in natural language. For example, a possible question might be, "Where are some hidden spots in Kyoto where I can relax in peace?" The device then sends this question directly to the server.
[0115] The server uses a natural language processing engine to analyze the received question. This analysis extracts the main point of the question and related keywords, and also uses an emotion engine to identify the user's emotions from their writing. For example, the phrase "I want to spend some time quietly" suggests that the user is seeking relaxation.
[0116] Based on the analyzed intent and emotional information, the server searches the database for the most relevant local information. Prioritizing information about quiet places frequently used by locals, the search might yield information about a quiet temple or garden in Kyoto.
[0117] The server generates responses based on search results, adjusting the content and format to reflect emotional information. For users seeking relaxation, information with calming images and music is recommended. For example, content might be created combining tranquil images of the relevant temple or garden with soothing music.
[0118] Finally, the server sends the generated response to the terminal. The terminal receives it and provides the information in a format that is easily understandable to the user. Through sight and hearing, the user can obtain detailed and valuable tourist information that addresses their specific emotional needs and can be used to plan their trip.
[0119] Thus, by providing information that takes user emotions into consideration, the present invention realizes a more personalized travel experience that goes beyond general tourist information.
[0120] The following describes the processing flow.
[0121] Step 1:
[0122] The user accesses the chat interface on their device and enters a question about sightseeing in natural language. For example, they might enter a question like, "Where are some quiet cafes in Hokkaido that locals recommend?"
[0123] Step 2:
[0124] The terminal sends the user's input directly to the server. This transmission process is asynchronous, ensuring that user input is delivered to the server quickly.
[0125] Step 3:
[0126] The server first analyzes the received question using a natural language processing engine. The analysis identifies the subject of the question and the type of information being sought, and extracts keywords. In this case, the keywords extracted are "recommended quiet cafes" and "Hokkaido."
[0127] Step 4:
[0128] After the server analyzes the subject of the question, it uses an emotion engine to identify the emotions underlying the user's words. For example, the phrase "quiet cafe" suggests that the user is seeking a quiet and calm atmosphere.
[0129] Step 5:
[0130] The server searches the database for relevant local information based on the analyzed keywords and sentiment information. Here, it selects information about cafes in Hokkaido with a quiet and relaxed atmosphere, and refers to ratings and reviews from local residents.
[0131] Step 6:
[0132] The server generates a response that reflects search results and sentiment information. If a user is looking for a "calm atmosphere," the server adjusts the text, photos, videos, and audio presentation to match that sentiment. Specifically, this might involve setting relaxing music in the background and providing images or videos of a "quiet and comfortable cafe."
[0133] Step 7:
[0134] The generated response is sent from the server to the terminal, where the user receives it. The terminal appropriately displays the transmitted information and, as a response to the user's input, provides information that meets their anticipated emotional needs. The user can then use this information to create a more fulfilling travel plan.
[0135] (Example 2)
[0136] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0137] Traditional tourist information systems only provide general information in response to user questions, making it difficult to address the emotional needs of individual users. This resulted in limited effectiveness of tourist information and a failure to adequately meet the experiences users desired.
[0138] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0139] In this invention, the server includes input means for the user to input a question in natural language, analysis means for receiving the question from the user and analyzing the question to understand its intent and emotion, and search means for retrieving relevant local information from a database based on the analyzed intent and emotion. This makes it possible to provide more personalized tourist information that meets the individual emotional needs of the user.
[0140] "Input means" refers to a method or device that enables a user to input a question in natural language.
[0141] "Analysis means" refers to a method or apparatus for analyzing questions received from a user and understanding their intent and emotions.
[0142] "Search means" refers to a method or apparatus for retrieving relevant local information from a database based on analyzed intentions and emotions.
[0143] "Generation means" refers to a method or apparatus for generating responses in text, video, and audio formats, adjusting the content and format of the response according to the user's emotions based on the retrieved regional information.
[0144] "Transmission means" refers to a method or apparatus for sending the generated response to the user.
[0145] "Local information" refers to detailed information about a specific area, and may include specialized information collected from local residents.
[0146] This invention is a system that provides personalized tourist information tailored to the emotional needs of users. The system consists of a terminal for users to input questions in natural language and a server that performs analysis and provides information.
[0147] Users enter questions or requests related to specific tourist destinations using the chat interface on their device. They are required to enter their questions in natural language. For example, if a user enters "Please tell me about quiet places to spend time in Kyoto," this will be recognized as a prompt by the device.
[0148] The terminal receives input from the user and sends that data to the server. The server uses a generative AI model and a natural language processing engine to analyze the content of the question and identify the user's intent and emotions. Emotion identification is performed using an emotion engine, which can identify what kind of emotional experience the user is seeking.
[0149] Based on identified intentions and emotions, the server searches the database for relevant local information. This local information includes specialized data provided by local residents, allowing the server to find information that meets the user's needs. For example, information about quiet temples and gardens in Kyoto might be retrieved from the server.
[0150] Furthermore, the server uses the generated AI model to create responses to the user in text, video, and audio formats based on the search results. During generation, it selects expressions and content that match the user's emotions and intentions, providing responses that combine calming video and music to users seeking relaxation.
[0151] Ultimately, the response generated by the server is sent to the terminal, and the user can view and listen to it to obtain valuable information that matches their emotional needs. This invention allows users to enjoy a tourism experience tailored to their individual emotions.
[0152] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0153] Step 1:
[0154] The user accesses the chat interface on their device and enters travel-related questions in natural language. For example, they might enter a prompt such as, "Please tell me about quiet places to stay in Kyoto." The device sends this input directly to the server. The input data is a natural language question, and the output to the server is this question data.
[0155] Step 2:
[0156] The server passes the received question to a natural language processing engine, which analyzes the gist of the question. Specifically, it extracts important keywords and identifies the information the user is seeking. This process takes a natural language question as input and outputs the analyzed intent and a set of keywords.
[0157] Step 3:
[0158] The server uses an emotion engine to identify the emotions expressed in the user's questions. For example, it senses that the user is seeking relaxation from the expression "I want to spend some time quietly." At this stage, the input is the analyzed question, and the output is emotion data.
[0159] Step 4:
[0160] The server uses the analyzed intent and sentiment data to search for appropriate local information from its information database. Here, data that particularly matches the user's emotional needs is prioritized. The input is a dataset of intents and sentiments, and the output provides relevant local information.
[0161] Step 5:
[0162] The server uses a generative AI model to generate responses to the user based on the searched local information. These responses are available in text, video, and audio formats, and their content and expression are adjusted according to the user's emotions. For example, they might include calming video and music. The input for this step is local information, and the output is emotionally responsive content.
[0163] Step 6:
[0164] The server sends the generated response to the terminal. The terminal displays the response in a user-friendly format. The input is the response content from the server, and the output is visual and auditory information provided to the user. Based on this, the user can then create a concrete travel plan.
[0165] (Application Example 2)
[0166] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0167] In modern society, users face the challenge of quickly and accurately obtaining information that matches their emotions and needs from a vast amount of data. Especially in virtual stores, there is a need to provide a more personalized experience by offering product suggestions based on the user's emotions.
[0168] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0169] In this invention, the server includes an input device for the user to input a question in natural language, an analysis device for analyzing the received question to understand intent and emotional information, and a search device for selecting the optimal data based on that information. This makes it possible to provide accurate information and product suggestions that respond to the user's emotions.
[0170] A "user" is a person who inputs information and receives services based on that information.
[0171] "Natural language" refers to the method of expressing information using the language that humans use on a daily basis.
[0172] An "input device" refers to a device used by a user to provide information to a system.
[0173] An "analysis device" is a device that analyzes input information and understands its intent and associated emotions.
[0174] "Emotional information" refers to data related to emotions extracted from information entered by the user.
[0175] A "search device" is a device used to select specific data based on analyzed information.
[0176] A "generation device" is a device that forms the optimal response for the user based on selected data.
[0177] A "display device" is a device that outputs a generated response visually or audibly.
[0178] This invention is a system that provides appropriate information based on the user's emotional state. This system allows the user to input information via a smart device and generate an optimal response based on that information.
[0179] The user first inputs a question in natural language through an input device such as smart glasses. This input can be done either by voice using the device's microphone or by text. This information is sent to a server. The server analyzes the user's question using natural language processing software (e.g., Google® Cloud Natural Language API). In this analysis process, a sentiment analysis engine (e.g., IBM Watson® Tone Analyzer) is used to identify the user's emotions in addition to the purpose of the question.
[0180] Based on the analyzed intent and emotional information, the server searches the database for relevant items. This narrows down the products and information to match the user's needs. The generated response is visually displayed on the smart glasses' screen. This display includes images and descriptions of relaxing products for users seeking relaxation.
[0181] For example, if a user enters "I feel like relaxing today," the system will visually suggest relaxing aromatherapy candles or houseplants. Another example of a prompt would be, "How can I provide personalized product suggestions based on the user's emotions?"
[0182] This approach allows users to instantly find the products and information that best suit their emotional state, resulting in a personalized shopping experience.
[0183] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0184] Step 1:
[0185] The user inputs a question in natural language through an input device such as smart glasses. The input language data is converted into text data through the device's speech recognition system. For example, if the voice input is "I feel like relaxing today," it will be converted into text format.
[0186] Step 2:
[0187] The terminal sends the converted text data to the server. The server receives this data and analyzes the sentence using a natural language processing engine (e.g., Google Cloud Natural Language API). This analysis infers the user's intent and extracts emotional information from the statement using an emotion engine. The output of the analysis includes the keyword "relax" and the emotional information "I want to relax."
[0188] Step 3:
[0189] Based on the analysis results, the server searches the database for relevant products and information. At this time, it filters the data for products with relaxing effects based on the emotional information "relaxation." For example, aromatherapy candles and houseplants might be selected as relevant products.
[0190] Step 4:
[0191] The server generates a visually appealing response based on the selected product information. The generated response consists of a format that includes text, images, and audio guidance. This includes images of the selected product, text describing its effects, and, in some cases, audio guidance.
[0192] Step 5:
[0193] The server sends the generated response to the terminal. The terminal receives this and displays the content of the response to the user in real time on the smart glasses' display. The user can visually check images and descriptions of products related to "relaxation" and decide on the next action.
[0194] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0195] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0196] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0197] [Second Embodiment]
[0198] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0199] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0200] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0201] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0202] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0203] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0204] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0205] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0206] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0207] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0208] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0209] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0210] This invention is a system in which a user inputs a question about tourism in natural language, the system analyzes the question, and generates and provides a detailed response based on local information. The processing of this system's program is described below in natural language.
[0211] The user begins by typing a question into the chat screen on their device. For example, they might type a question like, "What are some recommended hidden tourist spots in Kyoto?" The device then sends this question directly to the server.
[0212] The server analyzes the received question. Using natural language processing techniques, it identifies the intent of the question and information such as place names. Here, the intent "recommended hidden tourist spots" and the region "Kyoto" are extracted. Next, the server accesses the database and searches for the necessary local information. Utilizing information collected from local residents, it finds tourist spots that are not generally known.
[0213] When the server generates a response based on the information obtained from the search, it can provide information in a combination of text, video, and audio. For example, if it finds information about a small, popular local garden in Kyoto, it can provide photos and videos of the garden along with the text information. The generated response is then sent to the terminal.
[0214] The terminal provides the user with received information in a visually and audibly easy-to-understand format. The user can confirm the appeal of the tourist spots introduced on the screen and gain a deeper understanding through audio and video. In this form, the present invention provides detailed and valuable information that goes beyond general tourist information and is tailored to individual needs.
[0215] The following describes the processing flow.
[0216] Step 1:
[0217] The user accesses the chat interface on their device and enters a question about sightseeing in natural language. For example, they might enter the question, "What are some local recommendations for restaurants in Tokyo?"
[0218] Step 2:
[0219] The terminal sends the user's entered questions to the server. This transmission occurs in real time and is the first step in quickly processing the user's requests.
[0220] Step 3:
[0221] The server analyzes the received question using a natural language processing engine. This analysis identifies the main point and intent of the question and extracts important keywords and phrases. In this case, the keywords "recommended restaurants" and "Tokyo" are extracted.
[0222] Step 4:
[0223] The server searches a local information database based on the analyzed information. This database contains niche information collected from local residents, and the server narrows down the relevant information based on this. In this example, it retrieves information on a specific restaurant that is highly rated locally.
[0224] Step 5:
[0225] The server generates a response to provide to the user based on the search results. It combines text, image data, video links, and audio files to prepare information in a format that is easily understandable to the user. For example, it might prepare details and photos of a specific restaurant in Tokyo.
[0226] Step 6:
[0227] The server sends the generated response to the terminal. The transmitted data is displayed or played back in an appropriate format on the user interface.
[0228] Step 7:
[0229] The terminal displays responses received from the server, outputting text information to the screen as needed, while also playing video and audio. This allows users to obtain detailed local tourist information and use it to plan their trips.
[0230] (Example 1)
[0231] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0232] With the development of the information society, users are increasingly seeking detailed and personalized information specific to their region. However, traditional information systems face challenges in accurately understanding user intent and providing specific, region-specific information. Furthermore, there is a demand for information provision that appeals not only to text but also to visual and auditory senses, providing rich content. This necessitates improving the user experience.
[0233] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0234] In this invention, the server includes an interface means for the user to input a question in natural language; an analysis means for receiving the question input by the user and analyzing the question using generative AI technology to understand the intent and related information; an information retrieval means for searching for relevant geographic information from a data set based on the analyzed intent; and an information distribution means for providing the generated response to the user. This enables the user to obtain detailed and diverse content that is relevant to their intent and includes region-specific information.
[0235] "Interface means" refers to devices or software that allow users to input questions into a system using natural language.
[0236] "Generative AI technology" refers to a technology that uses artificial intelligence to analyze the intent of a user's question from natural language input and derive an appropriate response.
[0237] "Analysis means" refers to the process or method of analyzing received questions using artificial intelligence technology to identify intent and related information.
[0238] "Information retrieval methods" refer to the processes and technologies used to search for relevant geographic information from a data set based on analyzed intent.
[0239] "Information generation means" refers to functions and methods for generating responses in various formats (text, video, audio) based on retrieved geographic information.
[0240] "Information distribution means" refers to communication means and display devices used to deliver the generated response to the user.
[0241] This invention is a system in which users ask questions about tourism in natural language and obtain local information based on those questions. The specific techniques for implementing this system are described below.
[0242] The user first enters their question through an interface on their device. This interface is designed to allow users to easily enter questions in natural language. For example, they can enter a question such as, "What are some recommended hidden tourist spots in Kyoto?" An example of a prompt would be, "What are some recommended hidden tourist spots in Kyoto? I want to know about places that aren't well known to locals."
[0243] The terminal sends the entered question to the server. The server analyzes the question using generative AI technology, such as the GPT model. The analysis identifies the user's intent and relevant information such as place names. Based on this analysis, the server searches for relevant geographic information from the data set. The data set also includes unique information collected from local residents.
[0244] Based on the search results, the server uses information generation tools to create responses in text, video, and audio formats. These responses can combine diverse media to form rich content. For example, if the server finds information about a hidden garden in Kyoto, it will provide text descriptions of its history and highlights, along with photos and videos of the garden.
[0245] Finally, the generated response is provided to the user through the terminal's information distribution means. The user can visually and audibly understand the answer to the question through the displayed information. In this way, the present invention makes it possible to provide the user with detailed and specific tourist information they desire.
[0246] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0247] Step 1:
[0248] Users input travel-related questions in natural language using an interface on their device. For example, they can enter specific questions such as, "What are some recommended hidden tourist spots in Kyoto?" The device receives this input, formats the data, and sends it to the server via the network.
[0249] Step 2:
[0250] The server analyzes questions received from terminals using generation AI technology. In this analysis, the AI model processes text data to identify the intent of the question, place names, and areas of interest. It receives natural language question text as input and outputs the analysis results. These analysis results may include keywords such as "hidden tourist spots" and "Kyoto."
[0251] Step 3:
[0252] The server searches for geographical information within the data set based on the analyzed keywords. The data set contains unique information collected from local residents. The input is the keywords from the analysis results, and the output is information on related tourist spots. Information retrieval is performed through a high-speed database search algorithm.
[0253] Step 4:
[0254] The server generates a multimedia-based response based on the information obtained from the search. Search results are used as input, and the output is content combining text, images, and video. This generation process utilizes information generation methods to provide visually and aurally rich content.
[0255] Step 5:
[0256] The terminal receives multimedia responses sent from the server and provides them to the user. Its input is multimedia data from the server, and its output is the presentation of visual and auditory information to the user. Based on this information, the user can obtain detailed and personalized answers to their questions.
[0257] (Application Example 1)
[0258] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0259] Traditional tourist information systems only provide information based on user input, making it difficult to provide detailed, location-specific information or hidden gems in real time when a user visits a particular tourist destination. Therefore, there was a need for a method that could personalize the tourist experience and provide users with a deeper understanding.
[0260] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0261] In this invention, the server includes an interface means for the user to input a question in natural language, a natural language processing means for receiving the question from the user, analyzing the question, and understanding its intent, and a search means for retrieving relevant local information from a data store based on the analyzed intent and the user's current location information. This makes it possible to automatically provide detailed information and guidance on hidden attractions related to a place when the user visits a tourist destination.
[0262] An "interface means" is an input device or platform for users to input questions in natural language.
[0263] "Natural language processing means" refers to technologies and systems that analyze questions entered by users and understand their intent.
[0264] The "search method" is a function that searches for relevant regional information from the data store based on the analyzed intent and the user's current location information.
[0265] "Information processing means" refers to the process of generating responses in the form of text, images, and sound based on the searched regional information.
[0266] "Means of communication" refers to a method or system for transmitting a generated response to a user through a visual display or auditory device.
[0267] "Local information" refers to data and insights related to a specific geographical area, and in particular, information that includes hidden places and anecdotes collected from local residents.
[0268] "Visual information" refers to data related to videos and images presented to the user.
[0269] "Acoustic information" refers to audio and sound-related data presented to the user.
[0270] This invention provides a system for users to input tourism-related questions in natural language and obtain detailed information based on those inputs. The system has the following configuration:
[0271] Users input questions in natural language via devices such as smart glasses or mobile devices. For example, they might input a question like, "What are some recommended tourist spots nearby?" through the interface. In this process, visual displays and sound devices implemented on the device are used.
[0272] The server analyzes received questions using natural language processing to understand their intent. For example, common natural language processing libraries such as NLTK and Spacy can be used. The analyzed information is combined with the user's current location information to access a data store and search for relevant regional information. The search mechanism extracts useful data to present to the user.
[0273] The information processing system uses a generative AI model to generate responses to the user in text, image, and audio formats based on the explored local information. These generated responses are then communicated to the user through visual displays and auditory devices. For example, when a user visits a temple in Kyoto, the system can provide information about the temple's history, highlights, and nearby hidden gems.
[0274] The following prompt statements can be used with the generative AI model.
[0275] Example prompt: "What are some hidden gems in this area that everyone overlooks?"
[0276] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0277] Step 1:
[0278] The user inputs a question in natural language through the terminal. The input is in text form, and the voice input function or touch interface of the smart glasses can be used. The user confirms the question on the terminal and presses the send button, and the input information is sent to the server.
[0279] Step 2:
[0280] The server analyzes the received question using natural language processing means. By analyzing the input question text, the intention and keywords of the question are extracted. For example, intentions such as "tourist attractions" and "recommendations" are extracted here. Natural language processing libraries such as NLTK and Spacy are used for the analysis.
[0281] Step 3:
[0282] The server searches for regional information from the data store based on the analyzed intention and the user's current location information. The input is the extracted intention and GPS information. Information on related tourist attractions and region-specific information is obtained from the data store. A search algorithm is used to efficiently collect highly relevant information.
[0283] Step 4:
[0284] Using information processing means, a response is generated with an AI model based on the obtained regional information. The input is the searched regional information. For text generation, a generative AI model such as GPT-3 is used to create a response that combines visual information and audio information. The generated data is ready to be sent to the user's terminal.
[0285] Step 5:
[0286] The terminal presents the response received from the server to the user through a visual display and an acoustic device. The input is the response data sent from the server. The user can obtain detailed information about tourist attractions through the information displayed on the display and the guidance provided by voice.
[0287] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform specific processing using the user's emotion.
[0288] The present invention is a system that provides advanced information that can also meet the individual emotional needs of users by combining an emotion engine with an information providing system for tourist guidance. Below, the processing of the program of this system will be described in natural language and explained with specific examples.
[0289] The user first accesses the chat interface of the terminal and inputs questions and requests regarding travel in natural language. For example, a question such as "Where are the hidden spots where I can spend a quiet time in Kyoto?" can be considered. The terminal transmits this question as it is to the server.
[0290] The server uses a natural language processing engine to analyze the received question. Through this analysis, not only the main idea of the question and related keywords are extracted, but also the emotion is identified from the user's text by the emotion engine. For example, from the phrase "want to spend a quiet time", it is perceived that the user is seeking relaxation.
[0291] Based on the analyzed intention and emotion information, the server searches the database for the optimal local information. Here, information on quiet places frequently used by local residents is preferentially selected. For example, information on a certain quiet temple or garden in Kyoto can be obtained through the search.
[0292] The server generates a response based on the search results, and at this time, adjusts the content and format by reflecting the emotion information. For users seeking relaxation, information with videos and music that can convey a sense of tranquility is recommended. For example, content that includes quiet videos and calming music of the corresponding temple or garden is created.
[0293] Finally, the server sends the generated response to the terminal. The terminal receives it and provides the information in a format that is easily understandable to the user. Through sight and hearing, the user can obtain detailed and valuable tourist information that addresses their specific emotional needs and can be used to plan their trip.
[0294] Thus, by providing information that takes user emotions into consideration, the present invention realizes a more personalized travel experience that goes beyond general tourist information.
[0295] The following describes the processing flow.
[0296] Step 1:
[0297] The user accesses the chat interface on their device and enters a question about sightseeing in natural language. For example, they might enter a question like, "Where are some quiet cafes in Hokkaido that locals recommend?"
[0298] Step 2:
[0299] The terminal sends the user's input directly to the server. This transmission process is asynchronous, ensuring that user input is delivered to the server quickly.
[0300] Step 3:
[0301] The server first analyzes the received question using a natural language processing engine. The analysis identifies the subject of the question and the type of information being sought, and extracts keywords. In this case, the keywords extracted are "recommended quiet cafes" and "Hokkaido."
[0302] Step 4:
[0303] After the server analyzes the subject of the question, it uses an emotion engine to identify the emotions underlying the user's words. For example, the phrase "quiet cafe" suggests that the user is seeking a quiet and calm atmosphere.
[0304] Step 5:
[0305] Based on the analyzed keywords and sentiment information, the server searches the database for relevant local information. Here, information about cafes in Hokkaido with a quiet and calming atmosphere is selected, and evaluations and reviews from local residents are referred to.
[0306] Step 6:
[0307] The server generates a response that reflects the search results and sentiment information. When the user requests a "calming atmosphere", the presentation by text, photo, video, and audio is adjusted to match that sentiment. Specifically, adjustments are made such as using slow music in the background and preparing images or videos of "quiet and comfortable cafes".
[0308] Step 7:
[0309] The generated response is sent from the server to the terminal, and the user receives it on the terminal. The terminal appropriately displays the transmitted information and provides information that meets the expected emotional needs as a response to the user's input. The user can use this information to create a more fulfilling travel plan.
[0310] (Example 2)
[0311] Next, Example 2 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0312] Conventional tourism guidance systems only provide general information in response to questions from users, and it has been difficult to address the emotional needs of individual users. For this reason, there has been a problem that the effect of tourism guidance is limited and it cannot sufficiently respond to the experiences that users seek.
[0313] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0314] In this invention, the server includes input means for the user to input a question in natural language, analysis means for receiving the question from the user and analyzing the question to understand its intent and emotion, and search means for retrieving relevant local information from a database based on the analyzed intent and emotion. This makes it possible to provide more personalized tourist information that meets the individual emotional needs of the user.
[0315] "Input means" refers to a method or device that enables a user to input a question in natural language.
[0316] "Analysis means" refers to a method or apparatus for analyzing questions received from a user and understanding their intent and emotions.
[0317] "Search means" refers to a method or apparatus for retrieving relevant local information from a database based on analyzed intentions and emotions.
[0318] "Generation means" refers to a method or apparatus for generating responses in text, video, and audio formats, adjusting the content and format of the response according to the user's emotions based on the retrieved regional information.
[0319] "Transmission means" refers to a method or apparatus for sending the generated response to the user.
[0320] "Local information" refers to detailed information about a specific area, and may include specialized information collected from local residents.
[0321] This invention is a system that provides personalized tourist information tailored to the emotional needs of users. The system consists of a terminal for users to input questions in natural language and a server that performs analysis and provides information.
[0322] Users enter questions or requests related to specific tourist destinations using the chat interface on their device. They are required to enter their questions in natural language. For example, if a user enters "Please tell me about quiet places to spend time in Kyoto," this will be recognized as a prompt by the device.
[0323] The terminal receives input from the user and sends that data to the server. The server uses a generative AI model and a natural language processing engine to analyze the content of the question and identify the user's intent and emotions. Emotion identification is performed using an emotion engine, which can identify what kind of emotional experience the user is seeking.
[0324] Based on identified intentions and emotions, the server searches the database for relevant local information. This local information includes specialized data provided by local residents, allowing the server to find information that meets the user's needs. For example, information about quiet temples and gardens in Kyoto might be retrieved from the server.
[0325] Furthermore, the server uses the generated AI model to create responses to the user in text, video, and audio formats based on the search results. During generation, it selects expressions and content that match the user's emotions and intentions, providing responses that combine calming video and music to users seeking relaxation.
[0326] Ultimately, the response generated by the server is sent to the terminal, and the user can view and listen to it to obtain valuable information that matches their emotional needs. This invention allows users to enjoy a tourism experience tailored to their individual emotions.
[0327] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0328] Step 1:
[0329] The user accesses the chat interface on their device and enters travel-related questions in natural language. For example, they might enter a prompt such as, "Please tell me about quiet places to stay in Kyoto." The device sends this input directly to the server. The input data is a natural language question, and the output to the server is this question data.
[0330] Step 2:
[0331] The server passes the received question to a natural language processing engine, which analyzes the gist of the question. Specifically, it extracts important keywords and identifies the information the user is seeking. This process takes a natural language question as input and outputs the analyzed intent and a set of keywords.
[0332] Step 3:
[0333] The server uses an emotion engine to identify the emotions expressed in the user's questions. For example, it senses that the user is seeking relaxation from the expression "I want to spend some time quietly." At this stage, the input is the analyzed question, and the output is emotion data.
[0334] Step 4:
[0335] The server uses the analyzed intent and sentiment data to search for appropriate local information from its information database. Here, data that particularly matches the user's emotional needs is prioritized. The input is a dataset of intents and sentiments, and the output provides relevant local information.
[0336] Step 5:
[0337] The server uses a generative AI model to generate responses to the user based on the searched local information. These responses are available in text, video, and audio formats, and their content and expression are adjusted according to the user's emotions. For example, they might include calming video and music. The input for this step is local information, and the output is emotionally responsive content.
[0338] Step 6:
[0339] The server sends the generated response to the terminal. The terminal displays the response in a user-friendly format. The input is the response content from the server, and the output is visual and auditory information provided to the user. Based on this, the user can then create a concrete travel plan.
[0340] (Application Example 2)
[0341] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".
[0342] In modern society, users face the challenge of quickly and accurately obtaining information that matches their emotions and needs from a vast amount of data. Especially in virtual stores, there is a need to provide a more personalized experience by offering product suggestions based on the user's emotions.
[0343] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0344] In this invention, the server includes an input device for the user to input a question in natural language, an analysis device for analyzing the received question to understand intent and emotional information, and a search device for selecting the optimal data based on that information. This makes it possible to provide accurate information and product suggestions that respond to the user's emotions.
[0345] A "user" is a person who inputs information and receives services based on that information.
[0346] "Natural language" refers to the method of expressing information using the language that humans use on a daily basis.
[0347] An "input device" refers to a device used by a user to provide information to a system.
[0348] An "analysis device" is a device that analyzes input information and understands its intent and associated emotions.
[0349] "Emotional information" refers to data related to emotions extracted from information entered by the user.
[0350] A "search device" is a device used to select specific data based on analyzed information.
[0351] A "generation device" is a device that forms the optimal response for the user based on selected data.
[0352] A "display device" is a device that outputs a generated response visually or audibly.
[0353] This invention is a system that provides appropriate information based on the user's emotional state. This system allows the user to input information via a smart device and generate an optimal response based on that information.
[0354] The user first inputs a question in natural language through an input device such as smart glasses. This input can be done either by voice using the device's microphone or by text. This information is sent to a server. The server analyzes the user's question using natural language processing software (e.g., Google Cloud Natural Language API). In this analysis process, a sentiment analysis engine (e.g., IBM Watson Tone Analyzer) is used to identify the user's emotions in addition to the purpose of the question.
[0355] Based on the analyzed intent and emotional information, the server searches the database for relevant items. This narrows down the products and information to match the user's needs. The generated response is visually displayed on the smart glasses' screen. This display includes images and descriptions of relaxing products for users seeking relaxation.
[0356] For example, if a user enters "I feel like relaxing today," the system will visually suggest relaxing aromatherapy candles or houseplants. Another example of a prompt would be, "How can I provide personalized product suggestions based on the user's emotions?"
[0357] This approach allows users to instantly find the products and information that best suit their emotional state, resulting in a personalized shopping experience.
[0358] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0359] Step 1:
[0360] The user inputs a question in natural language through an input device such as smart glasses. The input language data is converted into text data through the device's speech recognition system. For example, if the voice input is "I feel like relaxing today," it will be converted into text format.
[0361] Step 2:
[0362] The terminal sends the converted text data to the server. The server receives this data and analyzes the sentence using a natural language processing engine (e.g., Google Cloud Natural Language API). This analysis infers the user's intent and extracts emotional information from the statement using an emotion engine. The output of the analysis includes the keyword "relax" and the emotional information "I want to relax."
[0363] Step 3:
[0364] Based on the analysis results, the server searches the database for relevant products and information. At this time, it filters the data for products with relaxing effects based on the emotional information "relaxation." For example, aromatherapy candles and houseplants might be selected as relevant products.
[0365] Step 4:
[0366] The server generates a visually appealing response based on the selected product information. The generated response consists of a format that includes text, images, and audio guidance. This includes images of the selected product, text describing its effects, and, in some cases, audio guidance.
[0367] Step 5:
[0368] The server sends the generated response to the terminal. The terminal receives this and displays the content of the response to the user in real time on the smart glasses' display. The user can visually check images and descriptions of products related to "relaxation" and decide on the next action.
[0369] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0370] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0371] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0372] [Third Embodiment]
[0373] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0374] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0375] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0376] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0377] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0378] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0379] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0380] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0381] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0382] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0383] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0384] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0385] This invention is a system in which a user inputs a question about tourism in natural language, the system analyzes the question, and generates and provides a detailed response based on local information. The processing of this system's program is described below in natural language.
[0386] The user begins by typing a question into the chat screen on their device. For example, they might type a question like, "What are some recommended hidden tourist spots in Kyoto?" The device then sends this question directly to the server.
[0387] The server analyzes the received question. Using natural language processing techniques, it identifies the intent of the question and information such as place names. Here, the intent "recommended hidden tourist spots" and the region "Kyoto" are extracted. Next, the server accesses the database and searches for the necessary local information. Utilizing information collected from local residents, it finds tourist spots that are not generally known.
[0388] When the server generates a response based on the information obtained from the search, it can provide information in a combination of text, video, and audio. For example, if it finds information about a small, popular local garden in Kyoto, it can provide photos and videos of the garden along with the text information. The generated response is then sent to the terminal.
[0389] The terminal provides the user with received information in a visually and audibly easy-to-understand format. The user can confirm the appeal of the tourist spots introduced on the screen and gain a deeper understanding through audio and video. In this form, the present invention provides detailed and valuable information that goes beyond general tourist information and is tailored to individual needs.
[0390] The following describes the processing flow.
[0391] Step 1:
[0392] The user accesses the chat interface on their device and enters a question about sightseeing in natural language. For example, they might enter the question, "What are some local recommendations for restaurants in Tokyo?"
[0393] Step 2:
[0394] The terminal sends the user's entered questions to the server. This transmission occurs in real time and is the first step in quickly processing the user's requests.
[0395] Step 3:
[0396] The server analyzes the received question using a natural language processing engine. This analysis identifies the main point and intent of the question and extracts important keywords and phrases. In this case, the keywords "recommended restaurants" and "Tokyo" are extracted.
[0397] Step 4:
[0398] The server searches a local information database based on the analyzed information. This database contains niche information collected from local residents, and the server narrows down the relevant information based on this. In this example, it retrieves information on a specific restaurant that is highly rated locally.
[0399] Step 5:
[0400] The server generates a response to provide to the user based on the search results. It combines text, image data, video links, and audio files to prepare information in a format that is easily understandable to the user. For example, it might prepare details and photos of a specific restaurant in Tokyo.
[0401] Step 6:
[0402] The server sends the generated response to the terminal. The transmitted data is displayed or played back in an appropriate format on the user interface.
[0403] Step 7:
[0404] The terminal displays responses received from the server, outputting text information to the screen as needed, while also playing video and audio. This allows users to obtain detailed local tourist information and use it to plan their trips.
[0405] (Example 1)
[0406] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0407] With the development of the information society, users are increasingly seeking detailed and personalized information specific to their region. However, traditional information systems face challenges in accurately understanding user intent and providing specific, region-specific information. Furthermore, there is a demand for information provision that appeals not only to text but also to visual and auditory senses, providing rich content. This necessitates improving the user experience.
[0408] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0409] In this invention, the server includes an interface means for the user to input a question in natural language; an analysis means for receiving the question input by the user and analyzing the question using generative AI technology to understand the intent and related information; an information retrieval means for searching for relevant geographic information from a data set based on the analyzed intent; and an information distribution means for providing the generated response to the user. This enables the user to obtain detailed and diverse content that is relevant to their intent and includes region-specific information.
[0410] "Interface means" refers to devices or software that allow users to input questions into a system using natural language.
[0411] "Generative AI technology" refers to a technology that uses artificial intelligence to analyze the intent of a user's question from natural language input and derive an appropriate response.
[0412] "Analysis means" refers to the process or method of analyzing received questions using artificial intelligence technology to identify intent and related information.
[0413] "Information retrieval methods" refer to the processes and technologies used to search for relevant geographic information from a data set based on analyzed intent.
[0414] "Information generation means" refers to functions and methods for generating responses in various formats (text, video, audio) based on retrieved geographic information.
[0415] "Information distribution means" refers to communication means and display devices used to deliver the generated response to the user.
[0416] This invention is a system in which users ask questions about tourism in natural language and obtain local information based on those questions. The specific techniques for implementing this system are described below.
[0417] The user first enters their question through an interface on their device. This interface is designed to allow users to easily enter questions in natural language. For example, they can enter a question such as, "What are some recommended hidden tourist spots in Kyoto?" An example of a prompt would be, "What are some recommended hidden tourist spots in Kyoto? I want to know about places that aren't well known to locals."
[0418] The terminal sends the entered question to the server. The server analyzes the question using generative AI technology, such as the GPT model. The analysis identifies the user's intent and relevant information such as place names. Based on this analysis, the server searches for relevant geographic information from the data set. The data set also includes unique information collected from local residents.
[0419] Based on the search results, the server uses information generation tools to create responses in text, video, and audio formats. These responses can combine diverse media to form rich content. For example, if the server finds information about a hidden garden in Kyoto, it will provide text descriptions of its history and highlights, along with photos and videos of the garden.
[0420] Finally, the generated response is provided to the user through the terminal's information distribution means. The user can visually and audibly understand the answer to the question through the displayed information. In this way, the present invention makes it possible to provide the user with detailed and specific tourist information they desire.
[0421] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0422] Step 1:
[0423] Users input travel-related questions in natural language using an interface on their device. For example, they can enter specific questions such as, "What are some recommended hidden tourist spots in Kyoto?" The device receives this input, formats the data, and sends it to the server via the network.
[0424] Step 2:
[0425] The server analyzes questions received from terminals using generation AI technology. In this analysis, the AI model processes text data to identify the intent of the question, place names, and areas of interest. It receives natural language question text as input and outputs the analysis results. These analysis results may include keywords such as "hidden tourist spots" and "Kyoto."
[0426] Step 3:
[0427] The server searches for geographical information within the data set based on the analyzed keywords. The data set contains unique information collected from local residents. The input is the keywords from the analysis results, and the output is information on related tourist spots. Information retrieval is performed through a high-speed database search algorithm.
[0428] Step 4:
[0429] The server generates a multimedia-based response based on the information obtained from the search. Search results are used as input, and the output is content combining text, images, and video. This generation process utilizes information generation methods to provide visually and aurally rich content.
[0430] Step 5:
[0431] The terminal receives multimedia responses sent from the server and provides them to the user. Its input is multimedia data from the server, and its output is the presentation of visual and auditory information to the user. Based on this information, the user can obtain detailed and personalized answers to their questions.
[0432] (Application Example 1)
[0433] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0434] Traditional tourist information systems only provide information based on user input, making it difficult to provide detailed, location-specific information or hidden gems in real time when a user visits a particular tourist destination. Therefore, there was a need for a method that could personalize the tourist experience and provide users with a deeper understanding.
[0435] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0436] In this invention, the server includes an interface means for the user to input a question in natural language, a natural language processing means for receiving the question from the user, analyzing the question, and understanding its intent, and a search means for retrieving relevant local information from a data store based on the analyzed intent and the user's current location information. This makes it possible to automatically provide detailed information and guidance on hidden attractions related to a place when the user visits a tourist destination.
[0437] An "interface means" is an input device or platform for users to input questions in natural language.
[0438] "Natural language processing means" refers to technologies and systems that analyze questions entered by users and understand their intent.
[0439] The "search method" is a function that searches for relevant regional information from the data store based on the analyzed intent and the user's current location information.
[0440] "Information processing means" refers to the process of generating responses in the form of text, images, and sound based on the searched regional information.
[0441] "Means of communication" refers to a method or system for transmitting a generated response to a user through a visual display or auditory device.
[0442] "Local information" refers to data and insights related to a specific geographical area, and in particular, information that includes hidden places and anecdotes collected from local residents.
[0443] "Visual information" refers to data related to videos and images presented to the user.
[0444] "Acoustic information" refers to audio and sound-related data presented to the user.
[0445] This invention provides a system for users to input tourism-related questions in natural language and obtain detailed information based on those inputs. The system has the following configuration:
[0446] Users input questions in natural language via devices such as smart glasses or mobile devices. For example, they might input a question like, "What are some recommended tourist spots nearby?" through the interface. In this process, visual displays and sound devices implemented on the device are used.
[0447] The server analyzes received questions using natural language processing to understand their intent. For example, common natural language processing libraries such as NLTK and Spacy can be used. The analyzed information is combined with the user's current location information to access a data store and search for relevant regional information. The search mechanism extracts useful data to present to the user.
[0448] The information processing system uses a generative AI model to generate responses to the user in text, image, and audio formats based on the explored local information. These generated responses are then communicated to the user through visual displays and auditory devices. For example, when a user visits a temple in Kyoto, the system can provide information about the temple's history, highlights, and nearby hidden gems.
[0449] The following prompt statements can be used with the generative AI model.
[0450] Example prompt: "What are some hidden gems in this area that everyone overlooks?"
[0451] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0452] Step 1:
[0453] The user enters the question in natural language through their device. The input is in text format and can be done using the voice input function or touch interface of smart glasses. The user confirms the question on the device and sends the entered information to the server by pressing the submit button.
[0454] Step 2:
[0455] The server analyzes received questions using natural language processing (NLTK) techniques. By analyzing the input question text, it extracts the intent and keywords of the question. For example, intents such as "tourist spots" and "recommendations" are extracted here. Natural language processing libraries such as NLTK and Spacy are used for the analysis.
[0456] Step 3:
[0457] The server searches for local information from the data store based on the analyzed intent and the user's current location. The inputs are the extracted intent and GPS information. The data store retrieves information on relevant tourist spots and region-specific information. A search algorithm is used to efficiently collect highly relevant information.
[0458] Step 4:
[0459] Using information processing tools, a generative AI model generates a response based on acquired regional information. The input is the explored regional information. A generative AI model such as GPT-3 is used for text generation, creating a response that combines visual and auditory information. The generated data is then ready to be sent to the user's terminal.
[0460] Step 5:
[0461] The terminal presents the response received from the server to the user through a visual display and audio device. The input is the response data sent from the server. The user can obtain detailed information about tourist attractions through the information displayed on the screen and the audio guides provided.
[0462] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0463] This invention provides a sophisticated information system that can respond to the individual emotional needs of users by combining an emotion engine with an information provision system for tourist information. Below, the program processing of this system will be explained in natural language, along with specific examples.
[0464] The user first accesses the chat interface on their device and enters travel-related questions and requests in natural language. For example, a possible question might be, "Where are some hidden spots in Kyoto where I can relax in peace?" The device then sends this question directly to the server.
[0465] The server uses a natural language processing engine to analyze the received question. This analysis extracts the main point of the question and related keywords, and also uses an emotion engine to identify the user's emotions from their writing. For example, the phrase "I want to spend some time quietly" suggests that the user is seeking relaxation.
[0466] Based on the analyzed intent and emotional information, the server searches the database for the most relevant local information. Prioritizing information about quiet places frequently used by locals, the search might yield information about a quiet temple or garden in Kyoto.
[0467] The server generates responses based on search results, adjusting the content and format to reflect emotional information. For users seeking relaxation, information with calming images and music is recommended. For example, content might be created combining tranquil images of the relevant temple or garden with soothing music.
[0468] Finally, the server sends the generated response to the terminal. The terminal receives it and provides the information in a format that is easily understandable to the user. Through sight and hearing, the user can obtain detailed and valuable tourist information that addresses their specific emotional needs and can be used to plan their trip.
[0469] Thus, by providing information that takes user emotions into consideration, the present invention realizes a more personalized travel experience that goes beyond general tourist information.
[0470] The following describes the processing flow.
[0471] Step 1:
[0472] The user accesses the chat interface on their device and enters a question about sightseeing in natural language. For example, they might enter a question like, "Where are some quiet cafes in Hokkaido that locals recommend?"
[0473] Step 2:
[0474] The terminal sends the user's input directly to the server. This transmission process is asynchronous, ensuring that user input is delivered to the server quickly.
[0475] Step 3:
[0476] The server first analyzes the received question using a natural language processing engine. The analysis identifies the subject of the question and the type of information being sought, and extracts keywords. In this case, the keywords extracted are "recommended quiet cafes" and "Hokkaido."
[0477] Step 4:
[0478] After the server analyzes the subject of the question, it uses an emotion engine to identify the emotions underlying the user's words. For example, the phrase "quiet cafe" suggests that the user is seeking a quiet and calm atmosphere.
[0479] Step 5:
[0480] The server searches the database for relevant local information based on the analyzed keywords and sentiment information. Here, it selects information about cafes in Hokkaido with a quiet and relaxed atmosphere, and refers to ratings and reviews from local residents.
[0481] Step 6:
[0482] The server generates a response that reflects search results and sentiment information. If a user is looking for a "calm atmosphere," the server adjusts the text, photos, videos, and audio presentation to match that sentiment. Specifically, this might involve setting relaxing music in the background and providing images or videos of a "quiet and comfortable cafe."
[0483] Step 7:
[0484] The generated response is sent from the server to the terminal, where the user receives it. The terminal appropriately displays the transmitted information and, as a response to the user's input, provides information that meets their anticipated emotional needs. The user can then use this information to create a more fulfilling travel plan.
[0485] (Example 2)
[0486] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0487] Traditional tourist information systems only provide general information in response to user questions, making it difficult to address the emotional needs of individual users. This resulted in limited effectiveness of tourist information and a failure to adequately meet the experiences users desired.
[0488] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0489] In this invention, the server includes input means for the user to input a question in natural language, analysis means for receiving the question from the user and analyzing the question to understand its intent and emotion, and search means for retrieving relevant local information from a database based on the analyzed intent and emotion. This makes it possible to provide more personalized tourist information that meets the individual emotional needs of the user.
[0490] "Input means" refers to a method or device that enables a user to input a question in natural language.
[0491] "Analysis means" refers to a method or apparatus for analyzing questions received from a user and understanding their intent and emotions.
[0492] "Search means" refers to a method or apparatus for retrieving relevant local information from a database based on analyzed intentions and emotions.
[0493] "Generation means" refers to a method or apparatus for generating responses in text, video, and audio formats, adjusting the content and format of the response according to the user's emotions based on the retrieved regional information.
[0494] "Transmission means" refers to a method or apparatus for sending the generated response to the user.
[0495] "Local information" refers to detailed information about a specific area, and may include specialized information collected from local residents.
[0496] This invention is a system that provides personalized tourist information tailored to the emotional needs of users. The system consists of a terminal for users to input questions in natural language and a server that performs analysis and provides information.
[0497] Users enter questions or requests related to specific tourist destinations using the chat interface on their device. They are required to enter their questions in natural language. For example, if a user enters "Please tell me about quiet places to spend time in Kyoto," this will be recognized as a prompt by the device.
[0498] The terminal receives input from the user and sends that data to the server. The server uses a generative AI model and a natural language processing engine to analyze the content of the question and identify the user's intent and emotions. Emotion identification is performed using an emotion engine, which can identify what kind of emotional experience the user is seeking.
[0499] Based on identified intentions and emotions, the server searches the database for relevant local information. This local information includes specialized data provided by local residents, allowing the server to find information that meets the user's needs. For example, information about quiet temples and gardens in Kyoto might be retrieved from the server.
[0500] Furthermore, the server uses the generated AI model to create responses to the user in text, video, and audio formats based on the search results. During generation, it selects expressions and content that match the user's emotions and intentions, providing responses that combine calming video and music to users seeking relaxation.
[0501] Ultimately, the response generated by the server is sent to the terminal, and the user can view and listen to it to obtain valuable information that matches their emotional needs. This invention allows users to enjoy a tourism experience tailored to their individual emotions.
[0502] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0503] Step 1:
[0504] The user accesses the chat interface on their device and enters travel-related questions in natural language. For example, they might enter a prompt such as, "Please tell me about quiet places to stay in Kyoto." The device sends this input directly to the server. The input data is a natural language question, and the output to the server is this question data.
[0505] Step 2:
[0506] The server passes the received question to a natural language processing engine, which analyzes the gist of the question. Specifically, it extracts important keywords and identifies the information the user is seeking. This process takes a natural language question as input and outputs the analyzed intent and a set of keywords.
[0507] Step 3:
[0508] The server uses an emotion engine to identify the emotions expressed in the user's questions. For example, it senses that the user is seeking relaxation from the expression "I want to spend some time quietly." At this stage, the input is the analyzed question, and the output is emotion data.
[0509] Step 4:
[0510] The server uses the analyzed intent and sentiment data to search for appropriate local information from its information database. Here, data that particularly matches the user's emotional needs is prioritized. The input is a dataset of intents and sentiments, and the output provides relevant local information.
[0511] Step 5:
[0512] The server uses a generative AI model to generate responses to the user based on the searched local information. These responses are available in text, video, and audio formats, and their content and expression are adjusted according to the user's emotions. For example, they might include calming video and music. The input for this step is local information, and the output is emotionally responsive content.
[0513] Step 6:
[0514] The server sends the generated response to the terminal. The terminal displays the response in a user-friendly format. The input is the response content from the server, and the output is visual and auditory information provided to the user. Based on this, the user can then create a concrete travel plan.
[0515] (Application Example 2)
[0516] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0517] In modern society, users face the challenge of quickly and accurately obtaining information that matches their emotions and needs from a vast amount of data. Especially in virtual stores, there is a need to provide a more personalized experience by offering product suggestions based on the user's emotions.
[0518] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0519] In this invention, the server includes an input device for the user to input a question in natural language, an analysis device for analyzing the received question to understand intent and emotional information, and a search device for selecting the optimal data based on that information. This makes it possible to provide accurate information and product suggestions that respond to the user's emotions.
[0520] A "user" is a person who inputs information and receives services based on that information.
[0521] "Natural language" refers to the method of expressing information using the language that humans use on a daily basis.
[0522] An "input device" refers to a device used by a user to provide information to a system.
[0523] An "analysis device" is a device that analyzes input information and understands its intent and associated emotions.
[0524] "Emotional information" refers to data related to emotions extracted from information entered by the user.
[0525] A "search device" is a device used to select specific data based on analyzed information.
[0526] A "generation device" is a device that forms the optimal response for the user based on selected data.
[0527] A "display device" is a device that outputs a generated response visually or audibly.
[0528] This invention is a system that provides appropriate information based on the user's emotional state. This system allows the user to input information via a smart device and generate an optimal response based on that information.
[0529] The user first inputs a question in natural language through an input device such as smart glasses. This input can be done either by voice using the device's microphone or by text. This information is sent to a server. The server analyzes the user's question using natural language processing software (e.g., Google Cloud Natural Language API). In this analysis process, a sentiment analysis engine (e.g., IBM Watson Tone Analyzer) is used to identify the user's emotions in addition to the purpose of the question.
[0530] Based on the analyzed intent and emotional information, the server searches the database for relevant items. This narrows down the products and information to match the user's needs. The generated response is visually displayed on the smart glasses' screen. This display includes images and descriptions of relaxing products for users seeking relaxation.
[0531] For example, if a user enters "I feel like relaxing today," the system will visually suggest relaxing aromatherapy candles or houseplants. Another example of a prompt would be, "How can I provide personalized product suggestions based on the user's emotions?"
[0532] This approach allows users to instantly find the products and information that best suit their emotional state, resulting in a personalized shopping experience.
[0533] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0534] Step 1:
[0535] The user inputs a question in natural language through an input device such as smart glasses. The input language data is converted into text data through the device's speech recognition system. For example, if the voice input is "I feel like relaxing today," it will be converted into text format.
[0536] Step 2:
[0537] The terminal sends the converted text data to the server. The server receives this data and analyzes the sentence using a natural language processing engine (e.g., Google Cloud Natural Language API). This analysis infers the user's intent and extracts emotional information from the statement using an emotion engine. The output of the analysis includes the keyword "relax" and the emotional information "I want to relax."
[0538] Step 3:
[0539] Based on the analysis results, the server searches the database for relevant products and information. At this time, it filters the data for products with relaxing effects based on the emotional information "relaxation." For example, aromatherapy candles and houseplants might be selected as relevant products.
[0540] Step 4:
[0541] The server generates a visually appealing response based on the selected product information. The generated response consists of a format that includes text, images, and audio guidance. This includes images of the selected product, text describing its effects, and, in some cases, audio guidance.
[0542] Step 5:
[0543] The server sends the generated response to the terminal. The terminal receives this and displays the content of the response to the user in real time on the smart glasses' display. The user can visually check images and descriptions of products related to "relaxation" and decide on the next action.
[0544] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0545] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0546] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0547] [Fourth Embodiment]
[0548] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0549] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0550] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0551] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0552] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0553] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0554] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0555] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0556] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0557] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0558] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0559] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0560] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0561] This invention is a system in which a user inputs a question about tourism in natural language, the system analyzes the question, and generates and provides a detailed response based on local information. The processing of this system's program is described below in natural language.
[0562] The user begins by typing a question into the chat screen on their device. For example, they might type a question like, "What are some recommended hidden tourist spots in Kyoto?" The device then sends this question directly to the server.
[0563] The server analyzes the received question. Using natural language processing techniques, it identifies the intent of the question and information such as place names. Here, the intent "recommended hidden tourist spots" and the region "Kyoto" are extracted. Next, the server accesses the database and searches for the necessary local information. Utilizing information collected from local residents, it finds tourist spots that are not generally known.
[0564] When the server generates a response based on the information obtained from the search, it can provide information in a combination of text, video, and audio. For example, if it finds information about a small, popular local garden in Kyoto, it can provide photos and videos of the garden along with the text information. The generated response is then sent to the terminal.
[0565] The terminal provides the user with received information in a visually and audibly easy-to-understand format. The user can confirm the appeal of the tourist spots introduced on the screen and gain a deeper understanding through audio and video. In this form, the present invention provides detailed and valuable information that goes beyond general tourist information and is tailored to individual needs.
[0566] The following describes the processing flow.
[0567] Step 1:
[0568] The user accesses the chat interface on their device and enters a question about sightseeing in natural language. For example, they might enter the question, "What are some local recommendations for restaurants in Tokyo?"
[0569] Step 2:
[0570] The terminal sends the user's entered questions to the server. This transmission occurs in real time and is the first step in quickly processing the user's requests.
[0571] Step 3:
[0572] The server analyzes the received question using a natural language processing engine. This analysis identifies the main point and intent of the question and extracts important keywords and phrases. In this case, the keywords "recommended restaurants" and "Tokyo" are extracted.
[0573] Step 4:
[0574] The server searches a local information database based on the analyzed information. This database contains niche information collected from local residents, and the server narrows down the relevant information based on this. In this example, it retrieves information on a specific restaurant that is highly rated locally.
[0575] Step 5:
[0576] The server generates a response to provide to the user based on the search results. It combines text, image data, video links, and audio files to prepare information in a format that is easily understandable to the user. For example, it might prepare details and photos of a specific restaurant in Tokyo.
[0577] Step 6:
[0578] The server sends the generated response to the terminal. The transmitted data is displayed or played back in an appropriate format on the user interface.
[0579] Step 7:
[0580] The terminal displays responses received from the server, outputting text information to the screen as needed, while also playing video and audio. This allows users to obtain detailed local tourist information and use it to plan their trips.
[0581] (Example 1)
[0582] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0583] With the development of the information society, users are increasingly seeking detailed and personalized information specific to their region. However, traditional information systems face challenges in accurately understanding user intent and providing specific, region-specific information. Furthermore, there is a demand for information provision that appeals not only to text but also to visual and auditory senses, providing rich content. This necessitates improving the user experience.
[0584] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0585] In this invention, the server includes an interface means for the user to input a question in natural language; an analysis means for receiving the question input by the user and analyzing the question using generative AI technology to understand the intent and related information; an information retrieval means for searching for relevant geographic information from a data set based on the analyzed intent; and an information distribution means for providing the generated response to the user. This enables the user to obtain detailed and diverse content that is relevant to their intent and includes region-specific information.
[0586] "Interface means" refers to devices or software that allow users to input questions into a system using natural language.
[0587] "Generative AI technology" refers to a technology that uses artificial intelligence to analyze the intent of a user's question from natural language input and derive an appropriate response.
[0588] "Analysis means" refers to the process or method of analyzing received questions using artificial intelligence technology to identify intent and related information.
[0589] "Information retrieval methods" refer to the processes and technologies used to search for relevant geographic information from a data set based on analyzed intent.
[0590] "Information generation means" refers to functions and methods for generating responses in various formats (text, video, audio) based on retrieved geographic information.
[0591] "Information distribution means" refers to communication means and display devices used to deliver the generated response to the user.
[0592] This invention is a system in which users ask questions about tourism in natural language and obtain local information based on those questions. The specific techniques for implementing this system are described below.
[0593] The user first enters their question through an interface on their device. This interface is designed to allow users to easily enter questions in natural language. For example, they can enter a question such as, "What are some recommended hidden tourist spots in Kyoto?" An example of a prompt would be, "What are some recommended hidden tourist spots in Kyoto? I want to know about places that aren't well known to locals."
[0594] The terminal sends the entered question to the server. The server analyzes the question using generative AI technology, such as the GPT model. The analysis identifies the user's intent and relevant information such as place names. Based on this analysis, the server searches for relevant geographic information from the data set. The data set also includes unique information collected from local residents.
[0595] Based on the search results, the server uses information generation tools to create responses in text, video, and audio formats. These responses can combine diverse media to form rich content. For example, if the server finds information about a hidden garden in Kyoto, it will provide text descriptions of its history and highlights, along with photos and videos of the garden.
[0596] Finally, the generated response is provided to the user through the terminal's information distribution means. The user can visually and audibly understand the answer to the question through the displayed information. In this way, the present invention makes it possible to provide the user with detailed and specific tourist information they desire.
[0597] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0598] Step 1:
[0599] Users input travel-related questions in natural language using an interface on their device. For example, they can enter specific questions such as, "What are some recommended hidden tourist spots in Kyoto?" The device receives this input, formats the data, and sends it to the server via the network.
[0600] Step 2:
[0601] The server analyzes questions received from terminals using generation AI technology. In this analysis, the AI model processes text data to identify the intent of the question, place names, and areas of interest. It receives natural language question text as input and outputs the analysis results. These analysis results may include keywords such as "hidden tourist spots" and "Kyoto."
[0602] Step 3:
[0603] The server searches for geographical information within the data set based on the analyzed keywords. The data set contains unique information collected from local residents. The input is the keywords from the analysis results, and the output is information on related tourist spots. Information retrieval is performed through a high-speed database search algorithm.
[0604] Step 4:
[0605] The server generates a multimedia-based response based on the information obtained from the search. Search results are used as input, and the output is content combining text, images, and video. This generation process utilizes information generation methods to provide visually and aurally rich content.
[0606] Step 5:
[0607] The terminal receives multimedia responses sent from the server and provides them to the user. Its input is multimedia data from the server, and its output is the presentation of visual and auditory information to the user. Based on this information, the user can obtain detailed and personalized answers to their questions.
[0608] (Application Example 1)
[0609] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0610] Traditional tourist information systems only provide information based on user input, making it difficult to provide detailed, location-specific information or hidden gems in real time when a user visits a particular tourist destination. Therefore, there was a need for a method that could personalize the tourist experience and provide users with a deeper understanding.
[0611] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0612] In this invention, the server includes an interface means for the user to input a question in natural language, a natural language processing means for receiving the question from the user, analyzing the question, and understanding its intent, and a search means for retrieving relevant local information from a data store based on the analyzed intent and the user's current location information. This makes it possible to automatically provide detailed information and guidance on hidden attractions related to a place when the user visits a tourist destination.
[0613] An "interface means" is an input device or platform for users to input questions in natural language.
[0614] "Natural language processing means" refers to technologies and systems that analyze questions entered by users and understand their intent.
[0615] The "search method" is a function that searches for relevant regional information from the data store based on the analyzed intent and the user's current location information.
[0616] "Information processing means" refers to the process of generating responses in the form of text, images, and sound based on the searched regional information.
[0617] "Means of communication" refers to a method or system for transmitting a generated response to a user through a visual display or auditory device.
[0618] "Local information" refers to data and insights related to a specific geographical area, and in particular, information that includes hidden places and anecdotes collected from local residents.
[0619] "Visual information" refers to data related to videos and images presented to the user.
[0620] "Acoustic information" refers to audio and sound-related data presented to the user.
[0621] This invention provides a system for users to input tourism-related questions in natural language and obtain detailed information based on those inputs. The system has the following configuration:
[0622] Users input questions in natural language via devices such as smart glasses or mobile devices. For example, they might input a question like, "What are some recommended tourist spots nearby?" through the interface. In this process, visual displays and sound devices implemented on the device are used.
[0623] The server analyzes received questions using natural language processing to understand their intent. For example, common natural language processing libraries such as NLTK and Spacy can be used. The analyzed information is combined with the user's current location information to access a data store and search for relevant regional information. The search mechanism extracts useful data to present to the user.
[0624] The information processing system uses a generative AI model to generate responses to the user in text, image, and audio formats based on the explored local information. These generated responses are then communicated to the user through visual displays and auditory devices. For example, when a user visits a temple in Kyoto, the system can provide information about the temple's history, highlights, and nearby hidden gems.
[0625] The following prompt statements can be used with the generative AI model.
[0626] Example prompt: "What are some hidden gems in this area that everyone overlooks?"
[0627] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0628] Step 1:
[0629] The user enters the question in natural language through their device. The input is in text format and can be done using the voice input function or touch interface of smart glasses. The user confirms the question on the device and sends the entered information to the server by pressing the submit button.
[0630] Step 2:
[0631] The server analyzes received questions using natural language processing (NLTK) techniques. By analyzing the input question text, it extracts the intent and keywords of the question. For example, intents such as "tourist spots" and "recommendations" are extracted here. Natural language processing libraries such as NLTK and Spacy are used for the analysis.
[0632] Step 3:
[0633] The server searches for local information from the data store based on the analyzed intent and the user's current location. The inputs are the extracted intent and GPS information. The data store retrieves information on relevant tourist spots and region-specific information. A search algorithm is used to efficiently collect highly relevant information.
[0634] Step 4:
[0635] Using information processing tools, a generative AI model generates a response based on acquired regional information. The input is the explored regional information. A generative AI model such as GPT-3 is used for text generation, creating a response that combines visual and auditory information. The generated data is then ready to be sent to the user's terminal.
[0636] Step 5:
[0637] The terminal presents the response received from the server to the user through a visual display and audio device. The input is the response data sent from the server. The user can obtain detailed information about tourist attractions through the information displayed on the screen and the audio guides provided.
[0638] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0639] This invention provides a sophisticated information system that can respond to the individual emotional needs of users by combining an emotion engine with an information provision system for tourist information. Below, the program processing of this system will be explained in natural language, along with specific examples.
[0640] The user first accesses the chat interface on their device and enters travel-related questions and requests in natural language. For example, a possible question might be, "Where are some hidden spots in Kyoto where I can relax in peace?" The device then sends this question directly to the server.
[0641] The server uses a natural language processing engine to analyze the received question. This analysis extracts the main point of the question and related keywords, and also uses an emotion engine to identify the user's emotions from their writing. For example, the phrase "I want to spend some time quietly" suggests that the user is seeking relaxation.
[0642] Based on the analyzed intent and emotional information, the server searches the database for the most relevant local information. Prioritizing information about quiet places frequently used by locals, the search might yield information about a quiet temple or garden in Kyoto.
[0643] The server generates responses based on search results, adjusting the content and format to reflect emotional information. For users seeking relaxation, information with calming images and music is recommended. For example, content might be created combining tranquil images of the relevant temple or garden with soothing music.
[0644] Finally, the server sends the generated response to the terminal. The terminal receives it and provides the information in a format that is easily understandable to the user. Through sight and hearing, the user can obtain detailed and valuable tourist information that addresses their specific emotional needs and can be used to plan their trip.
[0645] Thus, by providing information that takes user emotions into consideration, the present invention realizes a more personalized travel experience that goes beyond general tourist information.
[0646] The following describes the processing flow.
[0647] Step 1:
[0648] The user accesses the chat interface on their device and enters a question about sightseeing in natural language. For example, they might enter a question like, "Where are some quiet cafes in Hokkaido that locals recommend?"
[0649] Step 2:
[0650] The terminal sends the user's input directly to the server. This transmission process is asynchronous, ensuring that user input is delivered to the server quickly.
[0651] Step 3:
[0652] The server first analyzes the received question using a natural language processing engine. The analysis identifies the subject of the question and the type of information being sought, and extracts keywords. In this case, the keywords extracted are "recommended quiet cafes" and "Hokkaido."
[0653] Step 4:
[0654] After the server analyzes the subject of the question, it uses an emotion engine to identify the emotions underlying the user's words. For example, the phrase "quiet cafe" suggests that the user is seeking a quiet and calm atmosphere.
[0655] Step 5:
[0656] The server searches the database for relevant local information based on the analyzed keywords and sentiment information. Here, it selects information about cafes in Hokkaido with a quiet and relaxed atmosphere, and refers to ratings and reviews from local residents.
[0657] Step 6:
[0658] The server generates a response that reflects search results and sentiment information. If a user is looking for a "calm atmosphere," the server adjusts the text, photos, videos, and audio presentation to match that sentiment. Specifically, this might involve setting relaxing music in the background and providing images or videos of a "quiet and comfortable cafe."
[0659] Step 7:
[0660] The generated response is sent from the server to the terminal, where the user receives it. The terminal appropriately displays the transmitted information and, as a response to the user's input, provides information that meets their anticipated emotional needs. The user can then use this information to create a more fulfilling travel plan.
[0661] (Example 2)
[0662] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0663] Traditional tourist information systems only provide general information in response to user questions, making it difficult to address the emotional needs of individual users. This resulted in limited effectiveness of tourist information and a failure to adequately meet the experiences users desired.
[0664] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0665] In this invention, the server includes input means for the user to input a question in natural language, analysis means for receiving the question from the user and analyzing the question to understand its intent and emotion, and search means for retrieving relevant local information from a database based on the analyzed intent and emotion. This makes it possible to provide more personalized tourist information that meets the individual emotional needs of the user.
[0666] "Input means" refers to a method or device that enables a user to input a question in natural language.
[0667] "Analysis means" refers to a method or apparatus for analyzing questions received from a user and understanding their intent and emotions.
[0668] "Search means" refers to a method or apparatus for retrieving relevant local information from a database based on analyzed intentions and emotions.
[0669] "Generation means" refers to a method or apparatus for generating responses in text, video, and audio formats, adjusting the content and format of the response according to the user's emotions based on the retrieved regional information.
[0670] "Transmission means" refers to a method or apparatus for sending the generated response to the user.
[0671] "Local information" refers to detailed information about a specific area, and may include specialized information collected from local residents.
[0672] This invention is a system that provides personalized tourist information tailored to the emotional needs of users. The system consists of a terminal for users to input questions in natural language and a server that performs analysis and provides information.
[0673] Users enter questions or requests related to specific tourist destinations using the chat interface on their device. They are required to enter their questions in natural language. For example, if a user enters "Please tell me about quiet places to spend time in Kyoto," this will be recognized as a prompt by the device.
[0674] The terminal receives input from the user and sends that data to the server. The server uses a generative AI model and a natural language processing engine to analyze the content of the question and identify the user's intent and emotions. Emotion identification is performed using an emotion engine, which can identify what kind of emotional experience the user is seeking.
[0675] Based on identified intentions and emotions, the server searches the database for relevant local information. This local information includes specialized data provided by local residents, allowing the server to find information that meets the user's needs. For example, information about quiet temples and gardens in Kyoto might be retrieved from the server.
[0676] Furthermore, the server uses the generated AI model to create responses to the user in text, video, and audio formats based on the search results. During generation, it selects expressions and content that match the user's emotions and intentions, providing responses that combine calming video and music to users seeking relaxation.
[0677] Ultimately, the response generated by the server is sent to the terminal, and the user can view and listen to it to obtain valuable information that matches their emotional needs. This invention allows users to enjoy a tourism experience tailored to their individual emotions.
[0678] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0679] Step 1:
[0680] The user accesses the chat interface on their device and enters travel-related questions in natural language. For example, they might enter a prompt such as, "Please tell me about quiet places to stay in Kyoto." The device sends this input directly to the server. The input data is a natural language question, and the output to the server is this question data.
[0681] Step 2:
[0682] The server passes the received question to a natural language processing engine, which analyzes the gist of the question. Specifically, it extracts important keywords and identifies the information the user is seeking. This process takes a natural language question as input and outputs the analyzed intent and a set of keywords.
[0683] Step 3:
[0684] The server uses an emotion engine to identify the emotions expressed in the user's questions. For example, it senses that the user is seeking relaxation from the expression "I want to spend some time quietly." At this stage, the input is the analyzed question, and the output is emotion data.
[0685] Step 4:
[0686] The server uses the analyzed intent and sentiment data to search for appropriate local information from its information database. Here, data that particularly matches the user's emotional needs is prioritized. The input is a dataset of intents and sentiments, and the output provides relevant local information.
[0687] Step 5:
[0688] The server uses a generative AI model to generate responses to the user based on the searched local information. These responses are available in text, video, and audio formats, and their content and expression are adjusted according to the user's emotions. For example, they might include calming video and music. The input for this step is local information, and the output is emotionally responsive content.
[0689] Step 6:
[0690] The server sends the generated response to the terminal. The terminal displays the response in a user-friendly format. The input is the response content from the server, and the output is visual and auditory information provided to the user. Based on this, the user can then create a concrete travel plan.
[0691] (Application Example 2)
[0692] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0693] In modern society, users face the challenge of quickly and accurately obtaining information that matches their emotions and needs from a vast amount of data. Especially in virtual stores, there is a need to provide a more personalized experience by offering product suggestions based on the user's emotions.
[0694] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0695] In this invention, the server includes an input device for the user to input a question in natural language, an analysis device for analyzing the received question to understand intent and emotional information, and a search device for selecting the optimal data based on that information. This makes it possible to provide accurate information and product suggestions that respond to the user's emotions.
[0696] A "user" is a person who inputs information and receives services based on that information.
[0697] "Natural language" refers to the method of expressing information using the language that humans use on a daily basis.
[0698] An "input device" refers to a device used by a user to provide information to a system.
[0699] An "analysis device" is a device that analyzes input information and understands its intent and associated emotions.
[0700] "Emotional information" refers to data related to emotions extracted from information entered by the user.
[0701] A "search device" is a device used to select specific data based on analyzed information.
[0702] A "generation device" is a device that forms the optimal response for the user based on selected data.
[0703] A "display device" is a device that outputs a generated response visually or audibly.
[0704] This invention is a system that provides appropriate information based on the user's emotional state. This system allows the user to input information via a smart device and generate an optimal response based on that information.
[0705] The user first inputs a question in natural language through an input device such as smart glasses. This input can be done either by voice using the device's microphone or by text. This information is sent to a server. The server analyzes the user's question using natural language processing software (e.g., Google Cloud Natural Language API). In this analysis process, a sentiment analysis engine (e.g., IBM Watson Tone Analyzer) is used to identify the user's emotions in addition to the purpose of the question.
[0706] Based on the analyzed intent and emotional information, the server searches the database for relevant items. This narrows down the products and information to match the user's needs. The generated response is visually displayed on the smart glasses' screen. This display includes images and descriptions of relaxing products for users seeking relaxation.
[0707] For example, if a user enters "I feel like relaxing today," the system will visually suggest relaxing aromatherapy candles or houseplants. Another example of a prompt would be, "How can I provide personalized product suggestions based on the user's emotions?"
[0708] This approach allows users to instantly find the products and information that best suit their emotional state, resulting in a personalized shopping experience.
[0709] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0710] Step 1:
[0711] The user inputs a question in natural language through an input device such as smart glasses. The input language data is converted into text data through the device's speech recognition system. For example, if the voice input is "I feel like relaxing today," it will be converted into text format.
[0712] Step 2:
[0713] The terminal sends the converted text data to the server. The server receives this data and analyzes the sentence using a natural language processing engine (e.g., Google Cloud Natural Language API). This analysis infers the user's intent and extracts emotional information from the statement using an emotion engine. The output of the analysis includes the keyword "relax" and the emotional information "I want to relax."
[0714] Step 3:
[0715] Based on the analysis results, the server searches the database for relevant products and information. At this time, it filters the data for products with relaxing effects based on the emotional information "relaxation." For example, aromatherapy candles and houseplants might be selected as relevant products.
[0716] Step 4:
[0717] The server generates a visually appealing response based on the selected product information. The generated response consists of a format that includes text, images, and audio guidance. This includes images of the selected product, text describing its effects, and, in some cases, audio guidance.
[0718] Step 5:
[0719] The server sends the generated response to the terminal. The terminal receives this and displays the content of the response to the user in real time on the smart glasses' display. The user can visually check images and descriptions of products related to "relaxation" and decide on the next action.
[0720] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0721] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0722] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0723] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0724] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0725] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0726] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0727] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0728] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0729] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0730] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0731] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0732] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0733] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0734] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0735] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0736] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0737] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0738] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0739] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0740] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0741] The following is further disclosed regarding the embodiments described above.
[0742] (Claim 1)
[0743] An input method for users to enter questions in natural language,
[0744] An analytical means that receives a question entered by a user, analyzes the question, and understands its intent,
[0745] A search method for retrieving relevant local information from a database based on the analyzed intent,
[0746] A generation means that generates responses in text, video, and audio formats based on searched local information,
[0747] A transmission means for sending the generated response to the user,
[0748] A system that includes this.
[0749] (Claim 2)
[0750] The system according to claim 1, wherein the local information includes niche information collected from local residents.
[0751] (Claim 3)
[0752] The system according to claim 1, wherein the response to the user is provided in a combination of text information, video information, and audio information.
[0753] "Example 1"
[0754] (Claim 1)
[0755] An interface for users to input questions in natural language,
[0756] An analysis means that receives a question entered by a user, analyzes the question using generation AI technology to understand the intent and related information,
[0757] Information retrieval means for searching for relevant geographic information from a data set based on the analyzed intent,
[0758] Information generation means that generates a response by combining information in text, video, and audio formats based on retrieved geographic information,
[0759] Information distribution means that provides the generated response to the user,
[0760] A system that includes this.
[0761] (Claim 2)
[0762] The system according to claim 1, wherein the geographic information includes unique information collected from local communities.
[0763] (Claim 3)
[0764] The system according to claim 1, wherein the response to the user is provided in an integrated format of text information, video information, and audio information.
[0765] "Application Example 1"
[0766] (Claim 1)
[0767] An interface for users to input questions in natural language,
[0768] A natural language processing means that receives a question entered by a user, analyzes the question, and understands its intent,
[0769] A search means for retrieving relevant regional information from a data store based on the analyzed intent and the user's current location information,
[0770] Information processing means for generating responses in text, image, and sound formats based on explored regional information,
[0771] A communication means for transmitting the generated response to the user through a visual display or auditory device,
[0772] A system that includes this.
[0773] (Claim 2)
[0774] The system according to claim 1, wherein local information includes hidden places and anecdotes collected from local residents.
[0775] (Claim 3)
[0776] The system according to claim 1, wherein the response to the user is provided in a combination of visual information, image information, and acoustic information.
[0777] "Example 2 of combining an emotion engine"
[0778] (Claim 1)
[0779] An input method for users to enter questions in natural language,
[0780] An analytical means that receives a question entered by a user, analyzes the question to understand its intent and emotions,
[0781] A search method for retrieving relevant regional information from a database based on analyzed intentions and emotions,
[0782] A generation means that generates responses in text, video, and audio formats based on searched local information, and adjusts the content and format according to the user's emotions,
[0783] A transmission means for sending the generated response to the user,
[0784] A system that includes this.
[0785] (Claim 2)
[0786] The system according to claim 1, wherein the local information includes special information collected from local residents.
[0787] (Claim 3)
[0788] The system according to claim 1, wherein the response to the user is provided in a combination of text information, video information, and audio information, and is customized based on the user's emotions.
[0789] "Application example 2 of combining emotional engines"
[0790] (Claim 1)
[0791] An input device for users to enter questions in natural language,
[0792] An analysis device that receives a question entered by a user, analyzes the question, and understands its intent,
[0793] A search device for selecting appropriate items from a data structure based on emotional information associated with analyzed intentions,
[0794] A generating device that generates responses in text, image, and audio formats based on searched items,
[0795] A display device that displays the generated response on the user's device,
[0796] A system that includes this.
[0797] (Claim 2)
[0798] The system according to claim 1, which provides product suggestions based on the user's emotional state.
[0799] (Claim 3)
[0800] The system according to claim 1, wherein the response to the user is provided in a combination of text information, image information, and audio information. [Explanation of Symbols]
[0801] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. An input method for users to enter questions in natural language, An analytical means that receives a question entered by a user, analyzes the question, and understands its intent, A search method for retrieving relevant local information from a database based on the analyzed intent, A generation means that generates responses in text, video, and audio formats based on searched local information, A transmission means for sending the generated response to the user, A system that includes this.
2. The system according to claim 1, wherein the local information includes niche information collected from local residents.
3. The system according to claim 1, wherein the response to the user is provided in a combination of text information, video information, and audio information.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A