System

The system uses animated characters and a generative AI model to overcome language barriers and enhance the tourist experience by offering interactive and multilingual guidance at Japanese animation and manga sites.

JP2026034054APending Publication Date: 2026-02-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024137175
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-16
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Tourists visiting Japanese animation and manga holy sites face language barriers and lack of interactive experiences, limiting their understanding of the culture and history.

Method used

A system that uses animated characters to provide interactive guidance, supports multiple languages, and generates voice responses using a generative AI model, allowing users to engage with tourist spots and obtain detailed information.

Benefits of technology

Enhances the tourist experience by providing personalized, interactive, and multilingual guidance, improving the understanding of cultural and historical aspects of the sites.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026034054000001_ABST
    Figure 2026034054000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: This system includes a means for generating the guide voice of an animation character, a means for acquiring the information of a sightseeing spot, a means for interactively responding to the question of a user, and a means for dealing in multiple languages.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] The number of tourists wanting to visit the holy sites of Japanese animation and manga is increasing, but language barriers and a lack of information mean that they are unable to have a satisfying experience. Another issue is the lack of experiences that allow tourists to make pilgrimages to holy sites while interacting with characters. This makes it difficult for tourists to gain a deep understanding of the culture and history of the holy sites, and limits the experiences they get at the destinations they visit. [Means for solving the problem]

[0005] The present invention provides a system that includes a means for generating guidance voices for animated characters, a means for acquiring information about tourist spots, a means for interactively responding to user questions, and a means for providing multilingual support. This system allows users to make pilgrimages to sacred sites and obtain a wealth of information while interacting with the characters in real time. Furthermore, the system uses a generative AI model to generate the character's voice and provide appropriate responses to questions, improving the quality of the tourist experience. Furthermore, a translation function into a specified language overcomes language barriers and can accommodate foreign tourists.

[0006] An "animation character" is a fictional character that appears in animation or manga works.

[0007] "Guidance voice" is voice data that conveys information about tourist spots and specific locations by voice.

[0008] A "tourist destination" refers to a specific place that tourists visit and includes elements such as culture, history, and entertainment.

[0009] "Means of obtaining information" refers to the technical means for searching and obtaining the necessary information from databases or the Internet.

[0010] "User" refers to a person who uses a system or service.

[0011] "Means for interactive response" refers to technical means for generating appropriate responses in real time to questions or commands from the user.

[0012] "Multilingual response means" are technological means that enable communication in different languages.

[0013] A "system" is a collection of devices or programs consisting of multiple elements and means configured to achieve a specific purpose.

[0014] A "generative AI model" is an algorithm or software that uses artificial intelligence techniques to generate text or speech.

[0015] A "server" refers to a computer that provides services over a network.

[0016] "Terminal" refers to a device that is directly operated by a user, including smartphones and personal computers.

[0017] A "tour plan" is a plan that defines the specific itinerary and schedule for touring tourist spots.

[0018] A "database" is a collection of specific information that is systematically organized and can be managed and searched. [Brief explanation of the drawings]

[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9]1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0021] First, the terms used in the following description will be explained.

[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0027] [First embodiment]

[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0040] This invention is a system that generates guide voices for animated characters and provides real-time information about tourist spots. This system realizes an interactive pilgrimage tour based on the characters and tourist spots selected by the user. It also supports multiple languages, making it suitable for foreign tourists.

[0041] System configuration

[0042] 1. User Device

[0043] Function: Provides an interface for users to access the system and select characters and tourist attractions, plays audio guides for tourist attractions, and displays and plays responses to user questions in real time.

[0044] How it works: The user accesses a designated website or application using a device such as a smartphone or PC, selects a character and a destination on the interface, and uses text or voice input to ask any questions.

[0045] 2. Server

[0046] Function: Receives character and destination selection information sent from the user's device and retrieves tourist spot information from the database. Uses a generative AI model to generate character voice and text. Furthermore, responds interactively to user questions and supports multiple languages.

[0047] Processing method: Based on the user's selection, the server retrieves relevant sacred site information (such as name, location, and history) from the database. It then uses a generative AI model to generate audio data for the tour guide and sends it to the user's device. When the user enters a question, the server analyzes the question and generates the optimal response.

[0048] 3. Database

[0049] Function: Stores and manages detailed information about tourist spots (history, culture, episodes related to animation and manga, etc.). Provides necessary information in response to requests from the server.

[0050] Management method: The database is updated regularly to include new tourist destination information and information on animation and manga.

[0051] Specific examples

[0052] A specific example of the system is shown below.

[0053] Example 1: A park in Tokyo

[0054] 1. User selection: The user opens the smartphone application and selects Character A from the animation work and a certain park in Tokyo.

[0055] 2. Tour Guide: The server receives the selection information and retrieves detailed information about a certain park from a database. Using a generative AI model, it generates a narration in the voice of Character A, saying "This park is...", and sends it to the user's device. The user plays the narration on their device and tours the park.

[0056] 3. Question and Response: The user enters a question, such as "Tell me about the history of this place." The server analyzes the question and uses a generative AI model to generate a response, such as "This park was founded in ____...," which is sent to the device. The device then plays back the response in the voice of Character A.

[0057] 4. Multilingual support: If the user has English settings, the server will use artificial intelligence to translate the response into English and send it to the device as the English voice of Character A.

[0058] This allows users to have a rich experience touring tourist spots while interacting with animated characters.The purpose of the present invention is to improve the quality of users' sightseeing experiences through such an interactive guide system.

[0059] The processing flow will be explained below.

[0060] Step 1:

[0061] A user starts an application on a terminal and accesses the system. The terminal generates initial session information and sends it to the server.

[0062] Step 2:

[0063] The terminal displays the animated characters and a list of tourist spots on a user interface.

[0064] Step 3:

[0065] Through the interface, the user selects an animated guide character and the tourist sites they wish to visit.

[0066] Step 4:

[0067] The terminal transmits the user's selection (character ID, destination ID, etc.) to the server.

[0068] Step 5:

[0069] The server retrieves information about the selected character and tourist spot from a database.

[0070] Step 6:

[0071] The server generates a tour plan based on the acquired information and determines the spots to visit and the content of the narration.

[0072] Step 7:

[0073] The server generates narration text for the tour plan using the voice of the selected character using a generative AI model, and creates audio data.

[0074] Step 8:

[0075] The server transmits the generated voice data to the user's terminal.

[0076] Step 9:

[0077] The terminal starts playing the audio data received from the server as narration.

[0078] Step 10:

[0079] The user asks questions during the tour (e.g., "Tell me about something special about this place") via text or voice input.

[0080] Step 11:

[0081] The terminal sends the user's question to the server.

[0082] Step 12:

[0083] The server analyzes the question and retrieves relevant information from a database.

[0084] Step 13:

[0085] The server uses a generative AI model to generate the analyzed information as voice data for the character.

[0086] Step 14:

[0087] The server transmits the generated response voice data to the terminal.

[0088] Step 15:

[0089] The terminal reproduces the response voice data received from the server.

[0090] Step 16:

[0091] The server checks the language specified by the user in the settings, translates the voice data into that language if necessary, and sends the translated voice to the device.

[0092] Step 17:

[0093] The device plays the translated audio data and delivers it to the user.

[0094] Example 1

[0095] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0096] In the modern tourism industry, tourists are seeking more personalized experiences, but existing guidance systems are unable to meet these demands. Furthermore, information about tourist destinations is often provided in a one-way manner, creating a need for more interactive services. Furthermore, multilingual support is lacking, creating a significant language barrier for foreign tourists in particular.

[0097] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0098] In this invention, the server includes means for a user to select a character and a destination, means for transmitting the selected information to the server, means for the server to obtain destination information from a database, means for generating a voice guide for the character using a generative AI model, means for responding to user questions, and means for translating the generated response into multiple languages. This allows users to receive personalized interactive tourist information, and the multilingual support makes it easy for foreign tourists to obtain information.

[0099] "User character and destination selection means" refers to a device or software that provides an interface for a user to access the system and input or select a particular animated character and tourist destination.

[0100] The "means for transmitting selected information to the server" refers to a communication device or protocol for transmitting information about the character and destination selected by the user to the server.

[0101] "Means for the server to retrieve destination information from the database" refers to the process or software that the server uses to query a database that stores detailed information about tourist destinations based on the user's selection and retrieve the required information.

[0102] "Means for generating voice guidance for a character using a generative AI model" refers to a system or software that uses a generative AI model (e.g., a model using natural language generation technology) to generate voice guidance for a tourist in the voice of a character selected by the user.

[0103] The "means for responding to a user's question" refers to a system or software that analyzes a question entered by a user, generates an appropriate response, and provides that response to the user.

[0104] The "means for translating the generated response into multiple languages" refers to a translation system or software for translating the generated response into a language designated by the user.

[0105] "Means for the terminal to transmit information on the selected character and destination to the server" refers to communication functions or software that allow the user's terminal to transmit information on the selected character and tourist spot to the server.

[0106] "Means for the server to generate a guide plan based on a destination" refers to a system or software that enables the server to generate a detailed guide plan including descriptions of tourist spots and route guidance based on the destination selected by the user.

[0107] "Means for generating character voice using a generative AI model" refers to a system or software that utilizes a generative AI model to generate audio content in the voice of a character selected by a user.

[0108] "Means for providing responses to user questions through a character" refers to a system or software that provides responses to user questions through the voice of a character generated using a generative AI model.

[0109] This invention is a system that uses animated characters to guide tourists around tourist spots, and its features include being interactive and supporting multiple languages. This system is composed of multiple hardware and software components, including a user terminal, a server, a database, and a generative AI model.

[0110] System configuration

[0111] 1. User Device

[0112] Function: Provides an interface for users to access the system and select characters and tourist attractions, plays audio guides for tourist attractions, and responds to user questions in real time.

[0113] Details: Through a smartphone or PC application, users select a character and a destination on the interface, and ask questions by text or voice input. For example, a user opens a smartphone app and selects animated character A and a park in Tokyo.

[0114] 2. Server

[0115] Function: Receives selection information sent from the user's device, retrieves tourist destination information from the database, and generates character voice and text using a generative AI model. Furthermore, it responds interactively to user questions and provides multilingual support as needed.

[0116] Details: Based on the user's selection, the server queries the database for detailed information about related tourist attractions. It then uses a generative AI model (e.g., ChatGPT® by OpenAI®) to generate voice data and responses for the tour guide and sends them to the user's device.

[0117] 3. Database

[0118] Function: Stores and manages detailed information about tourist spots (such as information about history, culture, and anime). Provides necessary information in response to requests from the server.

[0119] Details: The database is regularly updated with new tourist information and animations. It also contains historical information and specific anecdotes about the tourist destinations.

[0120] Specific examples

[0121] Below is a concrete example of how the system works:

[0122] Example 1: A park in Tokyo

[0123] 1. User selection: The user opens the app on their smartphone and selects Character A from an animation work and a certain park in Tokyo.

[0124] 2. Tour Guide: The server receives the selection information and retrieves detailed information about a certain park from the database. Using the generative AI model, it generates a narration in the voice of Character A, saying "This park is...", and sends it to the user's device.

[0125] 3. Question and Response: The user enters a question such as "Tell me about the history of this place." The server analyzes the question and uses a generative AI model to generate a response such as "This park was founded in ____..." and sends it to the device. The device plays the response in the voice of Character A.

[0126] 4. Multilingual support: If the user has English settings, the server will use artificial intelligence to translate the response into English and send it to the device as the English voice of Character A.

[0127] Prompt Sentence Examples

[0128] "The selected character is Character A, and the tourist destination is a certain park in Tokyo. Please generate audio of a tourist guide about this park."

[0129] The purpose of this invention is to improve the user's sightseeing experience through such an interactive guide system. In addition, by supporting multiple languages, it can also accommodate foreign tourists, thereby meeting global tourism demand.

[0130] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0131] Step 1:

[0132] The user selects a character and a destination

[0133] Specific operation: The user opens the application on their smartphone or computer and selects Character A and a certain park in Tokyo on the interface.

[0134] Input: User selection of character and tourist spot.

[0135] Output: Information about the selected character and tourist spot (e.g., Character A, a certain park).

[0136] Step 2:

[0137] The device sends the selection information to the server

[0138] Specific operation: The device sends JSON data containing information about the selected character and tourist spot to the server as an HTTP request.

[0139] Input: Selected character and tourist attraction information.

[0140] Output: The selection information sent to the server.

[0141] Step 3:

[0142] The server retrieves tourist information from the database

[0143] Specific operation: The server sends a query to the database to obtain information about "a certain park," such as its history and characteristics.

[0144] Input: User selection information (Character A, a certain park).

[0145] Output: Tourist information retrieved from the database (e.g. park history, main attractions, etc.).

[0146] Step 4:

[0147] The server inputs the prompt sentence into the AI ​​model and generates the guidance voice.

[0148] Specific operation: Based on the tourist attraction information obtained by the server, the prompt sentence "The selected character is Character A, and the tourist attraction is a certain park in Tokyo. Please generate audio tourist information about this park" is input into the generation AI model.

[0149] Input: Obtained tourist attraction information and prompt sentence.

[0150] Output: Guidance speech data generated by the generative AI model.

[0151] Step 5:

[0152] Send the generated audio data to the device

[0153] Specific operation: The server sends the generated voice data to the user's terminal.

[0154] Input: Guidance speech data generated by a generative AI model.

[0155] Output: The audio data sent to the device.

[0156] Step 6:

[0157] The device plays a voice prompt to the user.

[0158] Specific behavior: The device plays the audio data received and lets the user listen to it. Uses the audio player in the app.

[0159] Input: Audio data sent to the device.

[0160] Output: The audio guidance the user hears.

[0161] Step 7:

[0162] The user enters a question

[0163] What happens: A user types a question into a text field in the app: "Tell me about the history of this place."

[0164] Input: The question entered by the user.

[0165] Output: The data from the question entered.

[0166] Step 8:

[0167] The device sends a question to the server

[0168] Specific operation: The device sends JSON data containing the question to the server as an HTTP request.

[0169] Input: The question data entered by the user.

[0170] Output: The query data sent to the server.

[0171] Step 9:

[0172] The server analyzes the question and inputs the prompt into the generative AI model to generate a response.

[0173] Specific operation: The server analyzes the question and inputs the prompt sentence, "The user asked me, 'Tell me about the history of this place.' Please generate an answer to this question." into the generative AI model.

[0174] Input: User question data and generated prompt text.

[0175] Output: The response data generated by the generative AI model.

[0176] Step 10:

[0177] The server sends the response data to the terminal.

[0178] Specific operation: The server sends the generated response data to the terminal.

[0179] Input: The response data generated by the generative AI model.

[0180] Output: The response data sent to the device.

[0181] Step 11:

[0182] The terminal plays the response to the user

[0183] Specific behavior: The device plays back the response data received and lets the user listen to it. Uses the in-app audio player.

[0184] Input: Response data sent to the terminal.

[0185] Output: The response audio that the user hears.

[0186] Step 12:

[0187] If necessary, the server translates the response for multilingual support.

[0188] Specific operation: The server checks the user's language setting and translates the generated response if necessary. For example, if the user's language setting is English, the server translates the response into English and generates English speech data using the generative AI model.

[0189] Input: The user's preferred language and the generated response data.

[0190] Output: The translated response audio data.

[0191] (Application example 1)

[0192] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0193] Currently, guidance at tourist spots and brick-and-mortar stores requires human intervention, and it is difficult to provide multilingual and interactive guidance. Furthermore, when users want to obtain detailed information on their smartphones or in-store terminals, there is a lack of immediate guidance in a user-friendly format. Guidance using animated characters is particularly important for providing entertainment and convenience to users, but the technology to achieve this is still limited.

[0194] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0195] In this invention, the server includes means for generating guidance voice from an animated character, means for acquiring information about tourist spots or brick-and-mortar stores, means for interactively responding to user questions, means for providing multilingual support, means for a designated character to provide audio guidance about in-store product descriptions and recommendations, and means for a user terminal or robot to interactively respond to product-related questions. This enables multilingual and interactive guidance in tourist spots and brick-and-mortar stores, significantly improving the user experience.

[0196] "Animated character guidance voice" refers to an audio guide in the form of a specific animated character speaking, generated using a computer program.

[0197] A "tourist destination" is a specific place or area visited by tourists, and refers to an area that has historical, cultural, or natural attractions.

[0198] "Brick and mortar store" refers to a store that sells goods or services in a physical location.

[0199] "Means of obtaining information" refers to the functions and methods for collecting the required data and information from databases and other sources.

[0200] "Means for interactive response" refers to a system or function that returns an immediate response to a question or input from a user.

[0201] A "multilingual response means" is a method or system for providing information or responding to questions in multiple languages.

[0202] "Means for providing product descriptions and recommended information by voice" refers to a function or system for providing product descriptions and recommended information to users by voice.

[0203] "User terminal" refers to a computing device used by a user, such as a smartphone or tablet.

[0204] A "robot" refers to an automated mechanical device that operates automatically and performs specific tasks based on user instructions.

[0205] An embodiment of the present invention is described below: This system generates guide voices by animated characters at tourist spots or brick-and-mortar stores, and provides users with a multilingual interactive guide service.

[0206] 1. System Overview

[0207] The system consists of three main components:

[0208] 1.1 User terminal

[0209] The user terminals are personal devices such as smartphones and tablets. Using these terminals, users can select their favorite animated characters and tourist spots or brick-and-mortar stores, and then obtain information through the character's voice guidance. The terminals are connected to a server via the Internet.

[0210] 1.2 Server

[0211] The server plays a central role in the system and has the following functions:

[0212] Character voice generation: Generates guidance voices for designated animated characters. Utilizing a generative AI model, the voices are tailored to the personality of the character selected by the user.

[0213] Information acquisition: Obtain information about tourist spots and brick-and-mortar stores from the database.

[0214] Interactive Response: Generates optimal responses to user questions and translates them into multiple languages ​​as needed.

[0215] Guide plan generation: Generate a travel plan or product introduction plan as needed.

[0216] 1.3 Database

[0217] The database contains detailed information about tourist destinations and brick-and-mortar stores, including their history, culture, product features, and prices. The database is updated regularly.

[0218] 2. Details of the processing

[0219] The server proceeds in the following way:

[0220] 2.1 Character voice generation

[0221] A generative AI model is used to generate a voice for the selected character, for example, using the gTTS (Google® Text-to-Speech) library to generate an audio file to provide voice guidance, which is then sent to the user's device, where the user can play it back.

[0222] 2.2 Information acquisition

[0223] The server retrieves necessary information from a database, such as historical information about tourist spots or product descriptions from physical stores, and generates voice guidance to provide to the user based on that information.

[0224] 2.3 Interactive Response

[0225] When a user enters a question in voice or text format, the server analyzes the question and generates the best possible response using a generative AI model. If necessary, the response is translated for multilingual support and sent to the user's device.

[0226] 3. Specific Examples

[0227] Here are some concrete examples of how the system can be used:

[0228] Example 1: Tourist information

[0229] The user selects a park as a tourist attraction and a specific anime character as a character. The server retrieves information about the park from a database and uses a generative AI model to generate a guide in the character's voice. When the user asks, "Tell me about the history of this park," the server generates an appropriate response to the question and sends it to the user's device. The user can then play the audio guide on their device and tour the park.

[0230] Prompt Sentence Examples

[0231] User: "I chose Pikachu."

[0232] App: "Pikachu's voice will begin prompting. What product would you like to know more about?"

[0233] User: "What product is 12345?"

[0234] App: "This item is of the highest quality. It costs 1000 yen."

[0235] User: "Where does this product originate from?"

[0236] App: "Origin: Japan."

[0237] In this way, users can enjoy being guided around tourist spots and brick-and-mortar stores together with animated characters.

[0238] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0239] Step 1:

[0240] Input: The user launches the application using their device and selects their favorite animated character and tourist destination or brick-and-mortar store.

[0241] Processing: The terminal receives as input information about the characters and tourist spots or brick-and-mortar stores selected by the user, and transmits the selected information to the server.

[0242] Output: The selection information sent to the server.

[0243] Specific operation: The user opens the smartphone app and taps to select a character and destination from the on-screen menu.

[0244] Step 2:

[0245] Input: The server receives the selection information sent by the user.

[0246] Processing: The server retrieves detailed information about tourist attractions or brick-and-mortar stores from the database based on the received selection information.

[0247] Output: Detailed information about a tourist attraction or brick-and-mortar store retrieved from the database.

[0248] What it does: The server queries a database to retrieve information about the destination, such as its history and product descriptions.

[0249] Step 3:

[0250] Input: Details retrieved from the database.

[0251] Processing: The server uses the generative AI model to generate a voice guide for the selected character based on the detailed information.

[0252] Output: Guidance voice data of the generated character.

[0253] Specific operation: Using a generative AI model (e.g., gTTS library), convert text information into the voice of a specified character and create an audio file.

[0254] Step 4:

[0255] Input: Generated guidance speech data.

[0256] Processing: The server transmits the generated voice data to the user terminal.

[0257] Output: Character guidance voice data sent to the user's terminal.

[0258] Specific operation: The server sends the audio file to the user's smartphone via the network.

[0259] Step 5:

[0260] Input: Voice data sent to the user's terminal.

[0261] Processing: The user terminal plays back the received guidance voice.

[0262] Output: The audio guidance that the user can hear.

[0263] Specific operation: The selected character's voice will play instructions through the smartphone speaker.

[0264] Step 6:

[0265] Input: The user types a question into the application using text or voice.

[0266] Processing: The terminal recognizes the user's question and sends it to the server.

[0267] Output: The user's question sent to the server.

[0268] Specific action: The user speaks a question or enters text using the device's microphone.

[0269] Step 7:

[0270] Input: The user's question received by the server.

[0271] Processing: The server analyzes the question and generates the best response using its database and generative AI models, translating it into multiple languages ​​if necessary.

[0272] Output: The generated response data.

[0273] Specific operation: The server analyzes the question using natural language processing technology, generates a corresponding answer, and then translates it into the specified language.

[0274] Step 8:

[0275] Input: The generated response data.

[0276] Processing: The server sends the generated response data to the user terminal.

[0277] Output: The response data sent to the user terminal.

[0278] Specific operation: The server sends response data to the user's smartphone via the network.

[0279] Step 9:

[0280] Input: Response data sent to the user terminal.

[0281] Processing: The user terminal plays back the answer in the character's voice based on the received response data.

[0282] Output: An interactive audio response that the user can hear.

[0283] Specific operation: The smartphone plays the received data as an audio file and provides the answer to the user.

[0284] This allows users to enjoy information and detailed information about tourist spots and brick-and-mortar stores through animated characters.

[0285] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0286] This invention is an interactive pilgrimage tour system using animated characters combined with an emotion engine that recognizes the user's emotions. This system not only provides detailed information about tourist spots, but also provides appropriate guidance and responses based on the user's emotions. It also supports multiple languages, making it suitable for foreign tourists.

[0287] System configuration

[0288] 1. User Device

[0289] Functions: Provides an interface for users to access the system and select characters and tourist attractions. Also plays audio guides for tourist attractions and displays and plays responses to user questions in real time. Furthermore, it has an emotion engine for recognizing emotions from user input (text and voice).

[0290] How it works: Users access a designated website or application using a device such as a smartphone or PC. They select a character and a destination on the interface, and if they have any questions, they input them using text or voice. The emotion engine analyzes the user's input and provides guidance accordingly.

[0291] 2. Server

[0292] Function: Receives character and destination selection information sent from the user's device and retrieves tourist spot information from a database. Uses a generative AI model to generate character voice and text. Additionally, responds interactively to user questions and supports multiple languages. Also, adjusts guidance and response content based on the user's emotional data.

[0293] Processing method: Based on the user's selection, the server retrieves relevant sacred site information (such as name, location, and history) from a database. It then uses a generative AI model to generate audio data for tour guidance and transmits it to the user's device. When the user enters a question, the server analyzes the question, generates an optimal response, and adjusts the response according to the emotion recognized by the emotion engine.

[0294] 3. Database

[0295] Function: Stores and manages detailed information about tourist spots (history, culture, episodes related to animation and manga, etc.). Provides necessary information in response to requests from the server.

[0296] Management method: The database is updated regularly to include new tourist destination information and information on animation and manga.

[0297] Specific examples

[0298] A specific example of the system is shown below.

[0299] Example 1: A park in Tokyo

[0300] 1. User selection: The user opens the smartphone application and selects Character A from the animation work and a certain park in Tokyo.

[0301] 2. Tour Guide: The server receives the selection information and retrieves detailed information about a certain park from a database. Using a generative AI model, it generates a narration in the voice of Character A, saying "This park is...", and sends it to the user's device. The user plays the narration on their device and tours the park.

[0302] 3. Question and Response: The user enters a question, such as "Tell me about the history of this place." The server analyzes the question and uses a generative AI model to generate a response, such as "This park was founded in ____...," which is sent to the device. The device then plays back the response in the voice of Character A.

[0303] 4. Emotion Recognition: When a user types "What a beautiful place," the emotion engine analyzes the emotion data and recognizes that the user is happy. The server generates a response based on the emotion, such as "I'm glad to hear that," and plays it on the device.

[0304] 5. Multilingual support: If the user has English settings, the server will use artificial intelligence to translate the response into English and send it to the device as the English voice of Character A.

[0305] This allows users to have a rich experience touring tourist spots while interacting with animated characters.The purpose of the present invention is to improve the quality of users' sightseeing experiences through such an interactive guide system.

[0306] The processing flow will be explained below.

[0307] Step 1:

[0308] A user starts an application on a terminal and accesses the system. The terminal generates initial session information and sends it to the server.

[0309] Step 2:

[0310] The terminal displays the animated characters and a list of tourist attractions on a user interface.

[0311] Step 3:

[0312] Through the interface, the user selects an animated guide character and the tourist sites they wish to visit.

[0313] Step 4:

[0314] The device sends the user's selection (character ID, destination ID, etc.) to the server.

[0315] Step 5:

[0316] The server retrieves information about the selected character and tourist spot from the database.

[0317] Step 6:

[0318] The server generates a tour plan based on the acquired information and determines the spots to visit and the content of the narration.

[0319] Step 7:

[0320] The server uses the generative AI model to generate narration text for the tour plan using the voice of the selected character, creating audio data.

[0321] Step 8:

[0322] The server transmits the generated voice data to the user's terminal.

[0323] Step 9:

[0324] The terminal starts playing the audio data received from the server as narration.

[0325] Step 10:

[0326] During the tour, the user can ask questions by text or voice, for example, "Tell me about something special about this place."

[0327] Step 11:

[0328] The terminal sends the user's question to the server.

[0329] Step 12:

[0330] The server analyzes the question and retrieves relevant information from a database.

[0331] Step 13:

[0332] The server uses the generative AI model to generate the analyzed information as voice data for the character, for example, "This park was established at ____."

[0333] Step 14:

[0334] The server transmits the generated response voice data to the terminal.

[0335] Step 15:

[0336] The terminal reproduces the response voice data received from the server.

[0337] Step 16:

[0338] When a user speaks, the emotion engine analyzes the user's input (text or voice) and recognizes the emotion. For example, if a user inputs "It's a very beautiful place," the device analyzes it and sends it to the emotion engine.

[0339] Step 17:

[0340] The emotion engine recognizes the user's emotion and sends the emotion data to the server. For example, it recognizes that the user is happy.

[0341] Step 18:

[0342] The server uses a generative AI model to generate an appropriate response based on the emotional data, such as "I'm glad to hear that."

[0343] Step 19:

[0344] The server transmits the generated emotion response voice data to the terminal.

[0345] Step 20:

[0346] The terminal reproduces the emotion response voice data received from the server.

[0347] Step 21:

[0348] The server checks the language specified by the user in the settings, translates the voice data into that language if necessary, and sends the translated voice to the device.

[0349] Step 22:

[0350] The device plays the translated audio data and delivers it to the user.

[0351] Example 2

[0352] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0353] Conventional tourist information systems provide only one-way guidance information, making it difficult to provide responses that take the user's emotions into consideration. Furthermore, they lack multilingual support, making it impossible to provide the same service to foreign tourists. Furthermore, interactive responses and real-time guidance generation are insufficient, leaving room for improvement in the user experience.

[0354] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for generating a guidance voice of an animated character, a means for acquiring information on tourist spots, a means for interactively responding to a user's questions, a means for providing responses in multiple languages, and a means for analyzing the user's emotions and adjusting the response based on the emotions. This makes it possible to provide interactive tourist information that takes the user's emotions into consideration in real time.

[0355] An "animated character" is a computer-generated character, such as a person or animal, that is displayed visually and interacts with the user through sound and movement.

[0356] "Guidance voice" is voice data for providing information on tourist spots and the like by voice, and is used to guide and explain to the user.

[0357] "Tourist destination information" refers to detailed data about a tourist destination, including its location, history, culture, and access methods.

[0358] An "interactive response means" is a device or program that has the function of responding in real time to questions or input from a user.

[0359] "Multilingual support" means providing services and responses in multiple languages ​​to users who speak different languages.

[0360] "Means for analyzing emotions" refers to technology or devices that identify emotions from user input or behavior and acquire them as data.

[0361] A "generative AI model" is an algorithm or program that uses artificial intelligence to generate text or speech.

[0362] A "database" is a system for efficiently storing and managing information, accumulating large amounts of data and allowing it to be quickly searched and retrieved as needed.

[0363] A "user terminal" is a device operated by a user, such as a smartphone, PC, or tablet.

[0364] A "server" is a computer system that receives requests from clients via a network and provides the necessary information.

[0365] This invention is an interactive tourist information system using animated characters combined with an emotion engine that recognizes the user's emotions. This system not only provides detailed information about tourist spots, but also provides appropriate guidance and responses based on the user's emotions. It also supports multiple languages, making it suitable for foreign tourists.

[0366] System configuration

[0367] User terminal

[0368] Users access a designated website or application using a device such as a smartphone or PC. Through this interface, users can select an animated character and a desired tourist destination and enter their request. The user device is equipped with an emotion engine that recognizes emotions from the user's input (text or voice). For example, if a user enters "I want to go to a certain park in Tokyo," the request is sent to the server.

[0369] server

[0370] The server receives the character and tourist attraction selection information sent from the user's device. The server accesses the database to obtain detailed information about the relevant tourist attraction. It then uses a generative AI model to generate audio guidance data. At this time, guidance such as "This park was established in XX..." is generated in the voice of the selected character. This audio data is sent to the user's device and played back.

[0371] Database

[0372] The database stores and manages detailed information about tourist destinations (location, history, culture, etc.). This information is used by the server and provided to the service as needed. The database is also regularly updated with new tourist destination information and information about animations and comics.

[0373] Specific examples

[0374] A specific example of the system's operation is shown below.

[0375] Example 1: A park in Tokyo

[0376] 1. User selection: The user opens the smartphone application and selects animated character A and a certain park in Tokyo.

[0377] 2. Tour Guide: The server receives the selection information and retrieves detailed information about a certain park from a database. Using a generative AI model, it generates a narration in the voice of Character A, saying "This park is...", and sends it to the user's device. The user plays the narration on their device and tours the park.

[0378] 3. Question and Response: The user enters a question such as "Tell me about the history of this place." The server analyzes the question and uses a generative AI model to generate a response such as "This park was founded in ____..." and sends it to the device. The device plays the response in the voice of Character A.

[0379] 4. Emotion Recognition: When a user types "What a beautiful place," the emotion engine analyzes the emotion data and recognizes that the user is happy. The server generates a response based on the emotion, such as "I'm glad to hear that," and plays it on the device.

[0380] 5. Multilingual support: If the user has English settings, the server will translate the response into English using the generative AI model and send it to the device as Character A's English voice to play.

[0381] Prompt Sentence Examples

[0382] 1. "I'd like you to show me around a certain park in Tokyo."

[0383] 2. "Tell me about the history of this park."

[0384] 3. "It's a beautiful place."

[0385] 4. “Can you provide the information in English?”

[0386] Through such a system, users can have a rich experience touring tourist spots while interacting with animated characters. The present invention aims to improve the quality of users' sightseeing experiences.

[0387] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0388] Step 1:

[0389] Users open the application on their smartphone or PC and select an animated character and a tourist destination of their choice. The user's selection and request are entered and sent from the device to the server. Specifically, the user taps or clicks on an option on the screen, and the selection is collected by the application and transmitted to the server via the network.

[0390] Step 2:

[0391] The server receives the character and tourist attraction selection information sent by the user. The server then queries the database to obtain detailed information about the corresponding tourist attraction. Specifically, a database query is executed using the character and tourist attraction identifiers. The query results in data such as the tourist attraction's name, location, history, and culture.

[0392] Step 3:

[0393] The server uses a generative AI model to generate voice guidance data based on the acquired tourist attraction information and user request information. Using the acquired tourist attraction information and request information as input, it processes and calculates data. Character voice guidance data is generated as output. Specifically, the generative AI model creates narration from text information and converts it into voice data.

[0394] Step 4:

[0395] The server transmits the generated voice guidance data to the user terminal. Specifically, a data packet containing the voice guidance data is transmitted to the user terminal via the network. The user terminal then plays back the received voice guidance data.

[0396] Step 5:

[0397] As users travel around tourist spots, they input questions and comments through their devices. The user's input is in the form of text or voice. The user's device then sends the input data to the server. Specifically, when the user inputs a question and taps the send button, the data is collected and transferred to the server.

[0398] Step 6:

[0399] The server receives and analyzes questions from users. It analyzes the question data received as input and generates an optimal response using a generative AI model. The generated response data is obtained as output. Specifically, a natural language processing algorithm analyzes the question and generates an appropriate response.

[0400] Step 7:

[0401] The generated response data is sent from the server to the user terminal. Specifically, a data packet containing the response data is sent to the user terminal via the network. The user terminal then plays back the received response data in the character's voice.

[0402] Step 8:

[0403] The emotion engine analyzes the user's input data (text and voice) and extracts emotion data. It uses the user's comment data as input and performs data analysis. Emotion data is obtained as output. Specifically, text mining and voice analysis algorithms recognize emotions.

[0404] Step 9:

[0405] The server adjusts the response content based on the emotional data. Using the emotional data and generated response data as input, it processes the data again using the generative AI model. The output is response data based on the emotion. Specifically, an algorithm is executed to generate response content that reflects the emotional data.

[0406] Step 10:

[0407] If the user has multilingual settings, the server translates the generated response data into the specified language. Using the generated response data as input, it performs data calculations using an AI translation model. The translated response data is obtained as output. Specifically, the translation algorithm converts the response data into another language.

[0408] In this way, users can experience interactive tourist information that responds to their emotions through animated characters. In addition, the multilingual support makes it easy for foreign tourists to use the service.

[0409] (Application example 2)

[0410] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0411] In recent years, the number of interactive systems that provide information about tourist destinations and restaurants has increased. However, current systems are unable to provide optimal responses based on user emotions. This results in problems such as not being able to provide the information and services users desire in a timely and appropriate manner. Furthermore, the lack of multilingual support makes these systems difficult for foreign tourists to use. To solve these issues, a system that analyzes user emotions and provides optimal recommendations based on those emotions is needed.

[0412] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0413] In this invention, the server includes a means for recognizing a user's emotions and adjusting the response content based on the emotions, a means for recommending optimal restaurants and dishes based on the emotions, and a means for generating recommendation text using a generative AI model, thereby enabling interactive responses according to the user's emotions.

[0414] The "means for generating guidance voice of an animated character" is a function for providing the user with information about tourist spots and restaurants in the voice of an animated character.

[0415] "Means for obtaining information on tourist destinations" refers to a function for collecting detailed data on tourist destinations from databases and external information sources.

[0416] The "means for interactively responding to user questions" is a function for analyzing questions from users and providing appropriate answers in real time.

[0417] "Means for providing responses in multiple languages" is a function for providing answers and guidance adapted to each language in scenarios requiring multilingual support.

[0418] "Means for recognizing the user's emotions and adjusting the response content based on those emotions" is a function for analyzing the emotions from the user's input content and voice and generating the optimal response accordingly.

[0419] The "means for recommending optimal restaurants and dishes based on emotions" is a function for taking into account the emotional state of the user and recommending restaurants and dishes that are suitable for them.

[0420] "Means for generating recommendation text using a generative AI model" is a function that utilizes an AI model to generate recommendation text based on a user's emotions and questions.

[0421] A "terminal" is a device operated by a user, such as a smartphone or tablet.

[0422] A "server" is a central processing unit that processes data from user terminals and generates and transmits necessary information.

[0423] This invention is an interactive food delivery application system that combines an emotion engine that recognizes the user's emotions. The system aims to recommend the most suitable restaurants and dishes to the user by linking the user terminal, server, and database.

[0424] System configuration

[0425] 1. User Device

[0426] Function: Provides an interface for users to access the application and select restaurants and dishes. It also recognizes user emotions from input (text and voice) and has an emotion engine. It also displays and plays recommendations based on emotions.

[0427] How it works: Users access the designated application using a smartphone or other device. They select a restaurant and a dish on the interface, and input their questions or preferences via text or voice. The emotion engine analyzes the emotions from these inputs and makes food and drink recommendations accordingly.

[0428] 2. Server

[0429] Function: Receives sentiment analysis results and question data sent from the user's device and retrieves relevant dining information from the database. Uses a generative AI model to generate recommendation text based on the user's sentiment and question and sends it to the user's device. Also supports multiple languages, providing translated responses into foreign languages.

[0430] Processing method: The server retrieves relevant restaurant information (store information, dish details, reviews, etc.) from a database based on the user's sentiment analysis results and question data. It then uses a generative AI model to generate optimal recommendation text and sends it to the user's device. The recommendation content is adjusted based on the sentiment recognized by the emotion engine.

[0431] 3. Database

[0432] Function: Stores and manages detailed information about restaurants and dishes (menus, prices, nutritional information, user reviews, etc.). Provides necessary information in response to requests from the server.

[0433] How it's maintained: The database is updated regularly to include information on new restaurants and cuisines.

[0434] Specific examples

[0435] A specific example of the system is shown below.

[0436] Example 1: Food and drink recommendations when the user is tired

[0437] 1. User makes a choice: The user opens the application on their smartphone and types "I'm tired."

[0438] 2. Emotion recognition and recommendation: The server receives the user's input, and the emotion engine analyzes the emotion "tired." Using the generative AI model, a recommendation text is generated, such as "We recommend the stamina bowl at nearby Restaurant A, which is a nutritious dish," and sent to the user's device.

[0439] 3. Multilingual support: If the user selects English, the generated recommendation text is translated into English and sent to the device, enabling multilingual support.

[0440] Example prompts for generative AI models

[0441] "What foods would you recommend for users when they're feeling low?"

[0442] "What food can cheer you up when you're feeling tired?"

[0443] In this way, through the system of the present invention, users can receive recommendations for optimal food and drink based on their emotions and enjoy a satisfying eating and drinking experience.

[0444] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0445] Step 1:

[0446] The user launches the smartphone application and inputs their question or request via text or voice.

[0447] Input: User text or voice input (e.g., "I'm tired today and want to eat something energizing.")

[0448] Output: The app gets the user's input data.

[0449] Specific actions: The user inputs their feelings and questions using the dedicated app interface, then presses the "Send" button.

[0450] Step 2:

[0451] The device sends the user's input data to an emotion recognition engine, which analyzes the emotion.

[0452] Input: User text or voice data

[0453] Output: Emotion recognition result (e.g. "tired")

[0454] Specific operation: The app's built-in emotion recognition engine analyzes the user's input and extracts emotions such as "tired" as labels.

[0455] Step 3:

[0456] The device sends the emotion recognition results to the server and requests optimal food and drink recommendations.

[0457] Input: Emotion recognition result (e.g., "tired")

[0458] Output: Send emotion recognition results to the server

[0459] Specific operation: The emotion recognition results are sent to the server, and the app calls the server's API using a network connection.

[0460] Step 4:

[0461] The server retrieves relevant dining information from the database.

[0462] Input: Emotion recognition results to the server

[0463] Output: Related food and drink information (e.g., food items such as "Stamina bowl")

[0464] Specific operation: The server queries the database and obtains food and drink information (e.g., foods that provide stamina) related to the emotion recognition results.

[0465] Step 5:

[0466] The server uses a generative AI model to generate recommendation text based on the user's sentiment and question.

[0467] Input: Food and drink information obtained from the database, emotion recognition results

[0468] Output: Recommendation text (e.g. "You seem tired today. I recommend a stamina bowl to give you energy.")

[0469] Specific operation: The server uses a generative AI model such as OpenAI's GPT-4 (registered trademark), inputs food and drink information and emotion recognition results as prompts, and generates recommendation text.

[0470] Step 6:

[0471] The server transmits the generated recommendation text to the user terminal.

[0472] Input: Recommendation text

[0473] Output: Send recommendation text to user device

[0474] Specific operation: The server sends the generated recommendation text to the user terminal via the network.

[0475] Step 7:

[0476] The terminal displays or plays the recommendation text to the user.

[0477] Input: Recommendation text

[0478] Output: Display to user or play as audio

[0479] What happens: The app displays the suggested text on the screen or plays it aloud using a text-to-speech engine.

[0480] Through these steps, users can receive optimal food and drink recommendations tailored to their emotional state. Using a generative AI model, we can provide more natural and specific recommendation text.

[0481] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0482] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0483] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0484] [Second embodiment]

[0485] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0486] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0487] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0488] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0489] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0490] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0491] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0492] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0493] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0494] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0495] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0496] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0497] This invention is a system that generates guide voices for animated characters and provides real-time information about tourist spots. This system realizes an interactive pilgrimage tour based on the characters and tourist spots selected by the user. It also supports multiple languages, making it suitable for foreign tourists.

[0498] System configuration

[0499] 1. User Device

[0500] Function: Provides an interface for users to access the system and select characters and tourist attractions, plays audio guides for tourist attractions, and displays and plays responses to user questions in real time.

[0501] How it works: The user accesses a designated website or application using a device such as a smartphone or PC, selects a character and a destination on the interface, and uses text or voice input to ask any questions.

[0502] 2. Server

[0503] Function: Receives character and destination selection information sent from the user's device and retrieves tourist spot information from the database. Uses a generative AI model to generate character voice and text. Furthermore, responds interactively to user questions and supports multiple languages.

[0504] Processing method: Based on the user's selection, the server retrieves relevant sacred site information (such as name, location, and history) from the database. It then uses a generative AI model to generate audio data for the tour guide and sends it to the user's device. When the user enters a question, the server analyzes the question and generates the optimal response.

[0505] 3. Database

[0506] Function: Stores and manages detailed information about tourist spots (history, culture, episodes related to animation and manga, etc.). Provides necessary information in response to requests from the server.

[0507] Management method: The database is updated regularly to include new tourist destination information and information on animation and manga.

[0508] Specific examples

[0509] A specific example of the system is shown below.

[0510] Example 1: A park in Tokyo

[0511] 1. User selection: The user opens the smartphone application and selects Character A from the animation work and a certain park in Tokyo.

[0512] 2. Tour Guide: The server receives the selection information and retrieves detailed information about a certain park from a database. Using a generative AI model, it generates a narration in the voice of Character A, saying "This park is...", and sends it to the user's device. The user plays the narration on their device and tours the park.

[0513] 3. Question and Response: The user enters a question, such as "Tell me about the history of this place." The server analyzes the question and uses a generative AI model to generate a response, such as "This park was founded in ____...," which is sent to the device. The device then plays back the response in the voice of Character A.

[0514] 4. Multilingual support: If the user has English settings, the server will use artificial intelligence to translate the response into English and send it to the device as the English voice of Character A.

[0515] This allows users to have a rich experience touring tourist spots while interacting with animated characters.The purpose of the present invention is to improve the quality of users' sightseeing experiences through such an interactive guide system.

[0516] The processing flow will be explained below.

[0517] Step 1:

[0518] A user starts an application on a terminal and accesses the system. The terminal generates initial session information and sends it to the server.

[0519] Step 2:

[0520] The terminal displays the animated characters and a list of tourist spots on a user interface.

[0521] Step 3:

[0522] Through the interface, the user selects an animated guide character and the tourist sites they wish to visit.

[0523] Step 4:

[0524] The terminal transmits the user's selection (character ID, destination ID, etc.) to the server.

[0525] Step 5:

[0526] The server retrieves information about the selected character and tourist spot from a database.

[0527] Step 6:

[0528] The server generates a tour plan based on the acquired information and determines the spots to visit and the content of the narration.

[0529] Step 7:

[0530] The server generates narration text for the tour plan using the voice of the selected character using a generative AI model, and creates audio data.

[0531] Step 8:

[0532] The server transmits the generated voice data to the user's terminal.

[0533] Step 9:

[0534] The terminal starts playing the audio data received from the server as narration.

[0535] Step 10:

[0536] The user asks questions during the tour (e.g., "Tell me about something special about this place") via text or voice input.

[0537] Step 11:

[0538] The terminal sends the user's question to the server.

[0539] Step 12:

[0540] The server analyzes the question and retrieves relevant information from a database.

[0541] Step 13:

[0542] The server uses a generative AI model to generate the analyzed information as voice data for the character.

[0543] Step 14:

[0544] The server transmits the generated response voice data to the terminal.

[0545] Step 15:

[0546] The terminal reproduces the response voice data received from the server.

[0547] Step 16:

[0548] The server checks the language specified by the user in the settings, translates the voice data into that language if necessary, and sends the translated voice to the device.

[0549] Step 17:

[0550] The device plays the translated audio data and delivers it to the user.

[0551] Example 1

[0552] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0553] In the modern tourism industry, tourists are seeking more personalized experiences, but existing guidance systems are unable to meet these demands. Furthermore, information about tourist destinations is often provided in a one-way manner, creating a need for more interactive services. Furthermore, multilingual support is lacking, creating a significant language barrier for foreign tourists in particular.

[0554] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0555] In this invention, the server includes means for a user to select a character and a destination, means for transmitting the selected information to the server, means for the server to obtain destination information from a database, means for generating a voice guide for the character using a generative AI model, means for responding to user questions, and means for translating the generated response into multiple languages. This allows users to receive personalized interactive tourist information, and the multilingual support makes it easy for foreign tourists to obtain information.

[0556] "User character and destination selection means" refers to a device or software that provides an interface for a user to access the system and input or select a particular animated character and tourist destination.

[0557] The "means for transmitting selected information to the server" refers to a communication device or protocol for transmitting information about the character and destination selected by the user to the server.

[0558] "Means for the server to retrieve destination information from the database" refers to the process or software that the server uses to query a database that stores detailed information about tourist destinations based on the user's selection and retrieve the required information.

[0559] "Means for generating voice guidance for a character using a generative AI model" refers to a system or software that uses a generative AI model (e.g., a model using natural language generation technology) to generate voice guidance for a tourist in the voice of a character selected by the user.

[0560] The "means for responding to a user's question" refers to a system or software that analyzes a question entered by a user, generates an appropriate response, and provides that response to the user.

[0561] The "means for translating the generated response into multiple languages" refers to a translation system or software for translating the generated response into a language designated by the user.

[0562] "Means for the terminal to transmit information on the selected character and destination to the server" refers to communication functions or software that allow the user's terminal to transmit information on the selected character and tourist spot to the server.

[0563] "Means for the server to generate a guide plan based on a destination" refers to a system or software that enables the server to generate a detailed guide plan including descriptions of tourist spots and route guidance based on the destination selected by the user.

[0564] "Means for generating character voice using a generative AI model" refers to a system or software that utilizes a generative AI model to generate audio content in the voice of a character selected by a user.

[0565] "Means for providing responses to user questions through a character" refers to a system or software that provides responses to user questions through the voice of a character generated using a generative AI model.

[0566] This invention is a system that uses animated characters to guide tourists around tourist spots, and its features include being interactive and supporting multiple languages. This system is composed of multiple hardware and software components, including a user terminal, a server, a database, and a generative AI model.

[0567] System configuration

[0568] 1. User Device

[0569] Function: Provides an interface for users to access the system and select characters and tourist attractions, plays audio guides for tourist attractions, and responds to user questions in real time.

[0570] Details: Through a smartphone or PC application, users select a character and a destination on the interface, and ask questions by text or voice input. For example, a user opens a smartphone app and selects animated character A and a park in Tokyo.

[0571] 2. Server

[0572] Function: Receives selection information sent from the user's device, retrieves tourist destination information from the database, and generates character voice and text using a generative AI model. Furthermore, it responds interactively to user questions and provides multilingual support as needed.

[0573] Details: Based on the user's selection, the server queries the database for detailed information about related tourist attractions. It then uses a generative AI model (e.g., OpenAI's ChatGPT) to generate audio data and responses for the tour guide and sends them to the user's device.

[0574] 3. Database

[0575] Function: Stores and manages detailed information about tourist spots (such as information about history, culture, and anime). Provides necessary information in response to requests from the server.

[0576] Details: The database is regularly updated with new tourist information and animations. It also contains historical information and specific anecdotes about the tourist destinations.

[0577] Specific examples

[0578] Below is a concrete example of how the system works:

[0579] Example 1: A park in Tokyo

[0580] 1. User selection: The user opens the app on their smartphone and selects Character A from an animation work and a certain park in Tokyo.

[0581] 2. Tour Guide: The server receives the selection information and retrieves detailed information about a certain park from the database. Using the generative AI model, it generates a narration in the voice of Character A, saying "This park is...", and sends it to the user's device.

[0582] 3. Question and Response: The user enters a question such as "Tell me about the history of this place." The server analyzes the question and uses a generative AI model to generate a response such as "This park was founded in ____..." and sends it to the device. The device plays the response in the voice of Character A.

[0583] 4. Multilingual support: If the user has English settings, the server will use artificial intelligence to translate the response into English and send it to the device as the English voice of Character A.

[0584] Prompt Sentence Examples

[0585] "The selected character is Character A, and the tourist destination is a certain park in Tokyo. Please generate audio of a tourist guide about this park."

[0586] The purpose of this invention is to improve the user's sightseeing experience through such an interactive guide system. In addition, by supporting multiple languages, it can also accommodate foreign tourists, thereby meeting global tourism demand.

[0587] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0588] Step 1:

[0589] The user selects a character and a destination

[0590] Specific operation: The user opens the application on their smartphone or computer and selects Character A and a certain park in Tokyo on the interface.

[0591] Input: User selection of character and tourist spot.

[0592] Output: Information about the selected character and tourist spot (e.g., Character A, a certain park).

[0593] Step 2:

[0594] The device sends the selection information to the server

[0595] Specific operation: The device sends JSON data containing information about the selected character and tourist spot to the server as an HTTP request.

[0596] Input: Selected character and tourist attraction information.

[0597] Output: The selection information sent to the server.

[0598] Step 3:

[0599] The server retrieves tourist information from the database

[0600] Specific operation: The server sends a query to the database to obtain information about "a certain park," such as its history and characteristics.

[0601] Input: User selection information (Character A, a certain park).

[0602] Output: Tourist information retrieved from the database (e.g. park history, main attractions, etc.).

[0603] Step 4:

[0604] The server inputs the prompt sentence into the AI ​​model and generates the guidance voice.

[0605] Specific operation: Based on the tourist attraction information obtained by the server, the prompt sentence "The selected character is Character A, and the tourist attraction is a certain park in Tokyo. Please generate audio tourist information about this park" is input into the generation AI model.

[0606] Input: Obtained tourist attraction information and prompt sentence.

[0607] Output: Guidance speech data generated by the generative AI model.

[0608] Step 5:

[0609] Send the generated audio data to the device

[0610] Specific operation: The server sends the generated voice data to the user's terminal.

[0611] Input: Guidance speech data generated by a generative AI model.

[0612] Output: The audio data sent to the device.

[0613] Step 6:

[0614] The device plays a voice prompt to the user.

[0615] Specific behavior: The device plays the audio data received and lets the user listen to it. Uses the audio player in the app.

[0616] Input: Audio data sent to the device.

[0617] Output: The audio guidance the user hears.

[0618] Step 7:

[0619] The user enters a question

[0620] What happens: A user types a question into a text field in the app: "Tell me about the history of this place."

[0621] Input: The question entered by the user.

[0622] Output: The data from the question entered.

[0623] Step 8:

[0624] The device sends a question to the server

[0625] Specific operation: The device sends JSON data containing the question to the server as an HTTP request.

[0626] Input: The question data entered by the user.

[0627] Output: The query data sent to the server.

[0628] Step 9:

[0629] The server analyzes the question and inputs the prompt into the generative AI model to generate a response.

[0630] Specific operation: The server analyzes the question and inputs the prompt sentence, "The user asked me, 'Tell me about the history of this place.' Please generate an answer to this question." into the generative AI model.

[0631] Input: User question data and generated prompt text.

[0632] Output: The response data generated by the generative AI model.

[0633] Step 10:

[0634] The server sends the response data to the terminal.

[0635] Specific operation: The server sends the generated response data to the terminal.

[0636] Input: The response data generated by the generative AI model.

[0637] Output: The response data sent to the device.

[0638] Step 11:

[0639] The terminal plays the response to the user

[0640] Specific behavior: The device plays back the response data received and lets the user listen to it. Uses the in-app audio player.

[0641] Input: Response data sent to the terminal.

[0642] Output: The response audio that the user hears.

[0643] Step 12:

[0644] If necessary, the server translates the response for multilingual support.

[0645] Specific operation: The server checks the user's language setting and translates the generated response if necessary. For example, if the user's language setting is English, the server translates the response into English and generates English speech data using the generative AI model.

[0646] Input: The user's preferred language and the generated response data.

[0647] Output: The translated response audio data.

[0648] (Application example 1)

[0649] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0650] Currently, guidance at tourist spots and brick-and-mortar stores requires human intervention, and it is difficult to provide multilingual and interactive guidance. Furthermore, when users want to obtain detailed information on their smartphones or in-store terminals, there is a lack of immediate guidance in a user-friendly format. Guidance using animated characters is particularly important for providing entertainment and convenience to users, but the technology to achieve this is still limited.

[0651] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0652] In this invention, the server includes means for generating guidance voice from an animated character, means for acquiring information about tourist spots or brick-and-mortar stores, means for interactively responding to user questions, means for providing multilingual support, means for a designated character to provide audio guidance about in-store product descriptions and recommendations, and means for a user terminal or robot to interactively respond to product-related questions. This enables multilingual and interactive guidance in tourist spots and brick-and-mortar stores, significantly improving the user experience.

[0653] "Animated character guidance voice" refers to an audio guide in the form of a specific animated character speaking, generated using a computer program.

[0654] A "tourist destination" is a specific place or area visited by tourists, and refers to an area that has historical, cultural, or natural attractions.

[0655] "Brick and mortar store" refers to a store that sells goods or services in a physical location.

[0656] "Means of obtaining information" refers to the functions and methods for collecting the required data and information from databases and other sources.

[0657] "Means for interactive response" refers to a system or function that returns an immediate response to a question or input from a user.

[0658] A "multilingual response means" is a method or system for providing information or responding to questions in multiple languages.

[0659] "Means for providing product descriptions and recommended information by voice" refers to a function or system for providing product descriptions and recommended information to users by voice.

[0660] "User terminal" refers to a computing device used by a user, such as a smartphone or tablet.

[0661] A "robot" refers to an automated mechanical device that operates automatically and performs specific tasks based on user instructions.

[0662] An embodiment of the present invention is described below: This system generates guide voices by animated characters at tourist spots or brick-and-mortar stores, and provides users with a multilingual interactive guide service.

[0663] 1. System Overview

[0664] The system consists of three main components:

[0665] 1.1 User terminal

[0666] The user terminals are personal devices such as smartphones and tablets. Using these terminals, users can select their favorite animated characters and tourist spots or brick-and-mortar stores, and then obtain information through the character's voice guidance. The terminals are connected to a server via the Internet.

[0667] 1.2 Server

[0668] The server plays a central role in the system and has the following functions:

[0669] Character voice generation: Generates guidance voices for designated animated characters. Utilizing a generative AI model, the voices are tailored to the personality of the character selected by the user.

[0670] Information acquisition: Obtain information about tourist spots and brick-and-mortar stores from the database.

[0671] Interactive Response: Generates optimal responses to user questions and translates them into multiple languages ​​as needed.

[0672] Guide plan generation: Generate a travel plan or product introduction plan as needed.

[0673] 1.3 Database

[0674] The database contains detailed information about tourist destinations and brick-and-mortar stores, including their history, culture, product features, and prices. The database is updated regularly.

[0675] 2. Details of the processing

[0676] The server proceeds in the following way:

[0677] 2.1 Character voice generation

[0678] A generative AI model is used to generate a voice for the selected character. For example, the gTTS (Google Text-to-Speech) library is used to generate an audio file to provide voice guidance. This audio file is sent to the user's device, where it can be played back.

[0679] 2.2 Information acquisition

[0680] The server retrieves necessary information from a database, such as historical information about tourist spots or product descriptions from physical stores, and generates voice guidance to provide to the user based on that information.

[0681] 2.3 Interactive Response

[0682] When a user enters a question in voice or text format, the server analyzes the question and generates the best possible response using a generative AI model. If necessary, the response is translated for multilingual support and sent to the user's device.

[0683] 3. Specific Examples

[0684] Here are some concrete examples of how the system can be used:

[0685] Example 1: Tourist information

[0686] The user selects a park as a tourist attraction and a specific anime character as a character. The server retrieves information about the park from a database and uses a generative AI model to generate a guide in the character's voice. When the user asks, "Tell me about the history of this park," the server generates an appropriate response to the question and sends it to the user's device. The user can then play the audio guide on their device and tour the park.

[0687] Prompt Sentence Examples

[0688] User: "I chose Pikachu."

[0689] App: "Pikachu's voice will begin prompting. What product would you like to know more about?"

[0690] User: "What product is 12345?"

[0691] App: "This item is of the highest quality. It costs 1000 yen."

[0692] User: "Where does this product originate from?"

[0693] App: "Origin: Japan."

[0694] In this way, users can enjoy being guided around tourist spots and brick-and-mortar stores together with animated characters.

[0695] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0696] Step 1:

[0697] Input: The user launches the application using their device and selects their favorite animated character and tourist destination or brick-and-mortar store.

[0698] Processing: The terminal receives as input information about the characters and tourist spots or brick-and-mortar stores selected by the user, and transmits the selected information to the server.

[0699] Output: The selection information sent to the server.

[0700] Specific operation: The user opens the smartphone app and taps to select a character and destination from the on-screen menu.

[0701] Step 2:

[0702] Input: The server receives the selection information sent by the user.

[0703] Processing: The server retrieves detailed information about tourist attractions or brick-and-mortar stores from the database based on the received selection information.

[0704] Output: Detailed information about a tourist attraction or brick-and-mortar store retrieved from the database.

[0705] What it does: The server queries a database to retrieve information about the destination, such as its history and product descriptions.

[0706] Step 3:

[0707] Input: Details retrieved from the database.

[0708] Processing: The server uses the generative AI model to generate a voice guide for the selected character based on the detailed information.

[0709] Output: Guidance voice data of the generated character.

[0710] Specific operation: Using a generative AI model (e.g., gTTS library), convert text information into the voice of a specified character and create an audio file.

[0711] Step 4:

[0712] Input: Generated guidance speech data.

[0713] Processing: The server transmits the generated voice data to the user terminal.

[0714] Output: Character guidance voice data sent to the user's terminal.

[0715] Specific operation: The server sends the audio file to the user's smartphone via the network.

[0716] Step 5:

[0717] Input: Voice data sent to the user's terminal.

[0718] Processing: The user terminal plays back the received guidance voice.

[0719] Output: The audio guidance that the user can hear.

[0720] Specific operation: The selected character's voice will play instructions through the smartphone speaker.

[0721] Step 6:

[0722] Input: The user types a question into the application using text or voice.

[0723] Processing: The terminal recognizes the user's question and sends it to the server.

[0724] Output: The user's question sent to the server.

[0725] Specific action: The user speaks a question or enters text using the device's microphone.

[0726] Step 7:

[0727] Input: The user's question received by the server.

[0728] Processing: The server analyzes the question and generates the best response using its database and generative AI models, translating it into multiple languages ​​if necessary.

[0729] Output: The generated response data.

[0730] Specific operation: The server analyzes the question using natural language processing technology, generates a corresponding answer, and then translates it into the specified language.

[0731] Step 8:

[0732] Input: The generated response data.

[0733] Processing: The server sends the generated response data to the user terminal.

[0734] Output: The response data sent to the user terminal.

[0735] Specific operation: The server sends response data to the user's smartphone via the network.

[0736] Step 9:

[0737] Input: Response data sent to the user terminal.

[0738] Processing: The user terminal plays back the answer in the character's voice based on the received response data.

[0739] Output: An interactive audio response that the user can hear.

[0740] Specific operation: The smartphone plays the received data as an audio file and provides the answer to the user.

[0741] This allows users to enjoy information and detailed information about tourist spots and brick-and-mortar stores through animated characters.

[0742] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0743] This invention is an interactive pilgrimage tour system using animated characters combined with an emotion engine that recognizes the user's emotions. This system not only provides detailed information about tourist spots, but also provides appropriate guidance and responses based on the user's emotions. It also supports multiple languages, making it suitable for foreign tourists.

[0744] System configuration

[0745] 1. User Device

[0746] Functions: Provides an interface for users to access the system and select characters and tourist attractions. Also plays audio guides for tourist attractions and displays and plays responses to user questions in real time. Furthermore, it has an emotion engine for recognizing emotions from user input (text and voice).

[0747] How it works: Users access a designated website or application using a device such as a smartphone or PC. They select a character and a destination on the interface, and if they have any questions, they input them using text or voice. The emotion engine analyzes the user's input and provides guidance accordingly.

[0748] 2. Server

[0749] Function: Receives character and destination selection information sent from the user's device and retrieves tourist spot information from a database. Uses a generative AI model to generate character voice and text. Additionally, responds interactively to user questions and supports multiple languages. Also, adjusts guidance and response content based on the user's emotional data.

[0750] Processing method: Based on the user's selection, the server retrieves relevant sacred site information (such as name, location, and history) from a database. It then uses a generative AI model to generate audio data for tour guidance and transmits it to the user's device. When the user enters a question, the server analyzes the question, generates an optimal response, and adjusts the response according to the emotion recognized by the emotion engine.

[0751] 3. Database

[0752] Function: Stores and manages detailed information about tourist spots (history, culture, episodes related to animation and manga, etc.). Provides necessary information in response to requests from the server.

[0753] Management method: The database is updated regularly to include new tourist destination information and information on animation and manga.

[0754] Specific examples

[0755] A specific example of the system is shown below.

[0756] Example 1: A park in Tokyo

[0757] 1. User selection: The user opens the smartphone application and selects Character A from the animation work and a certain park in Tokyo.

[0758] 2. Tour Guide: The server receives the selection information and retrieves detailed information about a certain park from a database. Using a generative AI model, it generates a narration in the voice of Character A, saying "This park is...", and sends it to the user's device. The user plays the narration on their device and tours the park.

[0759] 3. Question and Response: The user enters a question, such as "Tell me about the history of this place." The server analyzes the question and uses a generative AI model to generate a response, such as "This park was founded in ____...," which is sent to the device. The device then plays back the response in the voice of Character A.

[0760] 4. Emotion Recognition: When a user types "What a beautiful place," the emotion engine analyzes the emotion data and recognizes that the user is happy. The server generates a response based on the emotion, such as "I'm glad to hear that," and plays it on the device.

[0761] 5. Multilingual support: If the user has English settings, the server will use artificial intelligence to translate the response into English and send it to the device as the English voice of Character A.

[0762] This allows users to have a rich experience touring tourist spots while interacting with animated characters.The purpose of the present invention is to improve the quality of users' sightseeing experiences through such an interactive guide system.

[0763] The processing flow will be explained below.

[0764] Step 1:

[0765] A user starts an application on a terminal and accesses the system. The terminal generates initial session information and sends it to the server.

[0766] Step 2:

[0767] The terminal displays the animated characters and a list of tourist attractions on a user interface.

[0768] Step 3:

[0769] Through the interface, the user selects an animated guide character and the tourist sites they wish to visit.

[0770] Step 4:

[0771] The device sends the user's selection (character ID, destination ID, etc.) to the server.

[0772] Step 5:

[0773] The server retrieves information about the selected character and tourist spot from the database.

[0774] Step 6:

[0775] The server generates a tour plan based on the acquired information and determines the spots to visit and the content of the narration.

[0776] Step 7:

[0777] The server uses the generative AI model to generate narration text for the tour plan using the voice of the selected character, creating audio data.

[0778] Step 8:

[0779] The server transmits the generated voice data to the user's terminal.

[0780] Step 9:

[0781] The terminal starts playing the audio data received from the server as narration.

[0782] Step 10:

[0783] During the tour, the user can ask questions by text or voice, for example, "Tell me about something special about this place."

[0784] Step 11:

[0785] The terminal sends the user's question to the server.

[0786] Step 12:

[0787] The server analyzes the question and retrieves relevant information from a database.

[0788] Step 13:

[0789] The server uses the generative AI model to generate the analyzed information as voice data for the character, for example, "This park was established at ____."

[0790] Step 14:

[0791] The server transmits the generated response voice data to the terminal.

[0792] Step 15:

[0793] The terminal reproduces the response voice data received from the server.

[0794] Step 16:

[0795] When a user speaks, the emotion engine analyzes the user's input (text or voice) and recognizes the emotion. For example, if a user inputs "It's a very beautiful place," the device analyzes it and sends it to the emotion engine.

[0796] Step 17:

[0797] The emotion engine recognizes the user's emotion and sends the emotion data to the server. For example, it recognizes that the user is happy.

[0798] Step 18:

[0799] The server uses a generative AI model to generate an appropriate response based on the emotional data, such as "I'm glad to hear that."

[0800] Step 19:

[0801] The server transmits the generated emotion response voice data to the terminal.

[0802] Step 20:

[0803] The terminal reproduces the emotion response voice data received from the server.

[0804] Step 21:

[0805] The server checks the language specified by the user in the settings, translates the voice data into that language if necessary, and sends the translated voice to the device.

[0806] Step 22:

[0807] The device plays the translated audio data and delivers it to the user.

[0808] Example 2

[0809] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0810] Conventional tourist information systems provide only one-way guidance information, making it difficult to provide responses that take the user's emotions into consideration. Furthermore, they lack multilingual support, making it impossible to provide the same service to foreign tourists. Furthermore, interactive responses and real-time guidance generation are insufficient, leaving room for improvement in the user experience.

[0811] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for generating a guidance voice of an animated character, a means for acquiring information on tourist spots, a means for interactively responding to a user's questions, a means for providing responses in multiple languages, and a means for analyzing the user's emotions and adjusting the response based on the emotions. This makes it possible to provide interactive tourist information that takes the user's emotions into consideration in real time.

[0812] An "animated character" is a computer-generated character, such as a person or animal, that is displayed visually and interacts with the user through sound and movement.

[0813] "Guidance voice" is voice data for providing information on tourist spots and the like by voice, and is used to guide and explain to the user.

[0814] "Tourist destination information" refers to detailed data about a tourist destination, including its location, history, culture, and access methods.

[0815] An "interactive response means" is a device or program that has the function of responding in real time to questions or input from a user.

[0816] "Multilingual support" means providing services and responses in multiple languages ​​to users who speak different languages.

[0817] "Means for analyzing emotions" refers to technology or devices that identify emotions from user input or behavior and acquire them as data.

[0818] A "generative AI model" is an algorithm or program that uses artificial intelligence to generate text or speech.

[0819] A "database" is a system for efficiently storing and managing information, accumulating large amounts of data and allowing it to be quickly searched and retrieved as needed.

[0820] A "user terminal" is a device operated by a user, such as a smartphone, PC, or tablet.

[0821] A "server" is a computer system that receives requests from clients via a network and provides the necessary information.

[0822] This invention is an interactive tourist information system using animated characters combined with an emotion engine that recognizes the user's emotions. This system not only provides detailed information about tourist spots, but also provides appropriate guidance and responses based on the user's emotions. It also supports multiple languages, making it suitable for foreign tourists.

[0823] System configuration

[0824] User terminal

[0825] Users access a designated website or application using a device such as a smartphone or PC. Through this interface, users can select an animated character and a desired tourist destination and enter their request. The user device is equipped with an emotion engine that recognizes emotions from the user's input (text or voice). For example, if a user enters "I want to go to a certain park in Tokyo," the request is sent to the server.

[0826] server

[0827] The server receives the character and tourist attraction selection information sent from the user's device. The server accesses the database to obtain detailed information about the relevant tourist attraction. It then uses a generative AI model to generate audio guidance data. At this time, guidance such as "This park was established in XX..." is generated in the voice of the selected character. This audio data is sent to the user's device and played back.

[0828] Database

[0829] The database stores and manages detailed information about tourist destinations (location, history, culture, etc.). This information is used by the server and provided to the service as needed. The database is also regularly updated with new tourist destination information and information about animations and comics.

[0830] Specific examples

[0831] A specific example of the system's operation is shown below.

[0832] Example 1: A park in Tokyo

[0833] 1. User selection: The user opens the smartphone application and selects animated character A and a certain park in Tokyo.

[0834] 2. Tour Guide: The server receives the selection information and retrieves detailed information about a certain park from a database. Using a generative AI model, it generates a narration in the voice of Character A, saying "This park is...", and sends it to the user's device. The user plays the narration on their device and tours the park.

[0835] 3. Question and Response: The user enters a question such as "Tell me about the history of this place." The server analyzes the question and uses a generative AI model to generate a response such as "This park was founded in ____..." and sends it to the device. The device plays the response in the voice of Character A.

[0836] 4. Emotion Recognition: When a user types "What a beautiful place," the emotion engine analyzes the emotion data and recognizes that the user is happy. The server generates a response based on the emotion, such as "I'm glad to hear that," and plays it on the device.

[0837] 5. Multilingual support: If the user has English settings, the server will translate the response into English using the generative AI model and send it to the device as Character A's English voice to play.

[0838] Prompt Sentence Examples

[0839] 1. "I'd like you to show me around a certain park in Tokyo."

[0840] 2. "Tell me about the history of this park."

[0841] 3. "It's a beautiful place."

[0842] 4. “Can you provide the information in English?”

[0843] Through such a system, users can have a rich experience touring tourist spots while interacting with animated characters. The present invention aims to improve the quality of users' sightseeing experiences.

[0844] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0845] Step 1:

[0846] Users open the application on their smartphone or PC and select an animated character and a tourist destination of their choice. The user's selection and request are entered and sent from the device to the server. Specifically, the user taps or clicks on an option on the screen, and the selection is collected by the application and transmitted to the server via the network.

[0847] Step 2:

[0848] The server receives the character and tourist attraction selection information sent by the user. The server then queries the database to obtain detailed information about the corresponding tourist attraction. Specifically, a database query is executed using the character and tourist attraction identifiers. The query results in data such as the tourist attraction's name, location, history, and culture.

[0849] Step 3:

[0850] The server uses a generative AI model to generate voice guidance data based on the acquired tourist attraction information and user request information. Using the acquired tourist attraction information and request information as input, it processes and calculates data. Character voice guidance data is generated as output. Specifically, the generative AI model creates narration from text information and converts it into voice data.

[0851] Step 4:

[0852] The server transmits the generated voice guidance data to the user terminal. Specifically, a data packet containing the voice guidance data is transmitted to the user terminal via the network. The user terminal then plays back the received voice guidance data.

[0853] Step 5:

[0854] As users travel around tourist spots, they input questions and comments through their devices. The user's input is in the form of text or voice. The user's device then sends the input data to the server. Specifically, when the user inputs a question and taps the send button, the data is collected and transferred to the server.

[0855] Step 6:

[0856] The server receives and analyzes questions from users. It analyzes the question data received as input and generates an optimal response using a generative AI model. The generated response data is obtained as output. Specifically, a natural language processing algorithm analyzes the question and generates an appropriate response.

[0857] Step 7:

[0858] The generated response data is sent from the server to the user terminal. Specifically, a data packet containing the response data is sent to the user terminal via the network. The user terminal then plays back the received response data in the character's voice.

[0859] Step 8:

[0860] The emotion engine analyzes the user's input data (text and voice) and extracts emotion data. It uses the user's comment data as input and performs data analysis. Emotion data is obtained as output. Specifically, text mining and voice analysis algorithms recognize emotions.

[0861] Step 9:

[0862] The server adjusts the response content based on the emotional data. Using the emotional data and generated response data as input, it processes the data again using the generative AI model. The output is response data based on the emotion. Specifically, an algorithm is executed to generate response content that reflects the emotional data.

[0863] Step 10:

[0864] If the user has multilingual settings, the server translates the generated response data into the specified language. Using the generated response data as input, it performs data calculations using an AI translation model. The translated response data is obtained as output. Specifically, the translation algorithm converts the response data into another language.

[0865] In this way, users can experience interactive tourist information that responds to their emotions through animated characters. In addition, the multilingual support makes it easy for foreign tourists to use the service.

[0866] (Application example 2)

[0867] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0868] In recent years, the number of interactive systems that provide information about tourist destinations and restaurants has increased. However, current systems are unable to provide optimal responses based on user emotions. This results in problems such as not being able to provide the information and services users desire in a timely and appropriate manner. Furthermore, the lack of multilingual support makes these systems difficult for foreign tourists to use. To solve these issues, a system that analyzes user emotions and provides optimal recommendations based on those emotions is needed.

[0869] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0870] In this invention, the server includes a means for recognizing a user's emotions and adjusting the response content based on the emotions, a means for recommending optimal restaurants and dishes based on the emotions, and a means for generating recommendation text using a generative AI model, thereby enabling interactive responses according to the user's emotions.

[0871] The "means for generating guidance voice of an animated character" is a function for providing the user with information about tourist spots and restaurants in the voice of an animated character.

[0872] "Means for obtaining information on tourist destinations" refers to a function for collecting detailed data on tourist destinations from databases and external information sources.

[0873] The "means for interactively responding to user questions" is a function for analyzing questions from users and providing appropriate answers in real time.

[0874] "Means for providing responses in multiple languages" is a function for providing answers and guidance adapted to each language in scenarios requiring multilingual support.

[0875] "Means for recognizing the user's emotions and adjusting the response content based on those emotions" is a function for analyzing the emotions from the user's input content and voice and generating the optimal response accordingly.

[0876] The "means for recommending optimal restaurants and dishes based on emotions" is a function for taking into account the emotional state of the user and recommending restaurants and dishes that are suitable for them.

[0877] "Means for generating recommendation text using a generative AI model" is a function that utilizes an AI model to generate recommendation text based on a user's emotions and questions.

[0878] A "terminal" is a device operated by a user, such as a smartphone or tablet.

[0879] A "server" is a central processing unit that processes data from user terminals and generates and transmits necessary information.

[0880] This invention is an interactive food delivery application system that combines an emotion engine that recognizes the user's emotions. The system aims to recommend the most suitable restaurants and dishes to the user by linking the user terminal, server, and database.

[0881] System configuration

[0882] 1. User Device

[0883] Function: Provides an interface for users to access the application and select restaurants and dishes. It also recognizes user emotions from input (text and voice) and has an emotion engine. It also displays and plays recommendations based on emotions.

[0884] How it works: Users access the designated application using a smartphone or other device. They select a restaurant and a dish on the interface, and input their questions or preferences via text or voice. The emotion engine analyzes the emotions from these inputs and makes food and drink recommendations accordingly.

[0885] 2. Server

[0886] Function: Receives sentiment analysis results and question data sent from the user's device and retrieves relevant dining information from the database. Uses a generative AI model to generate recommendation text based on the user's sentiment and question and sends it to the user's device. Also supports multiple languages, providing translated responses into foreign languages.

[0887] Processing method: The server retrieves relevant restaurant information (store information, dish details, reviews, etc.) from a database based on the user's sentiment analysis results and question data. It then uses a generative AI model to generate optimal recommendation text and sends it to the user's device. The recommendation content is adjusted based on the sentiment recognized by the emotion engine.

[0888] 3. Database

[0889] Function: Stores and manages detailed information about restaurants and dishes (menus, prices, nutritional information, user reviews, etc.). Provides necessary information in response to requests from the server.

[0890] How it's maintained: The database is updated regularly to include information on new restaurants and cuisines.

[0891] Specific examples

[0892] A specific example of the system is shown below.

[0893] Example 1: Food and drink recommendations when the user is tired

[0894] 1. User makes a choice: The user opens the application on their smartphone and types "I'm tired."

[0895] 2. Emotion recognition and recommendation: The server receives the user's input, and the emotion engine analyzes the emotion "tired." Using the generative AI model, a recommendation text is generated, such as "We recommend the stamina bowl at nearby Restaurant A, which is a nutritious dish," and sent to the user's device.

[0896] 3. Multilingual support: If the user selects English, the generated recommendation text is translated into English and sent to the device, enabling multilingual support.

[0897] Example prompts for generative AI models

[0898] "What foods would you recommend for users when they're feeling low?"

[0899] "What food can cheer you up when you're feeling tired?"

[0900] In this way, through the system of the present invention, users can receive recommendations for optimal food and drink based on their emotions and enjoy a satisfying eating and drinking experience.

[0901] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0902] Step 1:

[0903] The user launches the smartphone application and inputs their question or request via text or voice.

[0904] Input: User text or voice input (e.g., "I'm tired today and want to eat something energizing.")

[0905] Output: The app gets the user's input data.

[0906] Specific actions: The user inputs their feelings and questions using the dedicated app interface, then presses the "Send" button.

[0907] Step 2:

[0908] The device sends the user's input data to an emotion recognition engine, which analyzes the emotion.

[0909] Input: User text or voice data

[0910] Output: Emotion recognition result (e.g. "tired")

[0911] Specific operation: The app's built-in emotion recognition engine analyzes the user's input and extracts emotions such as "tired" as labels.

[0912] Step 3:

[0913] The device sends the emotion recognition results to the server and requests optimal food and drink recommendations.

[0914] Input: Emotion recognition result (e.g., "tired")

[0915] Output: Send emotion recognition results to the server

[0916] Specific operation: The emotion recognition results are sent to the server, and the app calls the server's API using a network connection.

[0917] Step 4:

[0918] The server retrieves relevant dining information from the database.

[0919] Input: Emotion recognition results to the server

[0920] Output: Related food and drink information (e.g., food items such as "Stamina bowl")

[0921] Specific operation: The server queries the database and obtains food and drink information (e.g., foods that provide stamina) related to the emotion recognition results.

[0922] Step 5:

[0923] The server uses a generative AI model to generate recommendation text based on the user's sentiment and question.

[0924] Input: Food and drink information obtained from the database, emotion recognition results

[0925] Output: Recommendation text (e.g. "You seem tired today. I recommend a stamina bowl to give you energy.")

[0926] Specific operation: The server uses a generative AI model such as OpenAI's GPT-4, inputs food and drink information and emotion recognition results as prompts, and generates recommendation text.

[0927] Step 6:

[0928] The server transmits the generated recommendation text to the user terminal.

[0929] Input: Recommendation text

[0930] Output: Send recommendation text to user device

[0931] Specific operation: The server sends the generated recommendation text to the user terminal via the network.

[0932] Step 7:

[0933] The terminal displays or plays the recommendation text to the user.

[0934] Input: Recommendation text

[0935] Output: Display to user or play as audio

[0936] What happens: The app displays the suggested text on the screen or plays it aloud using a text-to-speech engine.

[0937] Through these steps, users can receive optimal food and drink recommendations tailored to their emotional state. Using a generative AI model, we can provide more natural and specific recommendation text.

[0938] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0939] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0940] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0941] [Third embodiment]

[0942] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0943] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0944] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0945] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0946] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0947] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0948] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0949] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0950] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0951] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0952] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0953] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0954] This invention is a system that generates guide voices for animated characters and provides real-time information about tourist spots. This system realizes an interactive pilgrimage tour based on the characters and tourist spots selected by the user. It also supports multiple languages, making it suitable for foreign tourists.

[0955] System configuration

[0956] 1. User Device

[0957] Function: Provides an interface for users to access the system and select characters and tourist attractions, plays audio guides for tourist attractions, and displays and plays responses to user questions in real time.

[0958] How it works: The user accesses a designated website or application using a device such as a smartphone or PC, selects a character and a destination on the interface, and uses text or voice input to ask any questions.

[0959] 2. Server

[0960] Function: Receives character and destination selection information sent from the user's device and retrieves tourist spot information from the database. Uses a generative AI model to generate character voice and text. Furthermore, responds interactively to user questions and supports multiple languages.

[0961] Processing method: Based on the user's selection, the server retrieves relevant sacred site information (such as name, location, and history) from the database. It then uses a generative AI model to generate audio data for the tour guide and sends it to the user's device. When the user enters a question, the server analyzes the question and generates the optimal response.

[0962] 3. Database

[0963] Function: Stores and manages detailed information about tourist spots (history, culture, episodes related to animation and manga, etc.). Provides necessary information in response to requests from the server.

[0964] Management method: The database is updated regularly to include new tourist destination information and information on animation and manga.

[0965] Specific examples

[0966] A specific example of the system is shown below.

[0967] Example 1: A park in Tokyo

[0968] 1. User selection: The user opens the smartphone application and selects Character A from the animation work and a certain park in Tokyo.

[0969] 2. Tour Guide: The server receives the selection information and retrieves detailed information about a certain park from a database. Using a generative AI model, it generates a narration in the voice of Character A, saying "This park is...", and sends it to the user's device. The user plays the narration on their device and tours the park.

[0970] 3. Question and Response: The user enters a question, such as "Tell me about the history of this place." The server analyzes the question and uses a generative AI model to generate a response, such as "This park was founded in ____...," which is sent to the device. The device then plays back the response in the voice of Character A.

[0971] 4. Multilingual support: If the user has English settings, the server will use artificial intelligence to translate the response into English and send it to the device as the English voice of Character A.

[0972] This allows users to have a rich experience touring tourist spots while interacting with animated characters.The purpose of the present invention is to improve the quality of users' sightseeing experiences through such an interactive guide system.

[0973] The processing flow will be explained below.

[0974] Step 1:

[0975] A user starts an application on a terminal and accesses the system. The terminal generates initial session information and sends it to the server.

[0976] Step 2:

[0977] The terminal displays the animated characters and a list of tourist spots on a user interface.

[0978] Step 3:

[0979] Through the interface, the user selects an animated guide character and the tourist sites they wish to visit.

[0980] Step 4:

[0981] The terminal transmits the user's selection (character ID, destination ID, etc.) to the server.

[0982] Step 5:

[0983] The server retrieves information about the selected character and tourist spot from a database.

[0984] Step 6:

[0985] The server generates a tour plan based on the acquired information and determines the spots to visit and the content of the narration.

[0986] Step 7:

[0987] The server generates narration text for the tour plan using the voice of the selected character using a generative AI model, and creates audio data.

[0988] Step 8:

[0989] The server transmits the generated voice data to the user's terminal.

[0990] Step 9:

[0991] The terminal starts playing the audio data received from the server as narration.

[0992] Step 10:

[0993] The user asks questions during the tour (e.g., "Tell me about something special about this place") via text or voice input.

[0994] Step 11:

[0995] The terminal sends the user's question to the server.

[0996] Step 12:

[0997] The server analyzes the question and retrieves relevant information from a database.

[0998] Step 13:

[0999] The server uses a generative AI model to generate the analyzed information as voice data for the character.

[1000] Step 14:

[1001] The server transmits the generated response voice data to the terminal.

[1002] Step 15:

[1003] The terminal reproduces the response voice data received from the server.

[1004] Step 16:

[1005] The server checks the language specified by the user in the settings, translates the voice data into that language if necessary, and sends the translated voice to the device.

[1006] Step 17:

[1007] The device plays the translated audio data and delivers it to the user.

[1008] Example 1

[1009] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1010] In the modern tourism industry, tourists are seeking more personalized experiences, but existing guidance systems are unable to meet these demands. Furthermore, information about tourist destinations is often provided in a one-way manner, creating a need for more interactive services. Furthermore, multilingual support is lacking, creating a significant language barrier for foreign tourists in particular.

[1011] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1012] In this invention, the server includes means for a user to select a character and a destination, means for transmitting the selected information to the server, means for the server to obtain destination information from a database, means for generating a voice guide for the character using a generative AI model, means for responding to user questions, and means for translating the generated response into multiple languages. This allows users to receive personalized interactive tourist information, and the multilingual support makes it easy for foreign tourists to obtain information.

[1013] "User character and destination selection means" refers to a device or software that provides an interface for a user to access the system and input or select a particular animated character and tourist destination.

[1014] The "means for transmitting selected information to the server" refers to a communication device or protocol for transmitting information about the character and destination selected by the user to the server.

[1015] "Means for the server to retrieve destination information from the database" refers to the process or software that the server uses to query a database that stores detailed information about tourist destinations based on the user's selection and retrieve the required information.

[1016] "Means for generating voice guidance for a character using a generative AI model" refers to a system or software that uses a generative AI model (e.g., a model using natural language generation technology) to generate voice guidance for a tourist in the voice of a character selected by the user.

[1017] The "means for responding to a user's question" refers to a system or software that analyzes a question entered by a user, generates an appropriate response, and provides that response to the user.

[1018] The "means for translating the generated response into multiple languages" refers to a translation system or software for translating the generated response into a language designated by the user.

[1019] "Means for the terminal to transmit information on the selected character and destination to the server" refers to communication functions or software that allow the user's terminal to transmit information on the selected character and tourist spot to the server.

[1020] "Means for the server to generate a guide plan based on a destination" refers to a system or software that enables the server to generate a detailed guide plan including descriptions of tourist spots and route guidance based on the destination selected by the user.

[1021] "Means for generating character voices using a generative AI model" refers to systems or software that utilize a generative AI model to generate audio content in the voice of a character selected by a user.

[1022] "Means for providing responses to user questions through a character" refers to a system or software that provides responses to user questions through the voice of a character generated using a generative AI model.

[1023] This invention is a system that uses animated characters to guide tourists around tourist spots, and its features include being interactive and supporting multiple languages. This system is composed of multiple hardware and software components, including a user terminal, a server, a database, and a generative AI model.

[1024] System configuration

[1025] 1. User Device

[1026] Function: Provides an interface for users to access the system and select characters and tourist attractions, plays audio guides for tourist attractions, and responds to user questions in real time.

[1027] Details: Through a smartphone or PC application, users select a character and a destination on the interface, and ask questions by text or voice input. For example, a user opens a smartphone app and selects animated character A and a park in Tokyo.

[1028] 2. Server

[1029] Function: Receives selection information sent from the user's device, retrieves tourist destination information from the database, and generates character voice and text using a generative AI model. Furthermore, it responds interactively to user questions and provides multilingual support as needed.

[1030] Details: Based on the user's selection, the server queries the database for detailed information about related tourist attractions. It then uses a generative AI model (e.g., OpenAI's ChatGPT) to generate audio data and responses for the tour guide and sends them to the user's device.

[1031] 3. Database

[1032] Function: Stores and manages detailed information about tourist spots (such as information about history, culture, and anime). Provides necessary information in response to requests from the server.

[1033] Details: The database is regularly updated with new tourist information and animations. It also contains historical information and specific anecdotes about the tourist destinations.

[1034] Specific examples

[1035] Below is a concrete example of how the system works:

[1036] Example 1: A park in Tokyo

[1037] 1. User selection: The user opens the app on their smartphone and selects Character A from an animation work and a certain park in Tokyo.

[1038] 2. Tour Guide: The server receives the selection information and retrieves detailed information about a certain park from the database. Using the generative AI model, it generates a narration in the voice of Character A, saying "This park is...", and sends it to the user's device.

[1039] 3. Question and Response: The user enters a question such as "Tell me about the history of this place." The server analyzes the question and uses a generative AI model to generate a response such as "This park was founded in ____..." and sends it to the device. The device plays the response in the voice of Character A.

[1040] 4. Multilingual support: If the user has English settings, the server will use artificial intelligence to translate the response into English and send it to the device as the English voice of Character A.

[1041] Prompt Sentence Examples

[1042] "The selected character is Character A, and the tourist destination is a certain park in Tokyo. Please generate audio of a tourist guide about this park."

[1043] The purpose of this invention is to improve the user's sightseeing experience through such an interactive guide system. In addition, by supporting multiple languages, it can also accommodate foreign tourists, thereby meeting global tourism demand.

[1044] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1045] Step 1:

[1046] The user selects a character and a destination

[1047] Specific operation: The user opens the application on their smartphone or computer and selects Character A and a certain park in Tokyo on the interface.

[1048] Input: User selection of character and tourist spot.

[1049] Output: Information about the selected character and tourist spot (e.g., Character A, a certain park).

[1050] Step 2:

[1051] The device sends the selection information to the server

[1052] Specific operation: The device sends JSON data containing information about the selected character and tourist spot to the server as an HTTP request.

[1053] Input: Selected character and tourist attraction information.

[1054] Output: The selection information sent to the server.

[1055] Step 3:

[1056] The server retrieves tourist information from the database

[1057] Specific operation: The server sends a query to the database to obtain information about "a certain park," such as its history and characteristics.

[1058] Input: User selection information (Character A, a certain park).

[1059] Output: Tourist information retrieved from the database (e.g. park history, main attractions, etc.).

[1060] Step 4:

[1061] The server inputs the prompt sentence into the AI ​​model and generates the guidance voice.

[1062] Specific operation: Based on the tourist attraction information obtained by the server, the prompt sentence "The selected character is Character A, and the tourist attraction is a certain park in Tokyo. Please generate audio tourist information about this park" is input into the generation AI model.

[1063] Input: Obtained tourist attraction information and prompt sentence.

[1064] Output: Guidance speech data generated by the generative AI model.

[1065] Step 5:

[1066] Send the generated audio data to the device

[1067] Specific operation: The server sends the generated voice data to the user's terminal.

[1068] Input: Guidance speech data generated by a generative AI model.

[1069] Output: The audio data sent to the device.

[1070] Step 6:

[1071] The device plays a voice prompt to the user.

[1072] Specific behavior: The device plays the audio data received and lets the user listen to it. Uses the audio player in the app.

[1073] Input: Audio data sent to the device.

[1074] Output: The audio guidance the user hears.

[1075] Step 7:

[1076] The user enters a question

[1077] What happens: A user types a question into a text field in the app: "Tell me about the history of this place."

[1078] Input: The question entered by the user.

[1079] Output: The data from the question entered.

[1080] Step 8:

[1081] The device sends a question to the server

[1082] Specific operation: The device sends JSON data containing the question to the server as an HTTP request.

[1083] Input: The question data entered by the user.

[1084] Output: The query data sent to the server.

[1085] Step 9:

[1086] The server analyzes the question and inputs the prompt into the generative AI model to generate a response.

[1087] Specific operation: The server analyzes the question and inputs the prompt sentence, "The user asked me, 'Tell me about the history of this place.' Please generate an answer to this question." into the generative AI model.

[1088] Input: User question data and generated prompt text.

[1089] Output: The response data generated by the generative AI model.

[1090] Step 10:

[1091] The server sends the response data to the terminal.

[1092] Specific operation: The server sends the generated response data to the terminal.

[1093] Input: The response data generated by the generative AI model.

[1094] Output: The response data sent to the device.

[1095] Step 11:

[1096] The terminal plays the response to the user

[1097] Specific behavior: The device plays back the response data received and lets the user listen to it. Uses the in-app audio player.

[1098] Input: Response data sent to the terminal.

[1099] Output: The response audio that the user hears.

[1100] Step 12:

[1101] If necessary, the server translates the response for multilingual support.

[1102] Specific operation: The server checks the user's language setting and translates the generated response if necessary. For example, if the user's language setting is English, the server translates the response into English and generates English speech data using the generative AI model.

[1103] Input: The user's preferred language and the generated response data.

[1104] Output: The translated response audio data.

[1105] (Application example 1)

[1106] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1107] Currently, guidance at tourist spots and brick-and-mortar stores requires human intervention, and it is difficult to provide multilingual and interactive guidance. Furthermore, when users want to obtain detailed information on their smartphones or in-store terminals, there is a lack of immediate guidance in a user-friendly format. Guidance using animated characters is particularly important for providing entertainment and convenience to users, but the technology to achieve this is still limited.

[1108] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1109] In this invention, the server includes means for generating guidance voice from an animated character, means for acquiring information about tourist spots or brick-and-mortar stores, means for interactively responding to user questions, means for providing multilingual support, means for a designated character to provide audio guidance about in-store product descriptions and recommendations, and means for a user terminal or robot to interactively respond to product-related questions. This enables multilingual and interactive guidance in tourist spots and brick-and-mortar stores, significantly improving the user experience.

[1110] "Animated character guidance voice" refers to an audio guide in the form of a specific animated character speaking, generated using a computer program.

[1111] A "tourist destination" is a specific place or area visited by tourists, and refers to an area that has historical, cultural, or natural attractions.

[1112] "Brick and mortar store" refers to a store that sells goods or services in a physical location.

[1113] "Means of obtaining information" refers to the functions and methods for collecting the required data and information from databases and other sources.

[1114] "Means for interactive response" refers to a system or function that returns an immediate response to a question or input from a user.

[1115] A "multilingual response means" is a method or system for providing information or responding to questions in multiple languages.

[1116] "Means for providing product descriptions and recommended information by voice" refers to a function or system for providing product descriptions and recommended information to users by voice.

[1117] "User terminal" refers to a computing device used by a user, such as a smartphone or tablet.

[1118] A "robot" refers to an automated mechanical device that operates automatically and performs specific tasks based on user instructions.

[1119] An embodiment of the present invention is described below: This system generates guide voices by animated characters at tourist spots or brick-and-mortar stores, and provides users with a multilingual interactive guide service.

[1120] 1. System Overview

[1121] The system consists of three main components:

[1122] 1.1 User terminal

[1123] The user terminals are personal devices such as smartphones and tablets. Using these terminals, users can select their favorite animated characters and tourist spots or brick-and-mortar stores, and then obtain information through the character's voice guidance. The terminals are connected to a server via the Internet.

[1124] 1.2 Server

[1125] The server plays a central role in the system and has the following functions:

[1126] Character voice generation: Generates guidance voices for designated animated characters. Utilizing a generative AI model, the voices are tailored to the personality of the character selected by the user.

[1127] Information acquisition: Obtain information about tourist spots and brick-and-mortar stores from the database.

[1128] Interactive Response: Generates optimal responses to user questions and translates them into multiple languages ​​as needed.

[1129] Guide plan generation: Generate a travel plan or product introduction plan as needed.

[1130] 1.3 Database

[1131] The database contains detailed information about tourist destinations and brick-and-mortar stores, including their history, culture, product features, and prices. The database is updated regularly.

[1132] 2. Details of the processing

[1133] The server proceeds in the following way:

[1134] 2.1 Character voice generation

[1135] A generative AI model is used to generate a voice for the selected character. For example, the gTTS (Google Text-to-Speech) library is used to generate an audio file to provide voice guidance. This audio file is sent to the user's device, where it can be played back.

[1136] 2.2 Information acquisition

[1137] The server retrieves necessary information from a database, such as historical information about tourist spots or product descriptions from physical stores, and generates voice guidance to provide to the user based on that information.

[1138] 2.3 Interactive Response

[1139] When a user enters a question in voice or text format, the server analyzes the question and generates the best possible response using a generative AI model. If necessary, the response is translated for multilingual support and sent to the user's device.

[1140] 3. Specific Examples

[1141] Here are some concrete examples of how the system can be used:

[1142] Example 1: Tourist information

[1143] The user selects a park as a tourist attraction and a specific anime character as a character. The server retrieves information about the park from a database and uses a generative AI model to generate a guide in the character's voice. When the user asks, "Tell me about the history of this park," the server generates an appropriate response to the question and sends it to the user's device. The user can then play the audio guide on their device and tour the park.

[1144] Prompt Sentence Examples

[1145] User: "I chose Pikachu."

[1146] App: "Pikachu's voice will begin prompting. What product would you like to know more about?"

[1147] User: "What product is 12345?"

[1148] App: "This item is of the highest quality. It costs 1000 yen."

[1149] User: "Where does this product originate from?"

[1150] App: "Origin: Japan."

[1151] In this way, users can enjoy being guided around tourist spots and brick-and-mortar stores together with animated characters.

[1152] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1153] Step 1:

[1154] Input: The user launches the application using their device and selects their favorite animated character and tourist destination or brick-and-mortar store.

[1155] Processing: The terminal receives as input information about the characters and tourist spots or brick-and-mortar stores selected by the user, and transmits the selected information to the server.

[1156] Output: The selection information sent to the server.

[1157] Specific operation: The user opens the smartphone app and taps to select a character and destination from the on-screen menu.

[1158] Step 2:

[1159] Input: The server receives the selection information sent by the user.

[1160] Processing: The server retrieves detailed information about tourist attractions or brick-and-mortar stores from the database based on the received selection information.

[1161] Output: Detailed information about a tourist attraction or brick-and-mortar store retrieved from the database.

[1162] What it does: The server queries a database to retrieve information about the destination, such as its history and product descriptions.

[1163] Step 3:

[1164] Input: Details retrieved from the database.

[1165] Processing: The server uses the generative AI model to generate a voice guide for the selected character based on the detailed information.

[1166] Output: Guidance voice data of the generated character.

[1167] Specific operation: Using a generative AI model (e.g., gTTS library), convert text information into the voice of a specified character and create an audio file.

[1168] Step 4:

[1169] Input: Generated guidance speech data.

[1170] Processing: The server transmits the generated voice data to the user terminal.

[1171] Output: Character guidance voice data sent to the user's terminal.

[1172] Specific operation: The server sends the audio file to the user's smartphone via the network.

[1173] Step 5:

[1174] Input: Voice data sent to the user's terminal.

[1175] Processing: The user terminal plays back the received guidance voice.

[1176] Output: The audio guidance that the user can hear.

[1177] Specific operation: The selected character's voice will play instructions through the smartphone speaker.

[1178] Step 6:

[1179] Input: The user types a question into the application using text or voice.

[1180] Processing: The terminal recognizes the user's question and sends it to the server.

[1181] Output: The user's question sent to the server.

[1182] Specific action: The user speaks a question or enters text using the device's microphone.

[1183] Step 7:

[1184] Input: The user's question received by the server.

[1185] Processing: The server analyzes the question and generates the best response using its database and generative AI models, translating it into multiple languages ​​if necessary.

[1186] Output: The generated response data.

[1187] Specific operation: The server analyzes the question using natural language processing technology, generates a corresponding answer, and then translates it into the specified language.

[1188] Step 8:

[1189] Input: The generated response data.

[1190] Processing: The server sends the generated response data to the user terminal.

[1191] Output: The response data sent to the user terminal.

[1192] Specific operation: The server sends response data to the user's smartphone via the network.

[1193] Step 9:

[1194] Input: Response data sent to the user terminal.

[1195] Processing: The user terminal plays back the answer in the character's voice based on the received response data.

[1196] Output: An interactive audio response that the user can hear.

[1197] Specific operation: The smartphone plays the received data as an audio file and provides the answer to the user.

[1198] This allows users to enjoy information and detailed information about tourist spots and brick-and-mortar stores through animated characters.

[1199] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1200] This invention is an interactive pilgrimage tour system using animated characters combined with an emotion engine that recognizes the user's emotions. This system not only provides detailed information about tourist spots, but also provides appropriate guidance and responses based on the user's emotions. It also supports multiple languages, making it suitable for foreign tourists.

[1201] System configuration

[1202] 1. User Device

[1203] Functions: Provides an interface for users to access the system and select characters and tourist attractions. Also plays audio guides for tourist attractions and displays and plays responses to user questions in real time. Furthermore, it has an emotion engine for recognizing emotions from user input (text and voice).

[1204] How it works: Users access a designated website or application using a device such as a smartphone or PC. They select a character and a destination on the interface, and if they have any questions, they input them using text or voice. The emotion engine analyzes the user's input and provides guidance accordingly.

[1205] 2. Server

[1206] Function: Receives character and destination selection information sent from the user's device and retrieves tourist spot information from a database. Uses a generative AI model to generate character voice and text. Additionally, responds interactively to user questions and supports multiple languages. Also, adjusts guidance and response content based on the user's emotional data.

[1207] Processing method: Based on the user's selection, the server retrieves relevant sacred site information (such as name, location, and history) from a database. It then uses a generative AI model to generate audio data for tour guidance and transmits it to the user's device. When the user enters a question, the server analyzes the question, generates an optimal response, and adjusts the response according to the emotion recognized by the emotion engine.

[1208] 3. Database

[1209] Function: Stores and manages detailed information about tourist spots (history, culture, episodes related to animation and manga, etc.). Provides necessary information in response to requests from the server.

[1210] Management method: The database is updated regularly to include new tourist destination information and information on animation and manga.

[1211] Specific examples

[1212] A specific example of the system is shown below.

[1213] Example 1: A park in Tokyo

[1214] 1. User selection: The user opens the smartphone application and selects Character A from the animation work and a certain park in Tokyo.

[1215] 2. Tour Guide: The server receives the selection information and retrieves detailed information about a certain park from a database. Using a generative AI model, it generates a narration in the voice of Character A, saying "This park is...", and sends it to the user's device. The user plays the narration on their device and tours the park.

[1216] 3. Question and Response: The user enters a question, such as "Tell me about the history of this place." The server analyzes the question and uses a generative AI model to generate a response, such as "This park was founded in ____...," which is sent to the device. The device then plays back the response in the voice of Character A.

[1217] 4. Emotion Recognition: When a user types "What a beautiful place," the emotion engine analyzes the emotion data and recognizes that the user is happy. The server generates a response based on the emotion, such as "I'm glad to hear that," and plays it on the device.

[1218] 5. Multilingual support: If the user has English settings, the server will use artificial intelligence to translate the response into English and send it to the device as the English voice of Character A.

[1219] This allows users to have a rich experience touring tourist spots while interacting with animated characters.The purpose of the present invention is to improve the quality of users' sightseeing experiences through such an interactive guide system.

[1220] The processing flow will be explained below.

[1221] Step 1:

[1222] A user starts an application on a terminal and accesses the system. The terminal generates initial session information and sends it to the server.

[1223] Step 2:

[1224] The terminal displays the animated characters and a list of tourist attractions on a user interface.

[1225] Step 3:

[1226] Through the interface, the user selects an animated guide character and the tourist sites they wish to visit.

[1227] Step 4:

[1228] The device sends the user's selection (character ID, destination ID, etc.) to the server.

[1229] Step 5:

[1230] The server retrieves information about the selected character and tourist spot from the database.

[1231] Step 6:

[1232] The server generates a tour plan based on the acquired information and determines the spots to visit and the content of the narration.

[1233] Step 7:

[1234] The server uses the generative AI model to generate narration text for the tour plan using the voice of the selected character, creating audio data.

[1235] Step 8:

[1236] The server transmits the generated voice data to the user's terminal.

[1237] Step 9:

[1238] The terminal starts playing the audio data received from the server as narration.

[1239] Step 10:

[1240] During the tour, the user can ask questions by text or voice, for example, "Tell me about something special about this place."

[1241] Step 11:

[1242] The terminal sends the user's question to the server.

[1243] Step 12:

[1244] The server analyzes the question and retrieves relevant information from a database.

[1245] Step 13:

[1246] The server uses the generative AI model to generate the analyzed information as voice data for the character, for example, "This park was established at ____."

[1247] Step 14:

[1248] The server transmits the generated response voice data to the terminal.

[1249] Step 15:

[1250] The terminal reproduces the response voice data received from the server.

[1251] Step 16:

[1252] When a user speaks, the emotion engine analyzes the user's input (text or voice) and recognizes the emotion. For example, if a user inputs "It's a very beautiful place," the device analyzes it and sends it to the emotion engine.

[1253] Step 17:

[1254] The emotion engine recognizes the user's emotion and sends the emotion data to the server. For example, it recognizes that the user is happy.

[1255] Step 18:

[1256] The server uses a generative AI model to generate an appropriate response based on the emotional data, such as "I'm glad to hear that."

[1257] Step 19:

[1258] The server transmits the generated emotion response voice data to the terminal.

[1259] Step 20:

[1260] The terminal reproduces the emotion response voice data received from the server.

[1261] Step 21:

[1262] The server checks the language specified by the user in the settings, translates the voice data into that language if necessary, and sends the translated voice to the device.

[1263] Step 22:

[1264] The device plays the translated audio data and delivers it to the user.

[1265] Example 2

[1266] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1267] Conventional tourist information systems provide only one-way guidance information, making it difficult to provide responses that take the user's emotions into consideration. Furthermore, they lack multilingual support, making it impossible to provide the same service to foreign tourists. Furthermore, interactive responses and real-time guidance generation are insufficient, leaving room for improvement in the user experience.

[1268] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for generating a guidance voice of an animated character, a means for acquiring information on tourist spots, a means for interactively responding to a user's questions, a means for providing responses in multiple languages, and a means for analyzing the user's emotions and adjusting the response based on the emotions. This makes it possible to provide interactive tourist information that takes the user's emotions into consideration in real time.

[1269] An "animated character" is a computer-generated character, such as a person or animal, that is displayed visually and interacts with the user through sound and movement.

[1270] "Guidance voice" is voice data for providing information on tourist spots and the like by voice, and is used to guide and explain to the user.

[1271] "Tourist destination information" refers to detailed data about a tourist destination, including its location, history, culture, and access methods.

[1272] An "interactive response means" is a device or program that has the function of responding in real time to questions or input from a user.

[1273] "Multilingual support" means providing services and responses in multiple languages ​​to users who speak different languages.

[1274] "Means for analyzing emotions" refers to technology or devices that identify emotions from user input or behavior and acquire them as data.

[1275] A "generative AI model" is an algorithm or program that uses artificial intelligence to generate text or speech.

[1276] A "database" is a system for efficiently storing and managing information, accumulating large amounts of data and allowing it to be quickly searched and retrieved as needed.

[1277] A "user terminal" is a device operated by a user, such as a smartphone, PC, or tablet.

[1278] A "server" is a computer system that receives requests from clients via a network and provides the necessary information.

[1279] This invention is an interactive tourist information system using animated characters combined with an emotion engine that recognizes the user's emotions. This system not only provides detailed information about tourist spots, but also provides appropriate guidance and responses based on the user's emotions. It also supports multiple languages, making it suitable for foreign tourists.

[1280] System configuration

[1281] User terminal

[1282] Users access a designated website or application using a device such as a smartphone or PC. Through this interface, users can select an animated character and a desired tourist destination and enter their request. The user device is equipped with an emotion engine that recognizes emotions from the user's input (text or voice). For example, if a user enters "I want to go to a certain park in Tokyo," the request is sent to the server.

[1283] server

[1284] The server receives the character and tourist attraction selection information sent from the user's device. The server accesses the database to obtain detailed information about the relevant tourist attraction. It then uses a generative AI model to generate audio guidance data. At this time, guidance such as "This park was established in XX..." is generated in the voice of the selected character. This audio data is sent to the user's device and played back.

[1285] Database

[1286] The database stores and manages detailed information about tourist destinations (location, history, culture, etc.). This information is used by the server and provided to the service as needed. The database is also regularly updated with new tourist destination information and information about animations and comics.

[1287] Specific examples

[1288] A specific example of the system's operation is shown below.

[1289] Example 1: A park in Tokyo

[1290] 1. User selection: The user opens the smartphone application and selects animated character A and a certain park in Tokyo.

[1291] 2. Tour Guide: The server receives the selection information and retrieves detailed information about a certain park from a database. Using a generative AI model, it generates a narration in the voice of Character A, saying "This park is...", and sends it to the user's device. The user plays the narration on their device and tours the park.

[1292] 3. Question and Response: The user enters a question such as "Tell me about the history of this place." The server analyzes the question and uses a generative AI model to generate a response such as "This park was founded in ____..." and sends it to the device. The device plays the response in the voice of Character A.

[1293] 4. Emotion Recognition: When a user types "What a beautiful place," the emotion engine analyzes the emotion data and recognizes that the user is happy. The server generates a response based on the emotion, such as "I'm glad to hear that," and plays it on the device.

[1294] 5. Multilingual support: If the user has English settings, the server will translate the response into English using the generative AI model and send it to the device as Character A's English voice to play.

[1295] Prompt Sentence Examples

[1296] 1. "I'd like you to show me around a certain park in Tokyo."

[1297] 2. "Tell me about the history of this park."

[1298] 3. "It's a beautiful place."

[1299] 4. “Can you provide the information in English?”

[1300] Through such a system, users can have a rich experience touring tourist spots while interacting with animated characters. The present invention aims to improve the quality of users' sightseeing experiences.

[1301] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1302] Step 1:

[1303] Users open the application on their smartphone or PC and select an animated character and a tourist destination of their choice. The user's selection and request are entered and sent from the device to the server. Specifically, the user taps or clicks on an option on the screen, and the selection is collected by the application and transmitted to the server via the network.

[1304] Step 2:

[1305] The server receives the character and tourist attraction selection information sent by the user. The server then queries the database to obtain detailed information about the corresponding tourist attraction. Specifically, a database query is executed using the character and tourist attraction identifiers. The query results in data such as the tourist attraction's name, location, history, and culture.

[1306] Step 3:

[1307] The server uses a generative AI model to generate voice guidance data based on the acquired tourist attraction information and user request information. Using the acquired tourist attraction information and request information as input, it processes and calculates data. Character voice guidance data is generated as output. Specifically, the generative AI model creates narration from text information and converts it into voice data.

[1308] Step 4:

[1309] The server transmits the generated voice guidance data to the user terminal. Specifically, a data packet containing the voice guidance data is transmitted to the user terminal via the network. The user terminal then plays back the received voice guidance data.

[1310] Step 5:

[1311] As users travel around tourist spots, they input questions and comments through their devices. The user's input is in the form of text or voice. The user's device then sends the input data to the server. Specifically, when the user inputs a question and taps the send button, the data is collected and transferred to the server.

[1312] Step 6:

[1313] The server receives and analyzes questions from users. It analyzes the question data received as input and generates an optimal response using a generative AI model. The generated response data is obtained as output. Specifically, a natural language processing algorithm analyzes the question and generates an appropriate response.

[1314] Step 7:

[1315] The generated response data is sent from the server to the user terminal. Specifically, a data packet containing the response data is sent to the user terminal via the network. The user terminal then plays back the received response data in the character's voice.

[1316] Step 8:

[1317] The emotion engine analyzes the user's input data (text and voice) and extracts emotion data. It uses the user's comment data as input and performs data analysis. Emotion data is obtained as output. Specifically, text mining and voice analysis algorithms recognize emotions.

[1318] Step 9:

[1319] The server adjusts the response content based on the emotional data. Using the emotional data and generated response data as input, it processes the data again using the generative AI model. The output is response data based on the emotion. Specifically, an algorithm is executed to generate response content that reflects the emotional data.

[1320] Step 10:

[1321] If the user has multilingual settings, the server translates the generated response data into the specified language. Using the generated response data as input, it performs data calculations using an AI translation model. The translated response data is obtained as output. Specifically, the translation algorithm converts the response data into another language.

[1322] In this way, users can experience interactive tourist information that responds to their emotions through animated characters. In addition, the multilingual support makes it easy for foreign tourists to use the service.

[1323] (Application example 2)

[1324] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1325] In recent years, the number of interactive systems that provide information about tourist destinations and restaurants has increased. However, current systems are unable to provide optimal responses based on user emotions. This results in problems such as not being able to provide the information and services users desire in a timely and appropriate manner. Furthermore, the lack of multilingual support makes these systems difficult for foreign tourists to use. To solve these issues, a system that analyzes user emotions and provides optimal recommendations based on those emotions is needed.

[1326] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1327] In this invention, the server includes a means for recognizing a user's emotions and adjusting the response content based on the emotions, a means for recommending optimal restaurants and dishes based on the emotions, and a means for generating recommendation text using a generative AI model, thereby enabling interactive responses according to the user's emotions.

[1328] The "means for generating guidance voice of an animated character" is a function for providing the user with information about tourist spots and restaurants in the voice of an animated character.

[1329] "Means for obtaining information on tourist destinations" refers to a function for collecting detailed data on tourist destinations from databases and external information sources.

[1330] The "means for interactively responding to user questions" is a function for analyzing questions from users and providing appropriate answers in real time.

[1331] "Means for providing responses in multiple languages" is a function for providing answers and guidance adapted to each language in scenarios requiring multilingual support.

[1332] "Means for recognizing the user's emotions and adjusting the response content based on those emotions" is a function for analyzing the emotions from the user's input content and voice and generating the optimal response accordingly.

[1333] The "means for recommending optimal restaurants and dishes based on emotions" is a function for taking into account the emotional state of the user and recommending restaurants and dishes that are suitable for them.

[1334] "Means for generating recommendation text using a generative AI model" is a function that utilizes an AI model to generate recommendation text based on a user's emotions and questions.

[1335] A "terminal" is a device operated by a user, such as a smartphone or tablet.

[1336] A "server" is a central processing unit that processes data from user terminals and generates and transmits necessary information.

[1337] This invention is an interactive food delivery application system that combines an emotion engine that recognizes the user's emotions. The system aims to recommend the most suitable restaurants and dishes to the user by linking the user terminal, server, and database.

[1338] System configuration

[1339] 1. User Device

[1340] Function: Provides an interface for users to access the application and select restaurants and dishes. It also recognizes user emotions from input (text and voice) and has an emotion engine. It also displays and plays recommendations based on emotions.

[1341] How it works: Users access the designated application using a smartphone or other device. They select a restaurant and a dish on the interface, and input their questions or preferences via text or voice. The emotion engine analyzes the emotions from these inputs and makes food and drink recommendations accordingly.

[1342] 2. Server

[1343] Function: Receives sentiment analysis results and question data sent from the user's device and retrieves relevant dining information from the database. Uses a generative AI model to generate recommendation text based on the user's sentiment and question and sends it to the user's device. Also supports multiple languages, providing translated responses into foreign languages.

[1344] Processing method: The server retrieves relevant restaurant information (store information, dish details, reviews, etc.) from a database based on the user's sentiment analysis results and question data. It then uses a generative AI model to generate optimal recommendation text and sends it to the user's device. The recommendation content is adjusted based on the sentiment recognized by the emotion engine.

[1345] 3. Database

[1346] Function: Stores and manages detailed information about restaurants and dishes (menus, prices, nutritional information, user reviews, etc.). Provides necessary information in response to requests from the server.

[1347] How it's maintained: The database is updated regularly to include information on new restaurants and cuisines.

[1348] Specific examples

[1349] A specific example of the system is shown below.

[1350] Example 1: Food and drink recommendations when the user is tired

[1351] 1. User makes a choice: The user opens the application on their smartphone and types "I'm tired."

[1352] 2. Emotion recognition and recommendation: The server receives the user's input, and the emotion engine analyzes the emotion "tired." Using the generative AI model, a recommendation text is generated, such as "We recommend the stamina bowl at nearby Restaurant A, which is a nutritious dish," and sent to the user's device.

[1353] 3. Multilingual support: If the user selects English, the generated recommendation text is translated into English and sent to the device, enabling multilingual support.

[1354] Example prompts for generative AI models

[1355] "What foods would you recommend for users when they're feeling low?"

[1356] "What food can cheer you up when you're feeling tired?"

[1357] In this way, through the system of the present invention, users can receive recommendations for optimal food and drink based on their emotions and enjoy a satisfying eating and drinking experience.

[1358] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1359] Step 1:

[1360] The user launches the smartphone application and inputs their question or request via text or voice.

[1361] Input: User text or voice input (e.g., "I'm tired today and want to eat something energizing.")

[1362] Output: The app gets the user's input data.

[1363] Specific actions: The user inputs their feelings and questions using the dedicated app interface, then presses the "Send" button.

[1364] Step 2:

[1365] The device sends the user's input data to an emotion recognition engine, which analyzes the emotion.

[1366] Input: User text or voice data

[1367] Output: Emotion recognition result (e.g. "tired")

[1368] Specific operation: The app's built-in emotion recognition engine analyzes the user's input and extracts emotions such as "tired" as labels.

[1369] Step 3:

[1370] The device sends the emotion recognition results to the server and requests optimal food and drink recommendations.

[1371] Input: Emotion recognition result (e.g., "tired")

[1372] Output: Send emotion recognition results to the server

[1373] Specific operation: The emotion recognition results are sent to the server, and the app calls the server's API using a network connection.

[1374] Step 4:

[1375] The server retrieves relevant dining information from the database.

[1376] Input: Emotion recognition results to the server

[1377] Output: Related food and drink information (e.g., food items such as "Stamina bowl")

[1378] Specific operation: The server queries the database and obtains food and drink information (e.g., foods that provide stamina) related to the emotion recognition results.

[1379] Step 5:

[1380] The server uses a generative AI model to generate recommendation text based on the user's sentiment and question.

[1381] Input: Food and drink information obtained from the database, emotion recognition results

[1382] Output: Recommendation text (e.g. "You seem tired today. I recommend a stamina bowl to give you energy.")

[1383] Specific operation: The server uses a generative AI model such as OpenAI's GPT-4, inputs food and drink information and emotion recognition results as prompts, and generates recommendation text.

[1384] Step 6:

[1385] The server transmits the generated recommendation text to the user terminal.

[1386] Input: Recommendation text

[1387] Output: Send recommendation text to user device

[1388] Specific operation: The server sends the generated recommendation text to the user terminal via the network.

[1389] Step 7:

[1390] The terminal displays or plays the recommendation text to the user.

[1391] Input: Recommendation text

[1392] Output: Display to user or play as audio

[1393] What happens: The app displays the suggested text on the screen or plays it aloud using a text-to-speech engine.

[1394] Through these steps, users can receive optimal food and drink recommendations tailored to their emotional state. Using a generative AI model, we can provide more natural and specific recommendation text.

[1395] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1396] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1397] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1398] [Fourth embodiment]

[1399] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1400] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1401] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1402] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1403] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1404] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1405] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1406] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1407] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1408] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1409] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1410] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1411] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1412] This invention is a system that generates guide voices for animated characters and provides real-time information about tourist spots. This system realizes an interactive pilgrimage tour based on the characters and tourist spots selected by the user. It also supports multiple languages, making it suitable for foreign tourists.

[1413] System configuration

[1414] 1. User Device

[1415] Function: Provides an interface for users to access the system and select characters and tourist attractions, plays audio guides for tourist attractions, and displays and plays responses to user questions in real time.

[1416] How it works: The user accesses a designated website or application using a device such as a smartphone or PC, selects a character and a destination on the interface, and uses text or voice input to ask any questions.

[1417] 2. Server

[1418] Function: Receives character and destination selection information sent from the user's device and retrieves tourist spot information from the database. Uses a generative AI model to generate character voice and text. Furthermore, responds interactively to user questions and supports multiple languages.

[1419] Processing method: Based on the user's selection, the server retrieves relevant sacred site information (such as name, location, and history) from the database. It then uses a generative AI model to generate audio data for the tour guide and sends it to the user's device. When the user enters a question, the server analyzes the question and generates the optimal response.

[1420] 3. Database

[1421] Function: Stores and manages detailed information about tourist spots (history, culture, episodes related to animation and manga, etc.). Provides necessary information in response to requests from the server.

[1422] Management method: The database is updated regularly to include new tourist destination information and information on animation and manga.

[1423] Specific examples

[1424] A specific example of the system is shown below.

[1425] Example 1: A park in Tokyo

[1426] 1. User selection: The user opens the smartphone application and selects Character A from the animation work and a certain park in Tokyo.

[1427] 2. Tour Guide: The server receives the selection information and retrieves detailed information about a certain park from a database. Using a generative AI model, it generates a narration in the voice of Character A, saying "This park is...", and sends it to the user's device. The user plays the narration on their device and tours the park.

[1428] 3. Question and Response: The user enters a question, such as "Tell me about the history of this place." The server analyzes the question and uses a generative AI model to generate a response, such as "This park was founded in ____...," which is sent to the device. The device then plays back the response in the voice of Character A.

[1429] 4. Multilingual support: If the user has English settings, the server will use artificial intelligence to translate the response into English and send it to the device as the English voice of Character A.

[1430] This allows users to have a rich experience touring tourist spots while interacting with animated characters.The purpose of the present invention is to improve the quality of users' sightseeing experiences through such an interactive guide system.

[1431] The processing flow will be explained below.

[1432] Step 1:

[1433] A user starts an application on a terminal and accesses the system. The terminal generates initial session information and sends it to the server.

[1434] Step 2:

[1435] The terminal displays the animated characters and a list of tourist spots on a user interface.

[1436] Step 3:

[1437] Through the interface, the user selects an animated guide character and the tourist sites they wish to visit.

[1438] Step 4:

[1439] The terminal transmits the user's selection (character ID, destination ID, etc.) to the server.

[1440] Step 5:

[1441] The server retrieves information about the selected character and tourist spot from a database.

[1442] Step 6:

[1443] The server generates a tour plan based on the acquired information and determines the spots to visit and the content of the narration.

[1444] Step 7:

[1445] The server generates narration text for the tour plan using the voice of the selected character using a generative AI model, and creates audio data.

[1446] Step 8:

[1447] The server transmits the generated voice data to the user's terminal.

[1448] Step 9:

[1449] The terminal starts playing the audio data received from the server as narration.

[1450] Step 10:

[1451] The user asks questions during the tour (e.g., "Tell me about something special about this place") via text or voice input.

[1452] Step 11:

[1453] The terminal sends the user's question to the server.

[1454] Step 12:

[1455] The server analyzes the question and retrieves relevant information from a database.

[1456] Step 13:

[1457] The server uses a generative AI model to generate the analyzed information as voice data for the character.

[1458] Step 14:

[1459] The server transmits the generated response voice data to the terminal.

[1460] Step 15:

[1461] The terminal reproduces the response voice data received from the server.

[1462] Step 16:

[1463] The server checks the language specified by the user in the settings, translates the voice data into that language if necessary, and sends the translated voice to the device.

[1464] Step 17:

[1465] The device plays the translated audio data and delivers it to the user.

[1466] Example 1

[1467] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1468] In the modern tourism industry, tourists are seeking more personalized experiences, but existing guidance systems are unable to meet these demands. Furthermore, information about tourist destinations is often provided in a one-way manner, creating a need for more interactive services. Furthermore, multilingual support is lacking, creating a significant language barrier for foreign tourists in particular.

[1469] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1470] In this invention, the server includes means for a user to select a character and a destination, means for transmitting the selected information to the server, means for the server to obtain destination information from a database, means for generating a voice guide for the character using a generative AI model, means for responding to user questions, and means for translating the generated response into multiple languages. This allows users to receive personalized interactive tourist information, and the multilingual support makes it easy for foreign tourists to obtain information.

[1471] "User character and destination selection means" refers to a device or software that provides an interface for a user to access the system and input or select a particular animated character and tourist destination.

[1472] The "means for transmitting selected information to the server" refers to a communication device or protocol for transmitting information about the character and destination selected by the user to the server.

[1473] "Means for the server to retrieve destination information from the database" refers to the process or software that the server uses to query a database that stores detailed information about tourist destinations based on the user's selection and retrieve the required information.

[1474] "Means for generating voice guidance for a character using a generative AI model" refers to a system or software that uses a generative AI model (e.g., a model using natural language generation technology) to generate voice guidance for a tourist in the voice of a character selected by the user.

[1475] The "means for responding to a user's question" refers to a system or software that analyzes a question entered by a user, generates an appropriate response, and provides that response to the user.

[1476] The "means for translating the generated response into multiple languages" refers to a translation system or software for translating the generated response into a language designated by the user.

[1477] "Means for the terminal to transmit information on the selected character and destination to the server" refers to communication functions or software that allow the user's terminal to transmit information on the selected character and tourist spot to the server.

[1478] "Means for the server to generate a guide plan based on a destination" refers to a system or software that enables the server to generate a detailed guide plan including descriptions of tourist spots and route guidance based on the destination selected by the user.

[1479] "Means for generating character voice using a generative AI model" refers to a system or software that utilizes a generative AI model to generate audio content in the voice of a character selected by a user.

[1480] "Means for providing responses to user questions through a character" refers to a system or software that provides responses to user questions through the voice of a character generated using a generative AI model.

[1481] This invention is a system that uses animated characters to guide tourists around tourist spots, and its features include being interactive and supporting multiple languages. This system is composed of multiple hardware and software components, including a user terminal, a server, a database, and a generative AI model.

[1482] System configuration

[1483] 1. User Device

[1484] Function: Provides an interface for users to access the system and select characters and tourist attractions, plays audio guides for tourist attractions, and responds to user questions in real time.

[1485] Details: Through a smartphone or PC application, users select a character and a destination on the interface, and ask questions by text or voice input. For example, a user opens a smartphone app and selects animated character A and a park in Tokyo.

[1486] 2. Server

[1487] Function: Receives selection information sent from the user's device, retrieves tourist destination information from the database, and generates character voice and text using a generative AI model. Furthermore, it responds interactively to user questions and provides multilingual support as needed.

[1488] Details: Based on the user's selection, the server queries the database for detailed information about related tourist attractions. It then uses a generative AI model (e.g., OpenAI's ChatGPT) to generate audio data and responses for the tour guide and sends them to the user's device.

[1489] 3. Database

[1490] Function: Stores and manages detailed information about tourist spots (such as information about history, culture, and anime). Provides necessary information in response to requests from the server.

[1491] Details: The database is regularly updated with new tourist information and animations. It also contains historical information and specific anecdotes about the tourist destinations.

[1492] Specific examples

[1493] Below is a concrete example of how the system works:

[1494] Example 1: A park in Tokyo

[1495] 1. User selection: The user opens the app on their smartphone and selects Character A from an animation work and a certain park in Tokyo.

[1496] 2. Tour Guide: The server receives the selection information and retrieves detailed information about a certain park from the database. Using the generative AI model, it generates a narration in the voice of Character A, saying "This park is...", and sends it to the user's device.

[1497] 3. Question and Response: The user enters a question such as "Tell me about the history of this place." The server analyzes the question and uses a generative AI model to generate a response such as "This park was founded in ____..." and sends it to the device. The device plays the response in the voice of Character A.

[1498] 4. Multilingual support: If the user has English settings, the server will use artificial intelligence to translate the response into English and send it to the device as the English voice of Character A.

[1499] Prompt Sentence Examples

[1500] "The selected character is Character A, and the tourist destination is a certain park in Tokyo. Please generate audio of a tourist guide about this park."

[1501] The purpose of this invention is to improve the user's sightseeing experience through such an interactive guide system. In addition, by supporting multiple languages, it can also accommodate foreign tourists, thereby meeting global tourism demand.

[1502] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1503] Step 1:

[1504] The user selects a character and a destination

[1505] Specific operation: The user opens the application on their smartphone or computer and selects Character A and a certain park in Tokyo on the interface.

[1506] Input: User selection of character and tourist spot.

[1507] Output: Information about the selected character and tourist spot (e.g., Character A, a certain park).

[1508] Step 2:

[1509] The device sends the selection information to the server

[1510] Specific operation: The device sends JSON data containing information about the selected character and tourist spot to the server as an HTTP request.

[1511] Input: Selected character and tourist attraction information.

[1512] Output: The selection information sent to the server.

[1513] Step 3:

[1514] The server retrieves tourist information from the database

[1515] Specific operation: The server sends a query to the database to obtain information about "a certain park," such as its history and characteristics.

[1516] Input: User selection information (Character A, a certain park).

[1517] Output: Tourist information retrieved from the database (e.g. park history, main attractions, etc.).

[1518] Step 4:

[1519] The server inputs the prompt sentence into the AI ​​model and generates the guidance voice.

[1520] Specific operation: Based on the tourist attraction information obtained by the server, the prompt sentence "The selected character is Character A, and the tourist attraction is a certain park in Tokyo. Please generate audio tourist information about this park" is input into the generation AI model.

[1521] Input: Obtained tourist attraction information and prompt sentence.

[1522] Output: Guidance speech data generated by the generative AI model.

[1523] Step 5:

[1524] Send the generated audio data to the device

[1525] Specific operation: The server sends the generated voice data to the user's terminal.

[1526] Input: Guidance speech data generated by a generative AI model.

[1527] Output: The audio data sent to the device.

[1528] Step 6:

[1529] The device plays a voice prompt to the user.

[1530] Specific behavior: The device plays the audio data received and lets the user listen to it. Uses the audio player in the app.

[1531] Input: Audio data sent to the device.

[1532] Output: The audio guidance the user hears.

[1533] Step 7:

[1534] The user enters a question

[1535] What happens: A user types a question into a text field in the app: "Tell me about the history of this place."

[1536] Input: The question entered by the user.

[1537] Output: The data from the question entered.

[1538] Step 8:

[1539] The device sends a question to the server

[1540] Specific operation: The device sends JSON data containing the question to the server as an HTTP request.

[1541] Input: The question data entered by the user.

[1542] Output: The query data sent to the server.

[1543] Step 9:

[1544] The server analyzes the question and inputs the prompt into the generative AI model to generate a response.

[1545] Specific operation: The server analyzes the question and inputs the prompt sentence, "The user asked me, 'Tell me about the history of this place.' Please generate an answer to this question." into the generative AI model.

[1546] Input: User question data and generated prompt text.

[1547] Output: The response data generated by the generative AI model.

[1548] Step 10:

[1549] The server sends the response data to the terminal.

[1550] Specific operation: The server sends the generated response data to the terminal.

[1551] Input: The response data generated by the generative AI model.

[1552] Output: The response data sent to the device.

[1553] Step 11:

[1554] The terminal plays the response to the user

[1555] Specific behavior: The device plays back the response data received and lets the user listen to it. Uses the in-app audio player.

[1556] Input: Response data sent to the terminal.

[1557] Output: The response audio that the user hears.

[1558] Step 12:

[1559] If necessary, the server translates the response for multilingual support.

[1560] Specific operation: The server checks the user's language setting and translates the generated response if necessary. For example, if the user's language setting is English, the server translates the response into English and generates English speech data using the generative AI model.

[1561] Input: The user's preferred language and the generated response data.

[1562] Output: The translated response audio data.

[1563] (Application example 1)

[1564] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1565] Currently, guidance at tourist spots and brick-and-mortar stores requires human intervention, and it is difficult to provide multilingual and interactive guidance. Furthermore, when users want to obtain detailed information on their smartphones or in-store terminals, there is a lack of immediate guidance in a user-friendly format. Guidance using animated characters is particularly important for providing entertainment and convenience to users, but the technology to achieve this is still limited.

[1566] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1567] In this invention, the server includes means for generating guidance voice from an animated character, means for acquiring information about tourist spots or brick-and-mortar stores, means for interactively responding to user questions, means for providing multilingual support, means for a designated character to provide audio guidance about in-store product descriptions and recommendations, and means for a user terminal or robot to interactively respond to product-related questions. This enables multilingual and interactive guidance in tourist spots and brick-and-mortar stores, significantly improving the user experience.

[1568] "Animated character guidance voice" refers to an audio guide in the form of a specific animated character speaking, generated using a computer program.

[1569] A "tourist destination" is a specific place or area visited by tourists, and refers to an area that has historical, cultural, or natural attractions.

[1570] "Brick and mortar store" refers to a store that sells goods or services in a physical location.

[1571] "Means of obtaining information" refers to the functions and methods for collecting the required data and information from databases and other sources.

[1572] "Means for interactive response" refers to a system or function that returns an immediate response to a question or input from a user.

[1573] A "multilingual response means" is a method or system for providing information or responding to questions in multiple languages.

[1574] "Means for providing product descriptions and recommended information by voice" refers to a function or system for providing product descriptions and recommended information to users by voice.

[1575] "User terminal" refers to a computing device used by a user, such as a smartphone or tablet.

[1576] A "robot" refers to an automated mechanical device that operates automatically and performs specific tasks based on user instructions.

[1577] An embodiment of the present invention is described below: This system generates guide voices by animated characters at tourist spots or brick-and-mortar stores, and provides users with a multilingual interactive guide service.

[1578] 1. System Overview

[1579] The system consists of three main components:

[1580] 1.1 User terminal

[1581] The user terminals are personal devices such as smartphones and tablets. Using these terminals, users can select their favorite animated characters and tourist spots or brick-and-mortar stores, and then obtain information through the character's voice guidance. The terminals are connected to a server via the Internet.

[1582] 1.2 Server

[1583] The server plays a central role in the system and has the following functions:

[1584] Character voice generation: Generates guidance voices for designated animated characters. Utilizing a generative AI model, the voices are tailored to the personality of the character selected by the user.

[1585] Information acquisition: Obtain information about tourist spots and brick-and-mortar stores from the database.

[1586] Interactive Response: Generates optimal responses to user questions and translates them into multiple languages ​​as needed.

[1587] Guide plan generation: Generate a travel plan or product introduction plan as needed.

[1588] 1.3 Database

[1589] The database contains detailed information about tourist destinations and brick-and-mortar stores, including their history, culture, product features, and prices. The database is updated regularly.

[1590] 2. Details of the processing

[1591] The server proceeds in the following way:

[1592] 2.1 Character voice generation

[1593] A generative AI model is used to generate a voice for the selected character. For example, the gTTS (Google Text-to-Speech) library is used to generate an audio file to provide voice guidance. This audio file is sent to the user's device, where it can be played back.

[1594] 2.2 Information acquisition

[1595] The server retrieves necessary information from a database, such as historical information about tourist spots or product descriptions from physical stores, and generates voice guidance to provide to the user based on that information.

[1596] 2.3 Interactive Response

[1597] When a user enters a question in voice or text format, the server analyzes the question and generates the best possible response using a generative AI model. If necessary, the response is translated for multilingual support and sent to the user's device.

[1598] 3. Specific Examples

[1599] Here are some concrete examples of how the system can be used:

[1600] Example 1: Tourist information

[1601] The user selects a park as a tourist attraction and a specific anime character as a character. The server retrieves information about the park from a database and uses a generative AI model to generate a guide in the character's voice. When the user asks, "Tell me about the history of this park," the server generates an appropriate response to the question and sends it to the user's device. The user can then play the audio guide on their device and tour the park.

[1602] Prompt Sentence Examples

[1603] User: "I chose Pikachu."

[1604] App: "Pikachu's voice will begin prompting. What product would you like to know more about?"

[1605] User: "What product is 12345?"

[1606] App: "This item is of the highest quality. It costs 1000 yen."

[1607] User: "Where does this product originate from?"

[1608] App: "Origin: Japan."

[1609] In this way, users can enjoy being guided around tourist spots and brick-and-mortar stores together with animated characters.

[1610] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1611] Step 1:

[1612] Input: The user launches the application using their device and selects their favorite animated character and tourist destination or brick-and-mortar store.

[1613] Processing: The terminal receives as input information about the characters and tourist spots or brick-and-mortar stores selected by the user, and transmits the selected information to the server.

[1614] Output: The selection information sent to the server.

[1615] Specific operation: The user opens the smartphone app and taps to select a character and destination from the on-screen menu.

[1616] Step 2:

[1617] Input: The server receives the selection information sent by the user.

[1618] Processing: The server retrieves detailed information about tourist attractions or brick-and-mortar stores from the database based on the received selection information.

[1619] Output: Detailed information about a tourist attraction or brick-and-mortar store retrieved from the database.

[1620] What it does: The server queries a database to retrieve information about the destination, such as its history and product descriptions.

[1621] Step 3:

[1622] Input: Details retrieved from the database.

[1623] Processing: The server uses the generative AI model to generate a voice guide for the selected character based on the detailed information.

[1624] Output: Guidance voice data of the generated character.

[1625] Specific operation: Using a generative AI model (e.g., gTTS library), convert text information into the voice of a specified character and create an audio file.

[1626] Step 4:

[1627] Input: Generated guidance speech data.

[1628] Processing: The server transmits the generated voice data to the user terminal.

[1629] Output: Character guidance voice data sent to the user's terminal.

[1630] Specific operation: The server sends the audio file to the user's smartphone via the network.

[1631] Step 5:

[1632] Input: Voice data sent to the user's terminal.

[1633] Processing: The user terminal plays back the received guidance voice.

[1634] Output: The audio guidance that the user can hear.

[1635] Specific operation: The selected character's voice will play instructions through the smartphone speaker.

[1636] Step 6:

[1637] Input: The user types a question into the application using text or voice.

[1638] Processing: The terminal recognizes the user's question and sends it to the server.

[1639] Output: The user's question sent to the server.

[1640] Specific action: The user speaks a question or enters text using the device's microphone.

[1641] Step 7:

[1642] Input: The user's question received by the server.

[1643] Processing: The server analyzes the question and generates the best response using its database and generative AI models, translating it into multiple languages ​​if necessary.

[1644] Output: The generated response data.

[1645] Specific operation: The server analyzes the question using natural language processing technology, generates a corresponding answer, and then translates it into the specified language.

[1646] Step 8:

[1647] Input: The generated response data.

[1648] Processing: The server sends the generated response data to the user terminal.

[1649] Output: The response data sent to the user terminal.

[1650] Specific operation: The server sends response data to the user's smartphone via the network.

[1651] Step 9:

[1652] Input: Response data sent to the user terminal.

[1653] Processing: The user terminal plays back the answer in the character's voice based on the received response data.

[1654] Output: An interactive audio response that the user can hear.

[1655] Specific operation: The smartphone plays the received data as an audio file and provides the answer to the user.

[1656] This allows users to enjoy information and detailed information about tourist spots and brick-and-mortar stores through animated characters.

[1657] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1658] This invention is an interactive pilgrimage tour system using animated characters combined with an emotion engine that recognizes the user's emotions. This system not only provides detailed information about tourist spots, but also provides appropriate guidance and responses based on the user's emotions. It also supports multiple languages, making it suitable for foreign tourists.

[1659] System configuration

[1660] 1. User Device

[1661] Functions: Provides an interface for users to access the system and select characters and tourist attractions. Also plays audio guides for tourist attractions and displays and plays responses to user questions in real time. Furthermore, it has an emotion engine for recognizing emotions from user input (text and voice).

[1662] How it works: Users access a designated website or application using a device such as a smartphone or PC. They select a character and a destination on the interface, and if they have any questions, they input them using text or voice. The emotion engine analyzes the user's input and provides guidance accordingly.

[1663] 2. Server

[1664] Function: Receives character and destination selection information sent from the user's device and retrieves tourist spot information from a database. Uses a generative AI model to generate character voice and text. Additionally, responds interactively to user questions and supports multiple languages. Also, adjusts guidance and response content based on the user's emotional data.

[1665] Processing method: Based on the user's selection, the server retrieves relevant sacred site information (such as name, location, and history) from a database. It then uses a generative AI model to generate audio data for tour guidance and transmits it to the user's device. When the user enters a question, the server analyzes the question, generates an optimal response, and adjusts the response according to the emotion recognized by the emotion engine.

[1666] 3. Database

[1667] Function: Stores and manages detailed information about tourist spots (history, culture, episodes related to animation and manga, etc.). Provides necessary information in response to requests from the server.

[1668] Management method: The database is updated regularly to include new tourist destination information and information on animation and manga.

[1669] Specific examples

[1670] A specific example of the system is shown below.

[1671] Example 1: A park in Tokyo

[1672] 1. User selection: The user opens the smartphone application and selects Character A from the animation work and a certain park in Tokyo.

[1673] 2. Tour Guide: The server receives the selection information and retrieves detailed information about a certain park from a database. Using a generative AI model, it generates a narration in the voice of Character A, saying "This park is...", and sends it to the user's device. The user plays the narration on their device and tours the park.

[1674] 3. Question and Response: The user enters a question, such as "Tell me about the history of this place." The server analyzes the question and uses a generative AI model to generate a response, such as "This park was founded in ____...," which is sent to the device. The device then plays back the response in the voice of Character A.

[1675] 4. Emotion Recognition: When a user types "What a beautiful place," the emotion engine analyzes the emotion data and recognizes that the user is happy. The server generates a response based on the emotion, such as "I'm glad to hear that," and plays it on the device.

[1676] 5. Multilingual support: If the user has English settings, the server will use artificial intelligence to translate the response into English and send it to the device as the English voice of Character A.

[1677] This allows users to have a rich experience touring tourist spots while interacting with animated characters.The purpose of the present invention is to improve the quality of users' sightseeing experiences through such an interactive guide system.

[1678] The processing flow will be explained below.

[1679] Step 1:

[1680] A user starts an application on a terminal and accesses the system. The terminal generates initial session information and sends it to the server.

[1681] Step 2:

[1682] The terminal displays the animated characters and a list of tourist attractions on a user interface.

[1683] Step 3:

[1684] Through the interface, the user selects an animated guide character and the tourist sites they wish to visit.

[1685] Step 4:

[1686] The device sends the user's selection (character ID, destination ID, etc.) to the server.

[1687] Step 5:

[1688] The server retrieves information about the selected character and tourist spot from the database.

[1689] Step 6:

[1690] The server generates a tour plan based on the acquired information and determines the spots to visit and the content of the narration.

[1691] Step 7:

[1692] The server uses the generative AI model to generate narration text for the tour plan using the voice of the selected character, creating audio data.

[1693] Step 8:

[1694] The server transmits the generated voice data to the user's terminal.

[1695] Step 9:

[1696] The terminal starts playing the audio data received from the server as narration.

[1697] Step 10:

[1698] During the tour, the user can ask questions by text or voice, for example, "Tell me about something special about this place."

[1699] Step 11:

[1700] The terminal sends the user's question to the server.

[1701] Step 12:

[1702] The server analyzes the question and retrieves relevant information from a database.

[1703] Step 13:

[1704] The server uses the generative AI model to generate the analyzed information as voice data for the character, for example, "This park was established at ____."

[1705] Step 14:

[1706] The server transmits the generated response voice data to the terminal.

[1707] Step 15:

[1708] The terminal reproduces the response voice data received from the server.

[1709] Step 16:

[1710] When a user speaks, the emotion engine analyzes the user's input (text or voice) and recognizes the emotion. For example, if a user inputs "It's a very beautiful place," the device analyzes it and sends it to the emotion engine.

[1711] Step 17:

[1712] The emotion engine recognizes the user's emotion and sends the emotion data to the server. For example, it recognizes that the user is happy.

[1713] Step 18:

[1714] The server uses a generative AI model to generate an appropriate response based on the emotional data, such as "I'm glad to hear that."

[1715] Step 19:

[1716] The server transmits the generated emotion response voice data to the terminal.

[1717] Step 20:

[1718] The terminal reproduces the emotion response voice data received from the server.

[1719] Step 21:

[1720] The server checks the language specified by the user in the settings, translates the voice data into that language if necessary, and sends the translated voice to the device.

[1721] Step 22:

[1722] The device plays the translated audio data and delivers it to the user.

[1723] Example 2

[1724] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1725] Conventional tourist information systems provide only one-way guidance information, making it difficult to provide responses that take the user's emotions into consideration. Furthermore, they lack multilingual support, making it impossible to provide the same service to foreign tourists. Furthermore, interactive responses and real-time guidance generation are insufficient, leaving room for improvement in the user experience.

[1726] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for generating a guidance voice of an animated character, a means for acquiring information on tourist spots, a means for interactively responding to a user's questions, a means for providing responses in multiple languages, and a means for analyzing the user's emotions and adjusting the response based on the emotions. This makes it possible to provide interactive tourist information that takes the user's emotions into consideration in real time.

[1727] An "animated character" is a computer-generated character, such as a person or animal, that is displayed visually and interacts with the user through sound and movement.

[1728] "Guidance voice" is voice data for providing information on tourist spots and the like by voice, and is used to guide and explain to the user.

[1729] "Tourist destination information" refers to detailed data about a tourist destination, including its location, history, culture, and access methods.

[1730] An "interactive response means" is a device or program that has the function of responding in real time to questions or input from a user.

[1731] "Multilingual support" means providing services and responses in multiple languages ​​to users who speak different languages.

[1732] "Means for analyzing emotions" refers to technology or devices that identify emotions from user input or behavior and acquire them as data.

[1733] A "generative AI model" is an algorithm or program that uses artificial intelligence to generate text or speech.

[1734] A "database" is a system for efficiently storing and managing information, accumulating large amounts of data and allowing it to be quickly searched and retrieved as needed.

[1735] A "user terminal" is a device operated by a user, such as a smartphone, PC, or tablet.

[1736] A "server" is a computer system that receives requests from clients via a network and provides the necessary information.

[1737] This invention is an interactive tourist information system using animated characters combined with an emotion engine that recognizes the user's emotions. This system not only provides detailed information about tourist spots, but also provides appropriate guidance and responses based on the user's emotions. It also supports multiple languages, making it suitable for foreign tourists.

[1738] System configuration

[1739] User terminal

[1740] Users access a designated website or application using a device such as a smartphone or PC. Through this interface, users can select an animated character and a desired tourist destination and enter their request. The user device is equipped with an emotion engine that recognizes emotions from the user's input (text or voice). For example, if a user enters "I want to go to a certain park in Tokyo," the request is sent to the server.

[1741] server

[1742] The server receives the character and tourist attraction selection information sent from the user's device. The server accesses the database to obtain detailed information about the relevant tourist attraction. It then uses a generative AI model to generate audio guidance data. At this time, guidance such as "This park was established in XX..." is generated in the voice of the selected character. This audio data is sent to the user's device and played back.

[1743] Database

[1744] The database stores and manages detailed information about tourist destinations (location, history, culture, etc.). This information is used by the server and provided to the service as needed. The database is also regularly updated with new tourist destination information and information about animations and comics.

[1745] Specific examples

[1746] A specific example of the system's operation is shown below.

[1747] Example 1: A park in Tokyo

[1748] 1. User selection: The user opens the smartphone application and selects animated character A and a certain park in Tokyo.

[1749] 2. Tour Guide: The server receives the selection information and retrieves detailed information about a certain park from a database. Using a generative AI model, it generates a narration in the voice of Character A, saying "This park is...", and sends it to the user's device. The user plays the narration on their device and tours the park.

[1750] 3. Question and Response: The user enters a question such as "Tell me about the history of this place." The server analyzes the question and uses a generative AI model to generate a response such as "This park was founded in ____..." and sends it to the device. The device plays the response in the voice of Character A.

[1751] 4. Emotion Recognition: When a user types "What a beautiful place," the emotion engine analyzes the emotion data and recognizes that the user is happy. The server generates a response based on the emotion, such as "I'm glad to hear that," and plays it on the device.

[1752] 5. Multilingual support: If the user has English settings, the server will translate the response into English using the generative AI model and send it to the device as Character A's English voice to play.

[1753] Prompt Sentence Examples

[1754] 1. "I'd like you to show me around a certain park in Tokyo."

[1755] 2. "Tell me about the history of this park."

[1756] 3. "It's a beautiful place."

[1757] 4. “Can you provide the information in English?”

[1758] Through such a system, users can have a rich experience touring tourist spots while interacting with animated characters. The present invention aims to improve the quality of users' sightseeing experiences.

[1759] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1760] Step 1:

[1761] Users open the application on their smartphone or PC and select an animated character and a tourist destination of their choice. The user's selection and request are entered and sent from the device to the server. Specifically, the user taps or clicks on an option on the screen, and the selection is collected by the application and transmitted to the server via the network.

[1762] Step 2:

[1763] The server receives the character and tourist attraction selection information sent by the user. The server then queries the database to obtain detailed information about the corresponding tourist attraction. Specifically, a database query is executed using the character and tourist attraction identifiers. The query results in data such as the tourist attraction's name, location, history, and culture.

[1764] Step 3:

[1765] The server uses a generative AI model to generate voice guidance data based on the acquired tourist attraction information and user request information. Using the acquired tourist attraction information and request information as input, it processes and calculates data. Character voice guidance data is generated as output. Specifically, the generative AI model creates narration from text information and converts it into voice data.

[1766] Step 4:

[1767] The server transmits the generated voice guidance data to the user terminal. Specifically, a data packet containing the voice guidance data is transmitted to the user terminal via the network. The user terminal then plays back the received voice guidance data.

[1768] Step 5:

[1769] As users travel around tourist spots, they input questions and comments through their devices. The user's input is in the form of text or voice. The user's device then sends the input data to the server. Specifically, when the user inputs a question and taps the send button, the data is collected and transferred to the server.

[1770] Step 6:

[1771] The server receives and analyzes questions from users. It analyzes the question data received as input and generates an optimal response using a generative AI model. The generated response data is obtained as output. Specifically, a natural language processing algorithm analyzes the question and generates an appropriate response.

[1772] Step 7:

[1773] The generated response data is sent from the server to the user terminal. Specifically, a data packet containing the response data is sent to the user terminal via the network. The user terminal then plays back the received response data in the character's voice.

[1774] Step 8:

[1775] The emotion engine analyzes the user's input data (text and voice) and extracts emotion data. It uses the user's comment data as input and performs data analysis. Emotion data is obtained as output. Specifically, text mining and voice analysis algorithms recognize emotions.

[1776] Step 9:

[1777] The server adjusts the response content based on the emotional data. Using the emotional data and generated response data as input, it processes the data again using the generative AI model. The output is response data based on the emotion. Specifically, an algorithm is executed to generate response content that reflects the emotional data.

[1778] Step 10:

[1779] If the user has multilingual settings, the server translates the generated response data into the specified language. Using the generated response data as input, it performs data calculations using an AI translation model. The translated response data is obtained as output. Specifically, the translation algorithm converts the response data into another language.

[1780] In this way, users can experience interactive tourist information that responds to their emotions through animated characters. In addition, the multilingual support makes it easy for foreign tourists to use the service.

[1781] (Application example 2)

[1782] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1783] In recent years, the number of interactive systems that provide information about tourist destinations and restaurants has increased. However, current systems are unable to provide optimal responses based on user emotions. This results in problems such as not being able to provide the information and services users desire in a timely and appropriate manner. Furthermore, the lack of multilingual support makes these systems difficult for foreign tourists to use. To solve these issues, a system that analyzes user emotions and provides optimal recommendations based on those emotions is needed.

[1784] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1785] In this invention, the server includes a means for recognizing a user's emotions and adjusting the response content based on the emotions, a means for recommending optimal restaurants and dishes based on the emotions, and a means for generating recommendation text using a generative AI model, thereby enabling interactive responses according to the user's emotions.

[1786] The "means for generating guidance voice of an animated character" is a function for providing the user with information about tourist spots and restaurants in the voice of an animated character.

[1787] "Means for obtaining information on tourist destinations" refers to a function for collecting detailed data on tourist destinations from databases and external information sources.

[1788] The "means for interactively responding to user questions" is a function for analyzing questions from users and providing appropriate answers in real time.

[1789] "Means for providing responses in multiple languages" is a function for providing answers and guidance adapted to each language in scenarios requiring multilingual support.

[1790] "Means for recognizing the user's emotions and adjusting the response content based on those emotions" is a function for analyzing the emotions from the user's input content and voice and generating the optimal response accordingly.

[1791] The "means for recommending optimal restaurants and dishes based on emotions" is a function for taking into account the emotional state of the user and recommending restaurants and dishes that are suitable for them.

[1792] "Means for generating recommendation text using a generative AI model" is a function that utilizes an AI model to generate recommendation text based on a user's emotions and questions.

[1793] A "terminal" is a device operated by a user, such as a smartphone or tablet.

[1794] A "server" is a central processing unit that processes data from user terminals and generates and transmits necessary information.

[1795] This invention is an interactive food delivery application system that combines an emotion engine that recognizes the user's emotions. The system aims to recommend the most suitable restaurants and dishes to the user by linking the user terminal, server, and database.

[1796] System configuration

[1797] 1. User Device

[1798] Function: Provides an interface for users to access the application and select restaurants and dishes. It also recognizes user emotions from input (text and voice) and has an emotion engine. It also displays and plays recommendations based on emotions.

[1799] How it works: Users access the designated application using a smartphone or other device. They select a restaurant and a dish on the interface, and input their questions or preferences via text or voice. The emotion engine analyzes the emotions from these inputs and makes food and drink recommendations accordingly.

[1800] 2. Server

[1801] Function: Receives sentiment analysis results and question data sent from the user's device and retrieves relevant dining information from the database. Uses a generative AI model to generate recommendation text based on the user's sentiment and question and sends it to the user's device. Also supports multiple languages, providing translated responses into foreign languages.

[1802] Processing method: The server retrieves relevant restaurant information (store information, dish details, reviews, etc.) from a database based on the user's sentiment analysis results and question data. It then uses a generative AI model to generate optimal recommendation text and sends it to the user's device. The recommendation content is adjusted based on the sentiment recognized by the emotion engine.

[1803] 3. Database

[1804] Function: Stores and manages detailed information about restaurants and dishes (menus, prices, nutritional information, user reviews, etc.). Provides necessary information in response to requests from the server.

[1805] How it's maintained: The database is updated regularly to include information on new restaurants and cuisines.

[1806] Specific examples

[1807] A specific example of the system is shown below.

[1808] Example 1: Food and drink recommendations when the user is tired

[1809] 1. User makes a choice: The user opens the application on their smartphone and types "I'm tired."

[1810] 2. Emotion recognition and recommendation: The server receives the user's input, and the emotion engine analyzes the emotion "tired." Using the generative AI model, a recommendation text is generated, such as "We recommend the stamina bowl at nearby Restaurant A, which is a nutritious dish," and sent to the user's device.

[1811] 3. Multilingual support: If the user selects English, the generated recommendation text is translated into English and sent to the device, enabling multilingual support.

[1812] Example prompts for generative AI models

[1813] "What foods would you recommend for users when they're feeling low?"

[1814] "What food can cheer you up when you're feeling tired?"

[1815] In this way, through the system of the present invention, users can receive recommendations for optimal food and drink based on their emotions and enjoy a satisfying eating and drinking experience.

[1816] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1817] Step 1:

[1818] The user launches the smartphone application and inputs their question or request via text or voice.

[1819] Input: User text or voice input (e.g., "I'm tired today and want to eat something energizing.")

[1820] Output: The app gets the user's input data.

[1821] Specific actions: The user inputs their feelings and questions using the dedicated app interface, then presses the "Send" button.

[1822] Step 2:

[1823] The device sends the user's input data to an emotion recognition engine, which analyzes the emotion.

[1824] Input: User text or voice data

[1825] Output: Emotion recognition result (e.g. "tired")

[1826] Specific operation: The app's built-in emotion recognition engine analyzes the user's input and extracts emotions such as "tired" as labels.

[1827] Step 3:

[1828] The device sends the emotion recognition results to the server and requests optimal food and drink recommendations.

[1829] Input: Emotion recognition result (e.g., "tired")

[1830] Output: Send emotion recognition results to the server

[1831] Specific operation: The emotion recognition results are sent to the server, and the app calls the server's API using a network connection.

[1832] Step 4:

[1833] The server retrieves relevant dining information from the database.

[1834] Input: Emotion recognition results to the server

[1835] Output: Related food and drink information (e.g., food items such as "Stamina bowl")

[1836] Specific operation: The server queries the database and obtains food and drink information (e.g., foods that provide stamina) related to the emotion recognition results.

[1837] Step 5:

[1838] The server uses a generative AI model to generate recommendation text based on the user's sentiment and question.

[1839] Input: Food and drink information obtained from the database, emotion recognition results

[1840] Output: Recommendation text (e.g. "You seem tired today. I recommend a stamina bowl to give you energy.")

[1841] Specific operation: The server uses a generative AI model such as OpenAI's GPT-4, inputs food and drink information and emotion recognition results as prompts, and generates recommendation text.

[1842] Step 6:

[1843] The server transmits the generated recommendation text to the user terminal.

[1844] Input: Recommendation text

[1845] Output: Send recommendation text to user device

[1846] Specific operation: The server sends the generated recommendation text to the user terminal via the network.

[1847] Step 7:

[1848] The terminal displays or plays the recommendation text to the user.

[1849] Input: Recommendation text

[1850] Output: Display to user or play as audio

[1851] What happens: The app displays the suggested text on the screen or plays it aloud using a text-to-speech engine.

[1852] Through these steps, users can receive optimal food and drink recommendations tailored to their emotional state. Using a generative AI model, we can provide more natural and specific recommendation text.

[1853] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1854] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1855] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1856] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1857] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1858] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1859] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1860] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1861] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1862] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1863] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1864] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1865] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1866] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1867] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1868] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1869] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1870] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1871] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1872] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1873] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1874] The following is further disclosed regarding the above embodiment.

[1875] (Claim 1)

[1876] means for generating a guidance voice of an animated character;

[1877] A means for obtaining information on tourist destinations;

[1878] means for interactively responding to user questions;

[1879] A system including a means for providing responses in multiple languages.

[1880] (Claim 2)

[1881] A means for generating guidance voices of animated characters in real time;

[1882] A means for obtaining information about tourist destinations from a database;

[1883] means for analyzing user questions and generating optimal responses;

[1884] 10. The system of claim 1, further comprising means for translating the generated response into a specified language.

[1885] (Claim 3)

[1886] A means for the terminal to transmit selection information of the animation character and the tourist spot to the server;

[1887] a means for the server to generate a tour plan based on the destination;

[1888] a means for generating a voice for the character using a generative AI model;

[1889] 2. The system of claim 1, further comprising means for returning a response to a user's question through the character.

[1890] "Example 1"

[1891] (Claim 1)

[1892] a means for a user to select a character and a destination;

[1893] means for transmitting the selection information to a server;

[1894] A means for the server to obtain destination information from a database;

[1895] A means for generating a voice guidance for a character using a generative AI model;

[1896] means for responding to user queries;

[1897] The system includes a means for translating the generated response into multiple languages.

[1898] (Claim 2)

[1899] A means for generating guidance voices of animated characters in real time;

[1900] a means for retrieving destination information from a database;

[1901] means for analyzing user questions and generating optimal responses;

[1902] 10. The system of claim 1, further comprising means for translating the generated response into multiple languages.

[1903] (Claim 3)

[1904] A means for the terminal to transmit character and destination selection information to the server;

[1905] a means for generating a guidance plan based on a destination by the server;

[1906] a means for generating a voice for the character using a generative AI model;

[1907] 2. The system of claim 1, further comprising means for returning a response to a user's question through the character.

[1908] "Application Example 1"

[1909] (Claim 1)

[1910] means for generating a guidance voice of an animated character;

[1911] A means of obtaining information about tourist spots or brick-and-mortar stores;

[1912] means for interactively responding to user questions;

[1913] A means of providing multilingual support,

[1914] A designated character will provide audio guidance on product descriptions and recommendations within the store,

[1915] A system including a means for a user terminal or a robot to interactively respond to questions about products.

[1916] (Claim 2)

[1917] A means for generating guidance voices of animated characters in real time;

[1918] A means for obtaining information about tourist attractions or brick-and-mortar stores from a database;

[1919] means for analyzing user questions and generating optimal responses;

[1920] 10. The system of claim 1, further comprising means for translating the generated response into a specified language.

[1921] (Claim 3)

[1922] A means for the terminal to transmit selection information of the animated character and the tourist spot or the brick-and-mortar store to the server;

[1923] A means for the server to generate a guide plan based on a destination or product description;

[1924] a means for generating a voice for the character using a generative AI model;

[1925] 2. The system of claim 1, further comprising means for returning a response to a user's question through the character.

[1926] "Example 2: Combining Emotion Engines"

[1927] (Claim 1)

[1928] means for generating a guidance voice of an animated character;

[1929] A means for obtaining information on tourist destinations;

[1930] means for interactively responding to user questions;

[1931] A means of providing multilingual support,

[1932] means for analyzing a user's emotions and adjusting a response based on the emotions;

[1933] A system including:

[1934] (Claim 2)

[1935] A means for generating guidance voices of animated characters in real time;

[1936] A means for obtaining information about tourist destinations from a database;

[1937] means for analyzing user questions and generating optimal responses;

[1938] means for translating the generated response into a specified language;

[1939] A means of recognizing emotions from user input using text and voice;

[1940] 10. The system of claim 1, comprising:

[1941] (Claim 3)

[1942] A means for the terminal to transmit selection information of the animation character and the tourist spot to the server;

[1943] a means for the server to generate a tour plan based on the destination;

[1944] a means for generating a voice for the character using a generative AI model;

[1945] A means for returning a response to a user's question through a character;

[1946] A means for adjusting response content based on user emotion data;

[1947] 10. The system of claim 1, comprising:

[1948] "Application example 2 when combining emotion engines"

[1949] (Claim 1)

[1950] means for generating a guidance voice of an animated character;

[1951] A means for obtaining information on tourist destinations;

[1952] means for interactively responding to user questions;

[1953] A means of providing multilingual support,

[1954] means for recognizing a user's emotion and tailoring responses accordingly;

[1955] A method for recommending the best restaurants and dishes based on emotions,

[1956] a means for generating recommendation text using a generative AI model;

[1957] The system includes means for transmitting the generated response to the terminal.

[1958] (Claim 2)

[1959] A means for generating guidance voices of animated characters in real time;

[1960] A means for obtaining information about tourist destinations from a database;

[1961] means for analyzing user questions and generating optimal responses;

[1962] means for translating the generated response into a specified language;

[1963] 10. The system of claim 1, further comprising means for analyzing a user's emotions and adjusting facilitation based thereon.

[1964] (Claim 3)

[1965] A means for the terminal to transmit selection information of the animation character and the tourist spot to the server;

[1966] a means for the server to generate a tour plan based on the destination;

[1967] a means for generating a voice for the character using a generative AI model;

[1968] A means for generating recommendation text according to user emotions;

[1969] 2. The system of claim 1, further comprising means for returning a response to a user's question through the character. [Explanation of symbols]

[1970] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for generating a guidance voice of an animated character; A means for obtaining information on tourist destinations; means for interactively responding to user questions; A system including a means for providing responses in multiple languages.

2. A means for generating guidance voices of animated characters in real time; A means for obtaining information about tourist destinations from a database; means for analyzing user questions and generating optimal responses; 10. The system of claim 1, further comprising means for translating the generated response into a specified language.

3. A means for the terminal to transmit selection information of the animation character and the tourist spot to the server; a means for the server to generate a tour plan based on the destination; a means for generating a voice for the character using a generative AI model; 2. The system according to claim 1, further comprising means for returning a response to a user's question through the character.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A