system

The system addresses the limitations of conventional metaverse agents by using user authentication, generative AI, and natural language processing to create personalized and sophisticated dialogue scenarios, improving user experience through flexible and effective interactions.

JP2026064758APending Publication Date: 2026-04-14SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-02
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Conventional metaverse agents are limited to fixed conversation scenarios and information provision, making it difficult to respond flexibly and personally to diverse user desires and actions, and the user's action history and preferences are not fully utilized, leading to a limited user experience.

Method used

A system that includes user authentication, autonomous agent generation using generative AI, natural language processing, personalization based on user input analysis, and response generation, incorporating user behavior history to create personalized and sophisticated dialogue scenarios.

Benefits of technology

Enables more natural and effective conversations by personalizing information based on user preferences and past behavior, enhancing the user experience in metaverse platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026064758000001_ABST
    Figure 2026064758000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] Means of performing user authentication, A means for generating an autonomous agent, A means of analyzing user input using natural language processing, A means of personalizing information based on analysis results, A means of generating an agent's response based on personalized information, A system including means for sending the generated response to a user terminal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Agents in conventional metaverses are limited to fixed conversation scenarios and information provision, so there is a problem that it is difficult to respond flexibly and personally to diverse desires and actions of users. There is also a problem that the user's action history and preferences cannot be fully utilized, and the user experience is limited. It is an object of the present invention to solve such problems and provide a system capable of more natural and effective conversations and personalized information provision.

Means for Solving the Problems

[0005] The present invention solves the above problems with a system that includes means for user authentication, means for generating an autonomous agent, means for analyzing user input using natural language processing, means for personalizing information based on the analysis results, means for generating an agent response based on the personalized information, and means for transmitting the generated response to the user terminal. In addition, by further including means for recording the user's behavior history and using it for agent generation and response personalization, advanced personalization based on the user's preferences and past behavior becomes possible. Furthermore, by including means for generating an autonomous agent using generative artificial intelligence, more natural and sophisticated dialogue scenarios can be realized.

[0006] User authentication is the process of verifying the identification information used by a user when logging into a system and granting them access permission.

[0007] An "autonomous agent" is a character or program with artificial intelligence designed to interact with users and automatically perform specific tasks.

[0008] "Generative AI" is a type of artificial intelligence technology that generates new content based on given data or patterns.

[0009] "Natural language processing" is a technology that analyzes natural language text and speech input by users to understand their meaning.

[0010] Personalization refers to customizing the information and experiences provided to users, taking into account their individual preferences and past behavior.

[0011] "Response generation" is the process of generating appropriate responses to user input and engaging in dialogue with the user.

[0012] A "user terminal" is a device that a user directly operates (e.g., a personal computer, smartphone, or tablet).

[0013] "Behavior history" refers to a record of operations and actions a user has performed in the past.

[0014] A "session ID" is a unique identifier used to identify a session between a user and the system.

[0015] A "database" is a system for efficiently storing, searching, and managing information. [Brief explanation of the drawing]

[0016] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12]It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.

Mode for Carrying Out the Invention

[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0020] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0021] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0024] [First Embodiment]

[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0037] This invention is a system in which an autonomous agent in the metaverse interacts with a user in natural language and provides information based on the user's actions and requests. This invention is realized through the processes of user authentication, agent generation, natural language processing, information personalization, response generation, and response transmission.

[0038] An example of this system is described in detail below.

[0039] User Authentication

[0040] When a user accesses the metaverse platform, they are first presented with a login screen. The user enters their user ID and password and submits them to the server. The server verifies this authentication information against its database, and if the credentials are correct, it grants the user access. Upon successful authentication, the server generates a session ID and retrieves the user's profile information.

[0041] Generation of autonomous agents

[0042] The server references the user's profile information and past behavior history, and generates an autonomous agent using generative AI. This agent holds conversational scenarios based on the user's preferences and interests. Information about the generated agent is stored in a database.

[0043] Natural Language Processing (NLP)

[0044] When a user asks an agent in the metaverse a question such as, "Hello, can you tell me about the latest smartphones?", that input is sent from the device to the server. The server uses a natural language processing engine to analyze the user input and extract keywords and intent.

[0045] Personalization of information

[0046] The server personalizes information based on extracted keywords and intent, referencing the user's past behavior history and preferences. Specifically, if a user has previously preferred purchasing smartphones of a particular brand, the server will prioritize generating information about the latest models of that brand.

[0047] Response generation

[0048] Based on personalized information, the server generates a natural language response from the agent. For example, it might say, "Hello! For the latest smartphone, I recommend the latest model from BrandX. It has excellent camera performance and is reasonably priced."

[0049] Sending and displaying responses

[0050] The generated response is sent from the server to the user's terminal, which then displays it to the user. The user can then review the response and continue interacting with the agent.

[0051] Specific example

[0052] For example, consider a scenario where a user is searching for information about a specific product in a shopping mall within the metaverse. When the user asks an agent, "Tell me about the new game console," the server provides personalized information based on the user's past behavior and preferences. In this case, the agent might respond, "The GameConsoleX is a popular new game console. It features realistic graphics and fast loading times."

[0053] Thus, the system of the present invention appropriately provides the information that users need and realizes an advanced user experience within the metaverse.

[0054] The following describes the processing flow.

[0055] Step 1:

[0056] The user opens the login screen for the metaverse platform using their device. The user enters their user ID and password.

[0057] Step 2:

[0058] The terminal sends the authentication information entered by the user to the server.

[0059] Step 3:

[0060] The server compares the received authentication information with the database, and if they match, it generates a session ID and grants the user permission to log in.

[0061] Step 4:

[0062] The server retrieves user profile information from the database and stores it as session information.

[0063] Step 5:

[0064] The server generates an autonomous agent using generative AI based on the user's profile information and past behavioral history. The information about the generated agent is then stored in a database.

[0065] Step 6:

[0066] The device displays an interface within the metaverse for the user to approach the agent.

[0067] Step 7:

[0068] The user enters "Hello, please tell me about the latest smartphones" into the agent via the interface.

[0069] Step 8:

[0070] The terminal sends the user's input text to the server.

[0071] Step 9:

[0072] The server passes the received text to a natural language processing (NLP) engine, which analyzes keywords and intent (e.g., "smartphone," "latest," "tell me," etc.).

[0073] Step 10:

[0074] Based on the analysis results, the server identifies the category of information the user is looking for (new smartphone).

[0075] Step 11:

[0076] The server uses generative AI to create suitable information by referencing the user's past behavior history and preferences.

[0077] Step 12:

[0078] The server generates natural language responses for the agent based on the generated recommendation information.

[0079] Step 13:

[0080] The server sends the agent's response to the user's terminal.

[0081] Step 14:

[0082] The terminal displays the received agent response in the user interface.

[0083] Step 15:

[0084] The user can review the agent's response and then ask further questions or take further action.

[0085] (Example 1)

[0086] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0087] Traditional metaverse platforms have suffered from insufficient information provision and a poor quality of conversational experience for users. In particular, the lack of features to personalize information based on user behavior history and individual preferences prevented users from quickly obtaining the information they needed. Furthermore, the immaturity of technologies for analyzing natural language input and generating responses that matched user intent limited the naturalness and usefulness of conversations. As a result, users were dissatisfied with their experience within the metaverse, making long-term continued use difficult.

[0088] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0089] In this invention, the server includes means for user authentication, means for generating an autonomous agent, means for analyzing user input using natural language processing, means for personalizing information based on the analysis results, means for generating agent responses based on the personalized information, means for sending the generated responses to the user terminal, means for generating an agent by referring to the user's profile information and behavioral history, means for maintaining a dialogue scenario with the agent, means for generating natural language responses for the agent using a generative AI model, and means for extracting keywords and intentions using a natural language processing engine. This makes it possible to provide personalized information based on the user's individual preferences and past behavioral history. Furthermore, by accurately analyzing user input in natural language dialogue and generating responses based on that analysis, it becomes possible to provide a more natural and useful dialogue experience.

[0090] User authentication is the process by which a user verifies that they are a legitimate user when accessing a system, using authentication information such as a user ID and password.

[0091] An "autonomous agent" is an artificial intelligence agent generated using a generative AI model, which maintains dialogue scenarios based on the user's profile information and behavioral history, and enables dialogue in natural language.

[0092] "Natural language processing" is a technology in which a server analyzes the natural language input by a user and extracts keywords and intent.

[0093] "Personalization" refers to individually optimizing information and responses based on a user's past behavior history and preferences.

[0094] A "generative AI model" is an artificial intelligence model trained on a large dataset, such as GPT-4 (registered trademark), which performs natural language generation.

[0095] A "session ID" is an identifier generated by the server to uniquely identify a user's session while they are logged into the system.

[0096] A "natural language processing engine" is software or a library that a server uses to analyze a user's natural language input and extract keywords and intent.

[0097] A "dialogue scenario" is a set of dialogue patterns and rules that an agent maintains in order to facilitate smooth conversations with users.

[0098] "User profile information" refers to data that includes the user's basic attribute information (e.g., name, age, interests, etc.).

[0099] "Behavioral history" refers to data on past actions and preferences that a user has taken within the system.

[0100] This invention is a system in which an autonomous agent in the metaverse interacts with a user in natural language and provides information based on the user's actions and requests. This invention is realized through the following processes: user authentication, agent generation, natural language processing (NLP), information personalization, response generation, and response transmission.

[0101] This system consists of the following hardware and software.

[0102] Hardware: High-performance cloud servers (e.g., AWS® EC2, Google® Cloud Compute Engine), user devices (e.g., PCs, smartphones)

[0103] Software: Relational database management systems (e.g., MySQL®, PostgreSQL), open-source NLP libraries (e.g., SpaCy, NLTK), generative AI models (e.g., GPT-4)

[0104] The specific form for implementing this system is as follows:

[0105] User authentication:

[0106] When a user accesses the metaverse platform, the server first displays a login screen. The user enters their user ID and password and sends them from their terminal to the server. The server verifies this authentication information against a relational database (e.g., MySQL), and if the authentication information is correct, it authenticates successfully and generates a session ID.

[0107] Agent generation:

[0108] The server retrieves the user's profile information and past behavior history from the database. Next, the server generates an autonomous agent using a generative AI model (e.g., GPT-4). This agent maintains conversational scenarios based on the user's preferences and interests, and the generated agent's information is stored in the database.

[0109] Natural Language Processing (NLP):

[0110] When a user enters a question into an agent within the metaverse, the content of that question is sent from the user's device to the server. The server analyzes the received user input using a natural language processing engine (e.g., SpaCy) to extract keywords and the user's intent.

[0111] Personalizing information:

[0112] The server retrieves data on the user's past behavior and preferences from a database, based on keywords and intents extracted by the NLP engine. The server uses this data to optimize and personalize the information for the user.

[0113] Response generation:

[0114] The server generates natural language responses using a generative AI model (e.g., GPT-4) based on personalized information. For example, it might generate a response like, "For the latest smartphone, I recommend the latest model from BrandX. It has excellent camera performance and is reasonably priced."

[0115] Sending and displaying responses:

[0116] The generated response is sent from the server to the user's terminal, which then displays it to the user. The user can then review this response and continue interacting with the agent.

[0117] Specific example:

[0118] Consider a scenario where a user is searching for information about a specific product in a shopping mall within the metaverse. When the user asks an agent, "Tell me about the new game console," the server provides personalized information based on the user's past behavior and preferences. In this case, the agent might respond, "The GameConsoleX is a popular new game console. It features realistic graphics and fast loading times."

[0119] Examples of prompts to input into a generative AI model:

[0120] 1. "Generate a response for when a user is looking for information about the latest smartphones."

[0121] 2. "Generate the dialogue in which the agent explains the new game console to the user."

[0122] This enables the system of the present invention to appropriately provide the information that users need and to realize an advanced user experience within the metaverse.

[0123] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0124] Step 1: User Authentication

[0125] The user accesses the metaverse platform using their device and enters their user ID and password on the login screen. This information is sent from the device to the server. The server verifies the received user ID and password against its database. If the authentication information is correct, the server authenticates successfully, generates a session ID, and retrieves the user's profile information. This retrieved profile information is used in the next step.

[0126] Input: User ID, Password

[0127] Output: Authentication result, session ID, profile information

[0128] Step 2: Agent Generation

[0129] The server retrieves user profile information and past behavioral history from a database. The server uses a generative AI model (e.g., GPT-4) to generate an autonomous agent based on the user's preferences and interests. This agent maintains conversational scenarios, and information about the generated agent is stored in the database.

[0130] Input: Profile information, past activity history

[0131] Output: Generated agent information

[0132] Step 3: Natural Language Processing (NLP)

[0133] The user enters a question into the agent within the metaverse. For example, a question such as "Tell me about the latest smartphones" is sent from the terminal to the server. The server analyzes the received user input using a natural language processing engine (e.g., SpaCy) to extract keywords and the user's intent.

[0134] Input: User's question

[0135] Output: Analyzed keywords, user intent

[0136] Step 4: Personalizing Information

[0137] The server retrieves past user behavior history and preference data from a database based on keywords and intentions extracted by the NLP engine. Using this data, the server generates information optimized for the user. For example, a user who has previously preferred purchasing smartphones of a particular brand will be prioritized in receiving information about the latest models of that brand.

[0138] Input: Analyzed keywords, user intent, behavioral history, preference data

[0139] Output: Personalized information

[0140] Step 5: Generating a response

[0141] The server generates natural language responses using a generative AI model (e.g., GPT-4) based on personalized information. For example, it might generate a response like, "For the latest smartphone, I recommend the latest model from BrandX. It has excellent camera performance and is reasonably priced."

[0142] Input: Personalized information

[0143] Output: Generated natural language response

[0144] Step 6: Sending and displaying the response

[0145] The server sends the generated response to the user's terminal. The terminal displays the received response to the user. The user can then review the response and continue interacting with the agent.

[0146] Input: Generated natural language response

[0147] Output: Response to the user

[0148] These steps enable the provision of highly personalized information to users within the metaverse, thereby improving the user experience.

[0149] (Application Example 1)

[0150] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0151] Modern brick-and-mortar stores need effective and timely ways to provide customers with the product information they need. However, in many cases, staff interaction and existing digital signage alone are insufficient to meet individual customer needs. This often leads to decreased customer satisfaction and missed sales opportunities. Furthermore, there is a lack of personalized information delivery systems that effectively utilize users' past purchase history and profile information. Therefore, there is a need to develop new interactive information delivery systems aimed at improving the customer experience and increasing store sales.

[0152] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0153] In this invention, the server includes means for user authentication, means for generating an autonomous agent, means for analyzing user input using natural language processing, means for personalizing information based on the analysis results, means for generating agent responses based on the personalized information, means for displaying the generated responses on a smart device, means for personalizing the agent using the user's past purchase history and profile information, means for customers to interact with the agent using a smart device in a store, and means for displaying the agent's responses on the smart device's display. This enables customers to obtain information tailored to their individual needs in real time through a smart device in a physical store, improving the customer experience and increasing store sales.

[0154] User authentication is the process of verifying the legitimacy of a user accessing a system.

[0155] An "autonomous agent" is a dialogue system that is generated based on the user's profile and past behavioral history, and automatically responds to the user's questions and requests.

[0156] "Natural language processing" is a technology that analyzes the language and text spoken by a user to understand its meaning and intent.

[0157] "Personalizing information" means tailoring the information provided to each user individually based on their past behavioral history and profile information.

[0158] "Agent response generation" is the process by which an autonomous agent creates a response to a user based on personalized information.

[0159] A "smart device" is a portable electronic device with advanced computing capabilities and connectivity, and it is a device that has an interface with the user.

[0160] "Purchase history" refers to a record of products and services that a user has purchased in the past.

[0161] "Profile information" refers to data about a user's personal information, behavioral patterns, and preferences.

[0162] "Customers interacting with agents using smart devices in-store" means that customers in a physical store use devices such as smartphones or smart glasses to communicate with agents via voice or text.

[0163] A "display" is a device used to visually display information.

[0164] This invention is a system in which customers in a physical store can use a smart device to interact with an autonomous agent in natural language and obtain personalized information.

[0165] Hardware and software to be used

[0166] Hardware: Smart glasses (e.g., Google Glass®)

[0167] Software: Server-side: Node.js, Natural Language Processing Engine: Google Cloud Natural Language API, Database: MongoDB, User Interface: React Native

[0168] Generative AI Model: Generative artificial intelligence of OpenAI (registered trademark)

[0169] Process Overview

[0170] 1. User Authentication:

[0171] The user wears smart glasses and enters their user ID and password through the login screen. The authentication information is sent to the server, which verifies the legitimacy of the login by comparing it with the database. If successful, the server issues a session ID and retrieves the user's profile information.

[0172] 2. Generation of autonomous agents:

[0173] The server matches the user's profile information with their past purchase history and uses a generative AI model to generate an autonomous agent optimized for the user. This agent maintains conversational scenarios based on the user's preferences and past behavior.

[0174] 3. Natural Language Processing:

[0175] When a user speaks to the agent through smart glasses, the voice input is sent to a server. The server uses the Google Cloud Natural Language API to analyze the user's speech and extract keywords and intent.

[0176] 4. Personalizing information:

[0177] Based on extracted keywords and intent, the server personalizes information by referencing the user's past purchase history and profile information. For example, if a user has previously purchased a specific sporting item, the server will prioritize providing information about new products from that brand.

[0178] 5. Generating the response:

[0179] Based on personalized information, the server generates natural language responses from autonomous agents. The responses are tailored to the user's preferences and may include statements such as, "For the latest running shoes, I recommend the BrandX model. It's lightweight and has excellent cushioning."

[0180] 6. Display of response:

[0181] The generated response is sent from the server to the smart glasses and displayed on the glasses' screen. The user can read this response and continue interacting with the agent.

[0182] Specific example

[0183] For example, if a customer in a physical store asks an agent through smart glasses, "Tell me about the latest cameras," the server will prioritize displaying camera information from their preferred brands based on their past purchase history. In this case, the agent might respond, "For the latest cameras, I recommend BrandX models. They offer high resolution and excellent performance even in low light."

[0184] Example of a prompt

[0185] "User profile: { 'interest': 'technology', 'preferredBrand': 'BrandX'}

[0186] User behavior history: [{ 'storeVisit': '2023-01-01', 'purchasedItems': ['BrandX Phone']}]

[0187] Generate a personalized agent response for the query 'Please tell me about the latest cameras' in Japanese.

[0188] This system allows customers to obtain information tailored to their individual needs in real time via smart devices within physical stores, which is expected to improve the customer experience and increase store sales.

[0189] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0190] Step 1:

[0191] User Authentication

[0192] Input: User ID and password

[0193] Process: The user wears smart glasses and enters their user ID and password on the login screen. This authentication information is sent from the device to the server. The server compares it with the authentication information stored in the database to verify the validity of the authentication.

[0194] Output: If authentication is successful, the server generates a session ID and retrieves the user's profile information. If authentication fails, an error message is returned.

[0195] Step 2:

[0196] Generation of autonomous agents

[0197] Input: User profile information and past purchase history

[0198] Processing: After successful authentication, the server references the user's profile information and past purchase history, and uses a generated AI model to create an autonomous agent suited to the user. The agent maintains conversational scenarios based on the user's preferences and past behavior.

[0199] Output: Data from the generated agent. This data is stored in a database on the server.

[0200] Step 3:

[0201] Natural Language Processing

[0202] Input: User voice input (e.g., "Tell me about the latest cameras.")

[0203] Processing: When a user speaks to the agent through smart glasses, the voice input is sent from the device to the server. The server analyzes the voice using the Google Cloud Natural Language API and extracts keywords and intent.

[0204] Output: Analyzed keywords and intent. This information will be used in the next step.

[0205] Step 4:

[0206] Personalization of information

[0207] Input: Analyzed keywords and intent, past purchase history, user profile information

[0208] Processing: Based on the analyzed keywords and intent, the server personalizes information by referencing past purchase history and user profile information. It prioritizes selecting content based on the user's preferences and interests.

[0209] Output: Personalized information. This information will be used in the next step.

[0210] Step 5:

[0211] Response generation

[0212] Input: Personalized information

[0213] Processing: The server uses a generative AI model to generate a natural language response based on personalized information. For example, it might create a response like, "For the latest camera, we recommend the BrandX model. It offers high resolution and excellent performance even in low light."

[0214] Output: The generated response. This response will be used in the next step.

[0215] Step 6:

[0216] Display of response

[0217] Input: Generated response

[0218] Processing: The generated response is sent from the server to the smart glasses. This response is displayed on the smart glasses' screen.

[0219] Output: The response displayed on the user's smart glasses. The user can review this response and continue interacting with the agent.

[0220] This series of processes allows users to receive individually personalized information from an autonomous agent in real time through smart glasses within a physical store.

[0221] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0222] This invention is a system in which an autonomous agent in the metaverse interacts with a user in natural language, recognizes not only the user's actions and requests but also their emotions, and provides personalized information based on that. This invention is realized through the following processes: user authentication, agent generation, natural language processing, information personalization, emotion recognition, response generation, and response transmission.

[0223] An example of this system is described in detail below.

[0224] User Authentication

[0225] When a user accesses the metaverse platform, they are first presented with a login screen. The user enters their user ID and password and submits them to the server. The server verifies this authentication information against its database, and if the credentials are correct, it grants the user access. Upon successful authentication, the server generates a session ID and retrieves the user's profile information.

[0226] Generation of autonomous agents

[0227] The server references the user's profile information and past behavior history, and generates an autonomous agent using generative AI. This agent holds conversational scenarios based on the user's preferences and interests. Information about the generated agent is stored in a database.

[0228] Natural Language Processing (NLP)

[0229] When a user asks an agent in the metaverse a question such as, "Hello, can you tell me about the latest smartphones?", that input is sent from the device to the server. The server uses a natural language processing engine to analyze the user input and extract keywords and intent.

[0230] Personalization of information

[0231] The server identifies the category of information the user is looking for (new smartphones) based on extracted keywords and intent. Furthermore, it personalizes the information by referring to the user's past behavior history and preferences. Specifically, if the user has previously preferred purchasing smartphones of a particular brand, the server will prioritize generating information about the latest models of that brand.

[0232] emotion recognition

[0233] The server passes user input data (text, voice, facial expressions) to the emotion engine, which analyzes the user's emotions. For example, if the system detects that the user is excited, it prepares a response appropriate to that emotion.

[0234] Response generation

[0235] Based on personalized information and sentiment recognition results, the server generates a natural language response for the agent. For example, it might say, "Hello! For the latest smartphone, I recommend the latest model from BrandX. It has excellent camera performance and is reasonably priced."

[0236] Sending and displaying responses

[0237] The generated response is sent from the server to the user's terminal, which then displays it to the user. The user can then review the response and continue interacting with the agent.

[0238] Specific example

[0239] For example, consider a scenario where a user is searching for information about a specific product in a shopping mall within the metaverse. When the user asks an agent, "Tell me about the new game console," the server provides personalized information based on the user's past behavior history, preferences, and the results of emotion recognition. In this case, the agent might respond, "The GameConsoleX is a popular new game console. It features realistic graphics and fast loading times. Also, your smile shows you're enjoying it, so please consider it."

[0240] Thus, the system of the present invention appropriately provides the information that the user needs, realizing an advanced user experience within the metaverse. By combining this with emotion recognition, even more natural and responsive dialogue becomes possible, providing a more user-centric interaction.

[0241] The following describes the processing flow.

[0242] Step 1:

[0243] The user uses their device to open the login screen for the metaverse platform and enters their user ID and password.

[0244] Step 2:

[0245] The terminal sends the authentication information entered by the user to the server.

[0246] Step 3:

[0247] The server compares the received authentication information with the database, and if they match, it generates a session ID and grants the user permission to log in.

[0248] Step 4:

[0249] The server retrieves user profile information from the database and stores it as session information.

[0250] Step 5:

[0251] The server generates an autonomous agent using generative AI based on the user's profile information and past behavioral history. The information about the generated agent is then stored in a database.

[0252] Step 6:

[0253] The device displays an interface within the metaverse for the user to approach the agent.

[0254] Step 7:

[0255] The user enters "Hello, please tell me about the latest smartphones" into the agent via the interface.

[0256] Step 8:

[0257] The terminal sends the user's input text to the server.

[0258] Step 9:

[0259] The server passes the received text to a natural language processing (NLP) engine, which analyzes keywords and intent (e.g., "smartphone," "latest," "tell me," etc.).

[0260] Step 10:

[0261] Based on the analysis results, the server identifies the category of information the user is looking for (new smartphone).

[0262] Step 11:

[0263] The server uses generative AI to create suitable information by referencing the user's past behavior history and preferences.

[0264] Step 12:

[0265] The server passes the user's input data to the emotion engine, which analyzes the user's emotions. For example, it determines emotions from text, tone of voice, facial expressions, and other factors.

[0266] Step 13:

[0267] The server generates natural language responses for the agent based on the results of emotion recognition and personalized information. For example, if the user is excited, it might prepare a response such as, "I can see you're having fun. This smartphone has a particularly good camera, so you can beautifully record your memories."

[0268] Step 14:

[0269] The server sends the response generated by the agent to the user's terminal.

[0270] Step 15:

[0271] The terminal displays the received agent response in the user interface.

[0272] Step 16:

[0273] The user can review the agent's response and continue with further questions or actions. For example, they might type, "That's great! What other features are there?"

[0274] (Example 2)

[0275] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0276] In modern society, there is a growing need for autonomous agents that can interact with users in natural language to improve the user experience within the metaverse. Conventional systems have the problem of not being able to adequately personalize user input and recognize emotions, and can only provide users with non-interactive and uniform information. This invention solves these problems and provides an advanced information provision system that takes into account the user's behavioral history and emotions.

[0277] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0278] In this invention, the server includes means for user authentication, means for generating an autonomous agent, means for analyzing user input using natural language processing, means for personalizing information based on the analysis results, means for recognizing emotions based on user input data, means for generating agent responses based on the personalized information and emotion recognition results, and means for transmitting the generated responses to the user terminal. This enables the provision of appropriate information and natural dialogue tailored to the user's individual preferences and emotions.

[0279] "User authentication" is a process of verifying the legitimacy of a user using a user ID and password when the user accesses the system.

[0280] "Autonomous agent" is an artificial intelligence agent that is generated based on a user's operation history and profile information and can conduct natural language conversations with the user.

[0281] "Natural language processing" is a technology that enables a computer to understand and analyze words (natural languages) commonly used by humans in daily life.

[0282] "Personalize information" means customizing the provided information based on an individual user's action history and preferences.

[0283] "Recognize emotions" means analyzing a user's input data (such as text, voice, expression, etc.) to identify the user's current emotional state.

[0284] "Generate agent responses" means that an autonomous agent creates appropriate responses expressed in natural language based on emotion recognition and personalized information.

[0285] "Generative artificial intelligence" is an advanced artificial intelligence that can generate natural language and other forms of data based on input data.

[0286] "User terminal" refers to electronic devices such as computers, smartphones, and tablets used by users.

[0287] This invention relates to a system in which a user interacts with an autonomous agent in natural language on a metaverse platform and provides personalized information based on the user's action history and emotions. This invention is implemented according to the following steps.

[0288] User authentication

[0289] When a user accesses the metaverse platform, the device displays a login screen. The user enters their user ID and password, which the device sends to the server. The server verifies the received authentication information against its database. If authentication is successful, the server generates a session ID and retrieves the user's profile information. This process utilizes common authentication and database management systems.

[0290] Generation of autonomous agents

[0291] The server references the user's profile information and past behavior history, and generates an autonomous agent using a generative AI (e.g., GPT-4). This agent holds conversational scenarios based on the user's preferences and interests. Information about the generated agent is stored in a database. Here, the generative AI API and various database management tools are used.

[0292] Natural Language Processing (NLP)

[0293] When a user asks an agent in the metaverse a question such as "Hello, can you tell me about the latest smartphones?", that input is sent from the device to the server. The server uses a natural language processing engine (e.g., NLTK or spaCy) to analyze the user input and extract keywords and intent. Specifically, grammatical and semantic analysis are performed.

[0294] Personalization of information

[0295] The server identifies the category of information the user is looking for (e.g., a new smartphone) based on extracted keywords and intent. Furthermore, it personalizes the information by referencing the user's past behavior history and preferences. For example, a user who has previously preferred purchasing smartphones from a specific brand will be prioritized in receiving information about the latest models from that brand. User profiling techniques and data analysis tools are utilized here.

[0296] emotion recognition

[0297] The server passes user input data (text, voice, facial expressions) to an emotion engine (for example, OpenAI's emotion recognition model) to analyze the user's emotions. For example, if the server detects that the user is excited, it prepares a response appropriate to that emotion. Specifically, voice analysis, facial expression recognition, and text analysis are used in an integrated manner.

[0298] Response generation

[0299] Based on personalized information and sentiment recognition results, the server generates natural language responses for the agent. For example, it might generate a response like, "Hello! For the latest smartphone, I recommend the latest model from BrandX. It has excellent camera performance and is reasonably priced." Generative AI and natural language generation (NLG) technologies are used here.

[0300] Sending and displaying responses

[0301] The generated response is sent from the server to the user's terminal, which then displays it to the user. The user can then review the response and continue interacting with the agent. Specific data communication methods include RESTful APIs using the HTTP / HTTPS protocol.

[0302] Examples of specific cases and prompt statements

[0303] For example, consider a scenario where a user is searching for information about a specific product in a shopping mall within the metaverse. When the user asks the agent, "Please tell me about the new game console," the server provides personalized information based on the user's past behavior history, preferences, and the results of emotion recognition. At this time, the agent responds with something like, "As the latest game console, the new model is very popular. It features realistic graphics and fast loading times. Also, since it can be seen that you are enjoying it from that smile, please do consider it."

[0304] Examples of prompt sentences:

[0305] "Hello, Agent. Please tell me about the latest smartphone."

[0306] "Please tell me about the new game console."

[0307] As described above, this system performs natural language processing on the user's questions, provides personalized information, and generates responses based on emotions, thereby realizing a high-level and natural conversation. By combining the user's behavior history and emotion recognition, it becomes possible to have an interaction that is more user-oriented.

[0308] The flow of the specific process in Example 2 will be described using FIG. 13.

[0309] Step 1:

[0310] User authentication

[0311] The terminal accesses the metaverse platform.

[0312] Input: User ID and password

[0313] Output: Display of the login screen

[0314] The terminal displays a login screen to the user, who then enters their user ID and password.

[0315] Specific operation: Display ID and password input fields using an HTML form.

[0316] The terminal sends the entered authentication information to the server.

[0317] Input: User input information

[0318] Output: Sending authentication information (HTTP request)

[0319] The server compares the received authentication information with the database to verify its accuracy.

[0320] Input: Authentication information

[0321] Output: Authentication result (success or failure)

[0322] Specific operation: Execute a database query and search for matching data.

[0323] Upon successful authentication, the server generates a session ID and retrieves the user's profile information.

[0324] Input: Authentication result

[0325] Output: Session ID and user profile

[0326] Specific operation: Generate a session ID using the session management system, and retrieve user information from the database.

[0327] Step 2:

[0328] Generation of autonomous agents

[0329] The server accesses the user's profile information and past activity history.

[0330] Input: User profile information, activity history

[0331] Output: Reference result

[0332] Specific operation: Execute database queries and retrieve relevant data.

[0333] The server generates autonomous agents using generative AI (e.g., GPT-4).

[0334] Input: Reference Result

[0335] Output: Autonomous agent

[0336] Specific operation: User information is passed to the generative AI API to generate an agent.

[0337] The information about the generated agents is stored in the database.

[0338] Input: Autonomous agent

[0339] Output: Agent logs in the database

[0340] Specific action: Execute a data insertion query to save agent information.

[0341] Step 3:

[0342] Natural Language Processing (NLP)

[0343] The user asks an agent a question within the metaverse.

[0344] Example: "Hello, could you tell me about the latest smartphones?"

[0345] Input: User's question

[0346] Output: Question data

[0347] The terminal sends this input to the server.

[0348] Input: User's question

[0349] Output: HTTP request (sends query data to the server)

[0350] Specific action: Send question data in JSON format.

[0351] The server uses a natural language processing engine (such as NLTK or spaCy) to analyze user input and extract keywords and intent.

[0352] Input: User's question data

[0353] Output: Analysis results (keywords, intent, etc.)

[0354] Specific operation: Calls the text analysis module to perform grammatical and semantic analysis.

[0355] Step 4:

[0356] Personalization of information

[0357] Based on the extracted keywords and intent, the server identifies the category of information the user is seeking.

[0358] Input: Analysis results

[0359] Output: Information Category

[0360] Specific operation: Uses a keyword matching algorithm.

[0361] The server further references the user's past behavior history and preferences to personalize the information.

[0362] Input: Information category, user behavior history, preferences

[0363] Output: Personalized information

[0364] Specific action: Execute the user profiling algorithm.

[0365] As a concrete example, users who have previously purchased a smartphone from a specific brand will be given priority in receiving information about the latest models from that brand.

[0366] Step 5:

[0367] emotion recognition

[0368] The server passes the user's input data to the emotion engine, which then analyzes the user's emotions.

[0369] Input: User's text, voice, and facial expression data

[0370] Output: Emotion analysis results

[0371] Specific actions: Input data into an emotion recognition model and identify emotions.

[0372] For example, if the system detects that the user is excited, it prepares a response that corresponds to that emotion.

[0373] Input: Sentiment analysis results

[0374] Output: Emotion-based response

[0375] Specific action: Apply an emotion filtering algorithm.

[0376] Step 6:

[0377] Response generation

[0378] Based on personalized information and sentiment recognition results, the server generates natural language responses for the agent.

[0379] Input: Personalized information, sentiment analysis results

[0380] Output: Natural language response

[0381] Specific operation: Utilizes the natural language generation (NLG) function of a generative AI.

[0382] For example, it can generate a response like, "Hello! For the latest smartphone, I recommend the latest model from BrandX. It has excellent camera performance and is reasonably priced."

[0383] Step 7:

[0384] Sending and displaying responses

[0385] The server sends the generated response to the terminal.

[0386] Input: Natural language response

[0387] Output: HTTP response (response data sent to the terminal)

[0388] Specific action: Send response data in JSON format.

[0389] The terminal displays this response to the user.

[0390] Input: Response data

[0391] Output: Screen display

[0392] Specific operation: Display the response content on the screen using HTML or GUI.

[0393] The user can review this response and continue interacting with the agent.

[0394] Through each of the steps described above, this system enables sophisticated information provision and natural dialogue tailored to the individual needs and emotions of the user.

[0395] (Application Example 2)

[0396] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0397] Conventional interactive agent systems in virtual stores lacked the ability to recognize user emotions, resulting in insufficient personalized information delivery. Consequently, they failed to adequately provide the information users sought, limiting the shopping experience. This invention aims to solve these problems by providing a system that analyzes user emotions and delivers more user-centric interactions.

[0398] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for user authentication, means for generating an autonomous agent, means for analyzing user input using natural language processing, means for personalizing information based on the analysis results, means for analyzing the user's emotions, means for generating an agent response based on the personalized information and the results of the emotion analysis, and means for transmitting the generated response to the user terminal. This makes it possible to provide personalized information that takes the user's emotions into consideration.

[0399] User authentication is the process of verifying whether a user attempting to access a system has the legitimate authority.

[0400] An "autonomous agent" is a program equipped with artificial intelligence that can generate natural language responses in response to user input and actions, and continue the conversation.

[0401] "Natural language processing" is a technology that enables computers to understand, analyze, and generate human language.

[0402] "Personalizing information" means selecting and providing the most relevant information to a user based on their past behavior history and preferences.

[0403] "Analyzing user emotions" refers to the technology of identifying the emotions a user is experiencing based on their statements, facial expressions, voice, and other factors.

[0404] "Generating agent responses" is the process of providing information to users in a natural conversational format based on analyzed and personalized information.

[0405] "Generative artificial intelligence" refers to artificial intelligence technology that has the ability to generate new information and content based on diverse data.

[0406] "User terminal" refers to devices used by the user, such as computers, smartphones, tablets, and head-mounted displays.

[0407] This invention provides an interactive agent system for improving the shopping experience in virtual stores. The system includes user authentication, autonomous agent generation, natural language processing, information personalization, sentiment recognition, and response generation and transmission as processes.

[0408] First, when a user accesses the virtual store, a login screen appears on the terminal. The user enters their ID and password and sends them to the server. The server checks the database and verifies the authentication information. If authentication is successful, the server generates a session ID and retrieves the user's profile information.

[0409] Next, the server uses generative artificial intelligence to generate autonomous agents. Based on the user's profile information and past behavioral history, it generates agents with dialogue scenarios tailored to the user's preferences and interests. The agent information is stored in a database.

[0410] When a user asks an agent a question in the metaverse, they might input something like, "Tell me about the new game console." This input is sent from the terminal to the server. The server uses a natural language processing engine (e.g., OpenAI GPT-3®) to analyze this input and extract keywords and intent.

[0411] Based on the analysis results, the server identifies the category of information the user is seeking. Furthermore, it personalizes the information by referring to the user's past behavior history and preferences. In this process, if the user prefers a particular brand, it will prioritize providing information about the latest products from that brand.

[0412] The server passes user input data (text, voice, facial expressions) to the emotion engine, which analyzes the user's emotions. For example, if the system detects that the user is excited, it prepares a response appropriate to that emotion.

[0413] Based on emotion recognition and personalized information, the server generates responses for the agent. For example, a response such as, "Among the latest game consoles, there is a popular one that features fast loading times and realistic graphics," might be generated.

[0414] The generated response is sent from the server to the user's terminal and displayed on the terminal. The user can then review this response and continue interacting with the agent.

[0415] Specific example

[0416] For example, if a user asks "Tell me about the new game console" in a virtual store, the system will provide appropriate information based on the user's preferences and browsing history, responding with something like, "It features realistic graphics and fast loading times."

[0417] Example of a prompt

[0418] "Analyze the following user input and extract key information and intent: I want to buy a new gaming console."

[0419] "Create a virtual shopping assistant with the following characteristics: User is interested in gaming and preferred brands are brand A and brand B."

[0420] The specific hardware used in this system includes smartphones, tablets, and head-mounted displays as user terminals, while the server requires a high-performance processor and a large-capacity database. For software, OpenAI GPT-3 is used for natural language processing, and third-party APIs are used for emotion recognition.

[0421] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0422] Step 1:

[0423] A user accesses a virtual store, and a login screen appears on their terminal. The user enters their user ID and password and sends them to the server. The server compares the entered authentication information with its database, and if the information is valid, it generates a session ID and retrieves the user's profile information.

[0424] Input: User ID, Password

[0425] Output: Authentication result, session ID, user profile information

[0426] Step 2:

[0427] The server uses generative artificial intelligence to generate autonomous agents based on the user's profile information and past behavioral history. The generated agents have conversational scenarios tailored to the user's preferences and interests, and this information is stored in a database.

[0428] Input: User profile information, past activity history

[0429] Output: Autonomous agent, dialogue scenario

[0430] Step 3:

[0431] The user enters a question within the virtual store. For example, an input such as "Tell me about the new game console" is sent from the terminal to the server. The server uses a natural language processing engine (e.g., OpenAI GPT-3) to analyze this input and extract keywords and intent.

[0432] Input: User's question (text)

[0433] Output: Keywords, intent

[0434] Step 4:

[0435] The server analyzes user input to identify the categories of information the user is seeking. Furthermore, it personalizes the information by referencing the user's past behavior and preferences. In this process, it prioritizes and organizes information that the user is likely to be interested in, such as the latest product information for a specific brand.

[0436] Input: Keywords, intent, past behavior history, preferences

[0437] Output: Personalized information

[0438] Step 5:

[0439] The server passes user input data (text, voice, facial expressions) to an emotion engine, which analyzes the user's emotions. For example, it can determine whether the user is excited based on the tone of their text or voice.

[0440] Input: User input data (text, voice, facial expressions)

[0441] Output: User's emotional state

[0442] Step 6:

[0443] The server generates agent responses based on personalized information and sentiment analysis results. This can produce responses such as, "As a modern gaming console, it features fast loading times and realistic graphics." These responses are adjusted according to the user's emotional state.

[0444] Input: Personalized information, user's emotional state

[0445] Output: Agent response (text)

[0446] Step 7:

[0447] The generated response is sent from the server to the user's terminal and displayed on the terminal. The user can then review this response and continue interacting with the agent.

[0448] Input: Agent's response (text)

[0449] Output: Display on the user terminal

[0450] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0451] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0452] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0453] [Second Embodiment]

[0454] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0455] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0456] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0457] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0458] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0459] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0460] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0461] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0462] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0463] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0464] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0465] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0466] This invention is a system in which an autonomous agent in the metaverse interacts with a user in natural language and provides information based on the user's actions and requests. This invention is realized through the processes of user authentication, agent generation, natural language processing, information personalization, response generation, and response transmission.

[0467] An example of this system is described in detail below.

[0468] User Authentication

[0469] When a user accesses the metaverse platform, they are first presented with a login screen. The user enters their user ID and password and submits them to the server. The server verifies this authentication information against its database, and if the credentials are correct, it grants the user access. Upon successful authentication, the server generates a session ID and retrieves the user's profile information.

[0470] Generation of autonomous agents

[0471] The server references the user's profile information and past behavior history, and generates an autonomous agent using generative AI. This agent holds conversational scenarios based on the user's preferences and interests. Information about the generated agent is stored in a database.

[0472] Natural Language Processing (NLP)

[0473] When a user asks an agent in the metaverse a question such as, "Hello, can you tell me about the latest smartphones?", that input is sent from the device to the server. The server uses a natural language processing engine to analyze the user input and extract keywords and intent.

[0474] Personalization of information

[0475] The server personalizes information based on extracted keywords and intent, referencing the user's past behavior history and preferences. Specifically, if a user has previously preferred purchasing smartphones of a particular brand, the server will prioritize generating information about the latest models of that brand.

[0476] Response generation

[0477] Based on personalized information, the server generates a natural language response from the agent. For example, it might say, "Hello! For the latest smartphone, I recommend the latest model from BrandX. It has excellent camera performance and is reasonably priced."

[0478] Sending and displaying responses

[0479] The generated response is sent from the server to the user's terminal, which then displays it to the user. The user can then review the response and continue interacting with the agent.

[0480] Specific example

[0481] For example, consider a scenario where a user is searching for information about a specific product in a shopping mall within the metaverse. When the user asks an agent, "Tell me about the new game console," the server provides personalized information based on the user's past behavior and preferences. In this case, the agent might respond, "The GameConsoleX is a popular new game console. It features realistic graphics and fast loading times."

[0482] Thus, the system of the present invention appropriately provides the information that users need and realizes an advanced user experience within the metaverse.

[0483] The following describes the processing flow.

[0484] Step 1:

[0485] The user opens the login screen for the metaverse platform using their device. The user enters their user ID and password.

[0486] Step 2:

[0487] The terminal sends the authentication information entered by the user to the server.

[0488] Step 3:

[0489] The server compares the received authentication information with the database, and if they match, it generates a session ID and grants the user permission to log in.

[0490] Step 4:

[0491] The server retrieves user profile information from the database and stores it as session information.

[0492] Step 5:

[0493] The server generates an autonomous agent using generative AI based on the user's profile information and past behavioral history. The information about the generated agent is then stored in a database.

[0494] Step 6:

[0495] The device displays an interface within the metaverse for the user to approach the agent.

[0496] Step 7:

[0497] The user enters "Hello, please tell me about the latest smartphones" into the agent via the interface.

[0498] Step 8:

[0499] The terminal sends the user's input text to the server.

[0500] Step 9:

[0501] The server passes the received text to a natural language processing (NLP) engine, which analyzes keywords and intent (e.g., "smartphone," "latest," "tell me," etc.).

[0502] Step 10:

[0503] Based on the analysis results, the server identifies the category of information the user is looking for (new smartphone).

[0504] Step 11:

[0505] The server uses generative AI to create suitable information by referencing the user's past behavior history and preferences.

[0506] Step 12:

[0507] The server generates natural language responses for the agent based on the generated recommendation information.

[0508] Step 13:

[0509] The server sends the agent's response to the user's terminal.

[0510] Step 14:

[0511] The terminal displays the received agent response in the user interface.

[0512] Step 15:

[0513] The user can review the agent's response and then ask further questions or take further action.

[0514] (Example 1)

[0515] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0516] Traditional metaverse platforms have suffered from insufficient information provision and a poor quality of conversational experience for users. In particular, the lack of features to personalize information based on user behavior history and individual preferences prevented users from quickly obtaining the information they needed. Furthermore, the immaturity of technologies for analyzing natural language input and generating responses that matched user intent limited the naturalness and usefulness of conversations. As a result, users were dissatisfied with their experience within the metaverse, making long-term continued use difficult.

[0517] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0518] In this invention, the server includes means for user authentication, means for generating an autonomous agent, means for analyzing user input using natural language processing, means for personalizing information based on the analysis results, means for generating agent responses based on the personalized information, means for sending the generated responses to the user terminal, means for generating an agent by referring to the user's profile information and behavioral history, means for maintaining a dialogue scenario with the agent, means for generating natural language responses for the agent using a generative AI model, and means for extracting keywords and intentions using a natural language processing engine. This makes it possible to provide personalized information based on the user's individual preferences and past behavioral history. Furthermore, by accurately analyzing user input in natural language dialogue and generating responses based on that analysis, it becomes possible to provide a more natural and useful dialogue experience.

[0519] User authentication is the process by which a user verifies that they are a legitimate user when accessing a system, using authentication information such as a user ID and password.

[0520] An "autonomous agent" is an artificial intelligence agent generated using a generative AI model, which maintains dialogue scenarios based on the user's profile information and behavioral history, and enables dialogue in natural language.

[0521] "Natural language processing" is a technology in which a server analyzes the natural language input by a user and extracts keywords and intent.

[0522] "Personalization" refers to individually optimizing information and responses based on a user's past behavior history and preferences.

[0523] A "generative AI model" is an artificial intelligence model trained on a large dataset, such as GPT-4, which performs natural language generation.

[0524] A "session ID" is an identifier generated by the server to uniquely identify a user's session while they are logged into the system.

[0525] A "natural language processing engine" is software or a library that a server uses to analyze a user's natural language input and extract keywords and intent.

[0526] A "dialogue scenario" is a set of dialogue patterns and rules that an agent maintains in order to facilitate smooth conversations with users.

[0527] "User profile information" refers to data that includes the user's basic attribute information (e.g., name, age, interests, etc.).

[0528] "Behavioral history" refers to data on past actions and preferences that a user has taken within the system.

[0529] This invention is a system in which an autonomous agent in the metaverse interacts with a user in natural language and provides information based on the user's actions and requests. This invention is realized through the following processes: user authentication, agent generation, natural language processing (NLP), information personalization, response generation, and response transmission.

[0530] This system consists of the following hardware and software.

[0531] Hardware: High-performance cloud servers (e.g., AWS EC2, Google Cloud Compute Engine), user devices (e.g., PCs, smartphones)

[0532] Software: Relational database management systems (e.g., MySQL, PostgreSQL), open-source NLP libraries (e.g., SpaCy, NLTK), generative AI models (e.g., GPT-4)

[0533] The specific form for implementing this system is as follows:

[0534] User authentication:

[0535] When a user accesses the metaverse platform, the server first displays a login screen. The user enters their user ID and password and sends them from their terminal to the server. The server verifies this authentication information against a relational database (e.g., MySQL), and if the authentication information is correct, it authenticates successfully and generates a session ID.

[0536] Agent generation:

[0537] The server retrieves the user's profile information and past behavior history from the database. Next, the server generates an autonomous agent using a generative AI model (e.g., GPT-4). This agent maintains conversational scenarios based on the user's preferences and interests, and the generated agent's information is stored in the database.

[0538] Natural Language Processing (NLP):

[0539] When a user enters a question into an agent within the metaverse, the content of that question is sent from the user's device to the server. The server analyzes the received user input using a natural language processing engine (e.g., SpaCy) to extract keywords and the user's intent.

[0540] Personalizing information:

[0541] The server retrieves data on the user's past behavior and preferences from a database, based on keywords and intents extracted by the NLP engine. The server uses this data to optimize and personalize the information for the user.

[0542] Response generation:

[0543] The server generates natural language responses using a generative AI model (e.g., GPT-4) based on personalized information. For example, it might generate a response like, "For the latest smartphone, I recommend the latest model from BrandX. It has excellent camera performance and is reasonably priced."

[0544] Sending and displaying responses:

[0545] The generated response is sent from the server to the user's terminal, which then displays it to the user. The user can then review this response and continue interacting with the agent.

[0546] Specific example:

[0547] Consider a scenario where a user is searching for information about a specific product in a shopping mall within the metaverse. When the user asks an agent, "Tell me about the new game console," the server provides personalized information based on the user's past behavior and preferences. In this case, the agent might respond, "The GameConsoleX is a popular new game console. It features realistic graphics and fast loading times."

[0548] Examples of prompts to input into a generative AI model:

[0549] 1. "Generate a response for when a user is looking for information about the latest smartphones."

[0550] 2. "Generate the dialogue in which the agent explains the new game console to the user."

[0551] This enables the system of the present invention to appropriately provide the information that users need and to realize an advanced user experience within the metaverse.

[0552] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0553] Step 1: User Authentication

[0554] The user accesses the metaverse platform using their device and enters their user ID and password on the login screen. This information is sent from the device to the server. The server verifies the received user ID and password against its database. If the authentication information is correct, the server authenticates successfully, generates a session ID, and retrieves the user's profile information. This retrieved profile information is used in the next step.

[0555] Input: User ID, Password

[0556] Output: Authentication result, session ID, profile information

[0557] Step 2: Agent Generation

[0558] The server retrieves user profile information and past behavioral history from a database. The server uses a generative AI model (e.g., GPT-4) to generate an autonomous agent based on the user's preferences and interests. This agent maintains conversational scenarios, and information about the generated agent is stored in the database.

[0559] Input: Profile information, past activity history

[0560] Output: Generated agent information

[0561] Step 3: Natural Language Processing (NLP)

[0562] The user enters a question into the agent within the metaverse. For example, a question such as "Tell me about the latest smartphones" is sent from the terminal to the server. The server analyzes the received user input using a natural language processing engine (e.g., SpaCy) to extract keywords and the user's intent.

[0563] Input: User's question

[0564] Output: Analyzed keywords, user intent

[0565] Step 4: Personalizing Information

[0566] The server retrieves past user behavior history and preference data from a database based on keywords and intentions extracted by the NLP engine. Using this data, the server generates information optimized for the user. For example, a user who has previously preferred purchasing smartphones of a particular brand will be prioritized in receiving information about the latest models of that brand.

[0567] Input: Analyzed keywords, user intent, behavioral history, preference data

[0568] Output: Personalized information

[0569] Step 5: Generating a response

[0570] The server generates natural language responses using a generative AI model (e.g., GPT-4) based on personalized information. For example, it might generate a response like, "For the latest smartphone, I recommend the latest model from BrandX. It has excellent camera performance and is reasonably priced."

[0571] Input: Personalized information

[0572] Output: Generated natural language response

[0573] Step 6: Sending and displaying the response

[0574] The server sends the generated response to the user's terminal. The terminal displays the received response to the user. The user can then review the response and continue interacting with the agent.

[0575] Input: Generated natural language response

[0576] Output: Response to the user

[0577] These steps enable the provision of highly personalized information to users within the metaverse, thereby improving the user experience.

[0578] (Application Example 1)

[0579] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0580] Modern brick-and-mortar stores need effective and timely ways to provide customers with the product information they need. However, in many cases, staff interaction and existing digital signage alone are insufficient to meet individual customer needs. This often leads to decreased customer satisfaction and missed sales opportunities. Furthermore, there is a lack of personalized information delivery systems that effectively utilize users' past purchase history and profile information. Therefore, there is a need to develop new interactive information delivery systems aimed at improving the customer experience and increasing store sales.

[0581] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0582] In this invention, the server includes means for user authentication, means for generating an autonomous agent, means for analyzing user input using natural language processing, means for personalizing information based on the analysis results, means for generating agent responses based on the personalized information, means for displaying the generated responses on a smart device, means for personalizing the agent using the user's past purchase history and profile information, means for customers to interact with the agent using a smart device in a store, and means for displaying the agent's responses on the smart device's display. This enables customers to obtain information tailored to their individual needs in real time through a smart device in a physical store, improving the customer experience and increasing store sales.

[0583] User authentication is the process of verifying the legitimacy of a user accessing a system.

[0584] An "autonomous agent" is a dialogue system that is generated based on the user's profile and past behavioral history, and automatically responds to the user's questions and requests.

[0585] "Natural language processing" is a technology that analyzes the language and text spoken by a user to understand its meaning and intent.

[0586] "Personalizing information" means tailoring the information provided to each user individually based on their past behavioral history and profile information.

[0587] "Agent response generation" is the process by which an autonomous agent creates a response to a user based on personalized information.

[0588] A "smart device" is a portable electronic device with advanced computing capabilities and connectivity, and it is a device that has an interface with the user.

[0589] "Purchase history" refers to a record of products and services that a user has purchased in the past.

[0590] "Profile information" refers to data about a user's personal information, behavioral patterns, and preferences.

[0591] "Customers interacting with agents using smart devices in-store" means that customers in a physical store use devices such as smartphones or smart glasses to communicate with agents via voice or text.

[0592] A "display" is a device used to visually display information.

[0593] This invention is a system in which customers in a physical store can use a smart device to interact with an autonomous agent in natural language and obtain personalized information.

[0594] Hardware and software to be used

[0595] Hardware: Smart glasses (e.g., Google Glass)

[0596] Software: Server-side: Node.js, Natural Language Processing Engine: Google Cloud Natural Language API, Database: MongoDB, User Interface: React Native

[0597] Generative AI Models: Generative artificial intelligence from OpenAI

[0598] Process Overview

[0599] 1. User Authentication:

[0600] The user wears smart glasses and enters their user ID and password through the login screen. The authentication information is sent to the server, which verifies the legitimacy of the login by comparing it with the database. If successful, the server issues a session ID and retrieves the user's profile information.

[0601] 2. Generation of autonomous agents:

[0602] The server matches the user's profile information with their past purchase history and uses a generative AI model to generate an autonomous agent optimized for the user. This agent maintains conversational scenarios based on the user's preferences and past behavior.

[0603] 3. Natural Language Processing:

[0604] When a user speaks to the agent through smart glasses, the voice input is sent to a server. The server uses the Google Cloud Natural Language API to analyze the user's speech and extract keywords and intent.

[0605] 4. Personalizing information:

[0606] Based on extracted keywords and intent, the server personalizes information by referencing the user's past purchase history and profile information. For example, if a user has previously purchased a specific sporting item, the server will prioritize providing information about new products from that brand.

[0607] 5. Generating the response:

[0608] Based on personalized information, the server generates natural language responses from autonomous agents. The responses are tailored to the user's preferences and may include statements such as, "For the latest running shoes, I recommend the BrandX model. It's lightweight and has excellent cushioning."

[0609] 6. Display of response:

[0610] The generated response is sent from the server to the smart glasses and displayed on the glasses' screen. The user can read this response and continue interacting with the agent.

[0611] Specific example

[0612] For example, if a customer in a physical store asks an agent through smart glasses, "Tell me about the latest cameras," the server will prioritize displaying camera information from their preferred brands based on their past purchase history. In this case, the agent might respond, "For the latest cameras, I recommend BrandX models. They offer high resolution and excellent performance even in low light."

[0613] Example of a prompt

[0614] "User profile: { 'interest': 'technology', 'preferredBrand': 'BrandX'}

[0615] User behavior history: [{ 'storeVisit': '2023-01-01', 'purchasedItems': ['BrandX Phone']}]

[0616] Generate a personalized agent response for the query 'Please tell me about the latest cameras' in Japanese.

[0617] This system allows customers to obtain information tailored to their individual needs in real time via smart devices within physical stores, which is expected to improve the customer experience and increase store sales.

[0618] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0619] Step 1:

[0620] User Authentication

[0621] Input: User ID and password

[0622] Process: The user wears smart glasses and enters their user ID and password on the login screen. This authentication information is sent from the device to the server. The server compares it with the authentication information stored in the database to verify the validity of the authentication.

[0623] Output: If authentication is successful, the server generates a session ID and retrieves the user's profile information. If authentication fails, an error message is returned.

[0624] Step 2:

[0625] Generation of autonomous agents

[0626] Input: User profile information and past purchase history

[0627] Processing: After successful authentication, the server references the user's profile information and past purchase history, and uses a generated AI model to create an autonomous agent suited to the user. The agent maintains conversational scenarios based on the user's preferences and past behavior.

[0628] Output: Data from the generated agent. This data is stored in a database on the server.

[0629] Step 3:

[0630] Natural Language Processing

[0631] Input: User voice input (e.g., "Tell me about the latest cameras.")

[0632] Processing: When a user speaks to the agent through smart glasses, the voice input is sent from the device to the server. The server analyzes the voice using the Google Cloud Natural Language API and extracts keywords and intent.

[0633] Output: Analyzed keywords and intent. This information will be used in the next step.

[0634] Step 4:

[0635] Personalization of information

[0636] Input: Analyzed keywords and intent, past purchase history, user profile information

[0637] Processing: Based on the analyzed keywords and intent, the server personalizes information by referencing past purchase history and user profile information. It prioritizes selecting content based on the user's preferences and interests.

[0638] Output: Personalized information. This information will be used in the next step.

[0639] Step 5:

[0640] Response generation

[0641] Input: Personalized information

[0642] Processing: The server uses a generative AI model to generate a natural language response based on personalized information. For example, it might create a response like, "For the latest camera, we recommend the BrandX model. It offers high resolution and excellent performance even in low light."

[0643] Output: The generated response. This response will be used in the next step.

[0644] Step 6:

[0645] Display of response

[0646] Input: Generated response

[0647] Processing: The generated response is sent from the server to the smart glasses. This response is displayed on the smart glasses' screen.

[0648] Output: The response displayed on the user's smart glasses. The user can review this response and continue interacting with the agent.

[0649] This series of processes allows users to receive individually personalized information from an autonomous agent in real time through smart glasses within a physical store.

[0650] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0651] This invention is a system in which an autonomous agent in the metaverse interacts with a user in natural language, recognizes not only the user's actions and requests but also their emotions, and provides personalized information based on that. This invention is realized through the following processes: user authentication, agent generation, natural language processing, information personalization, emotion recognition, response generation, and response transmission.

[0652] An example of this system is described in detail below.

[0653] User Authentication

[0654] When a user accesses the metaverse platform, they are first presented with a login screen. The user enters their user ID and password and submits them to the server. The server verifies this authentication information against its database, and if the credentials are correct, it grants the user access. Upon successful authentication, the server generates a session ID and retrieves the user's profile information.

[0655] Generation of autonomous agents

[0656] The server references the user's profile information and past behavior history, and generates an autonomous agent using generative AI. This agent holds conversational scenarios based on the user's preferences and interests. Information about the generated agent is stored in a database.

[0657] Natural Language Processing (NLP)

[0658] When a user asks an agent in the metaverse a question such as, "Hello, can you tell me about the latest smartphones?", that input is sent from the device to the server. The server uses a natural language processing engine to analyze the user input and extract keywords and intent.

[0659] Personalization of information

[0660] The server identifies the category of information the user is looking for (new smartphones) based on extracted keywords and intent. Furthermore, it personalizes the information by referring to the user's past behavior history and preferences. Specifically, if the user has previously preferred purchasing smartphones of a particular brand, the server will prioritize generating information about the latest models of that brand.

[0661] emotion recognition

[0662] The server passes user input data (text, voice, facial expressions) to the emotion engine, which analyzes the user's emotions. For example, if the system detects that the user is excited, it prepares a response appropriate to that emotion.

[0663] Response generation

[0664] Based on personalized information and sentiment recognition results, the server generates a natural language response for the agent. For example, it might say, "Hello! For the latest smartphone, I recommend the latest model from BrandX. It has excellent camera performance and is reasonably priced."

[0665] Sending and displaying responses

[0666] The generated response is sent from the server to the user's terminal, which then displays it to the user. The user can then review the response and continue interacting with the agent.

[0667] Specific example

[0668] For example, consider a scenario where a user is searching for information about a specific product in a shopping mall within the metaverse. When the user asks an agent, "Tell me about the new game console," the server provides personalized information based on the user's past behavior history, preferences, and the results of emotion recognition. In this case, the agent might respond, "The GameConsoleX is a popular new game console. It features realistic graphics and fast loading times. Also, your smile shows you're enjoying it, so please consider it."

[0669] Thus, the system of the present invention appropriately provides the information that the user needs, realizing an advanced user experience within the metaverse. By combining this with emotion recognition, even more natural and responsive dialogue becomes possible, providing a more user-centric interaction.

[0670] The following describes the processing flow.

[0671] Step 1:

[0672] The user uses their device to open the login screen for the metaverse platform and enters their user ID and password.

[0673] Step 2:

[0674] The terminal sends the authentication information entered by the user to the server.

[0675] Step 3:

[0676] The server compares the received authentication information with the database, and if they match, it generates a session ID and grants the user permission to log in.

[0677] Step 4:

[0678] The server retrieves user profile information from the database and stores it as session information.

[0679] Step 5:

[0680] The server generates an autonomous agent using generative AI based on the user's profile information and past behavioral history. The information about the generated agent is then stored in a database.

[0681] Step 6:

[0682] The device displays an interface within the metaverse for the user to approach the agent.

[0683] Step 7:

[0684] The user enters "Hello, please tell me about the latest smartphones" into the agent via the interface.

[0685] Step 8:

[0686] The terminal sends the user's input text to the server.

[0687] Step 9:

[0688] The server passes the received text to a natural language processing (NLP) engine, which analyzes keywords and intent (e.g., "smartphone," "latest," "tell me," etc.).

[0689] Step 10:

[0690] Based on the analysis results, the server identifies the category of information the user is looking for (new smartphone).

[0691] Step 11:

[0692] The server uses generative AI to create suitable information by referencing the user's past behavior history and preferences.

[0693] Step 12:

[0694] The server passes the user's input data to the emotion engine, which analyzes the user's emotions. For example, it determines emotions from text, tone of voice, facial expressions, and other factors.

[0695] Step 13:

[0696] The server generates natural language responses for the agent based on the results of emotion recognition and personalized information. For example, if the user is excited, it might prepare a response such as, "I can see you're having fun. This smartphone has a particularly good camera, so you can beautifully record your memories."

[0697] Step 14:

[0698] The server sends the response generated by the agent to the user's terminal.

[0699] Step 15:

[0700] The terminal displays the received agent response in the user interface.

[0701] Step 16:

[0702] The user can review the agent's response and continue with further questions or actions. For example, they might type, "That's great! What other features are there?"

[0703] (Example 2)

[0704] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0705] In modern society, there is a growing need for autonomous agents that can interact with users in natural language to improve the user experience within the metaverse. Conventional systems have the problem of not being able to adequately personalize user input and recognize emotions, and can only provide users with non-interactive and uniform information. This invention solves these problems and provides an advanced information provision system that takes into account the user's behavioral history and emotions.

[0706] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0707] In this invention, the server includes means for user authentication, means for generating an autonomous agent, means for analyzing user input using natural language processing, means for personalizing information based on the analysis results, means for recognizing emotions based on user input data, means for generating agent responses based on the personalized information and emotion recognition results, and means for transmitting the generated responses to the user terminal. This enables the provision of appropriate information and natural dialogue tailored to the user's individual preferences and emotions.

[0708] "User authentication" is the process of verifying the legitimacy of a user using their user ID and password when they access a system.

[0709] An "autonomous agent" is an artificial intelligence agent that is generated based on a user's activity history and profile information, and is capable of engaging in natural language dialogue with the user.

[0710] "Natural language processing" is a technology that enables computers to understand and analyze the language that humans use in everyday life (natural language).

[0711] "Personalizing information" means customizing the information provided based on each user's individual behavioral history and preferences.

[0712] "Recognizing emotions" means analyzing user input data (text, voice, facial expressions, etc.) to identify the user's current emotional state.

[0713] "Generating agent responses" means that an autonomous agent creates appropriate responses expressed in natural language based on emotion recognition and personalized information.

[0714] "Generative artificial intelligence" refers to advanced artificial intelligence that can generate natural language and other forms of data based on input data.

[0715] A "user terminal" refers to an electronic device such as a computer, smartphone, or tablet used by a user.

[0716] This invention relates to a system in which a user interacts with an autonomous agent in natural language on a metaverse platform, and personalized information is provided based on the user's behavioral history and emotions. The invention is carried out according to the following procedure.

[0717] User Authentication

[0718] When a user accesses the metaverse platform, the device displays a login screen. The user enters their user ID and password, which the device sends to the server. The server verifies the received authentication information against its database. If authentication is successful, the server generates a session ID and retrieves the user's profile information. This process utilizes common authentication and database management systems.

[0719] Generation of autonomous agents

[0720] The server references the user's profile information and past behavior history, and generates an autonomous agent using a generative AI (e.g., GPT-4). This agent holds conversational scenarios based on the user's preferences and interests. Information about the generated agent is stored in a database. Here, the generative AI API and various database management tools are used.

[0721] Natural Language Processing (NLP)

[0722] When a user asks an agent in the metaverse a question such as "Hello, can you tell me about the latest smartphones?", that input is sent from the device to the server. The server uses a natural language processing engine (e.g., NLTK or spaCy) to analyze the user input and extract keywords and intent. Specifically, grammatical and semantic analysis are performed.

[0723] Personalization of information

[0724] The server identifies the category of information the user is looking for (e.g., a new smartphone) based on extracted keywords and intent. Furthermore, it personalizes the information by referencing the user's past behavior history and preferences. For example, a user who has previously preferred purchasing smartphones from a specific brand will be prioritized in receiving information about the latest models from that brand. User profiling techniques and data analysis tools are utilized here.

[0725] emotion recognition

[0726] The server passes user input data (text, voice, facial expressions) to an emotion engine (for example, OpenAI's emotion recognition model) to analyze the user's emotions. For example, if the server detects that the user is excited, it prepares a response appropriate to that emotion. Specifically, voice analysis, facial expression recognition, and text analysis are used in an integrated manner.

[0727] Response generation

[0728] Based on personalized information and sentiment recognition results, the server generates natural language responses for the agent. For example, it might generate a response like, "Hello! For the latest smartphone, I recommend the latest model from BrandX. It has excellent camera performance and is reasonably priced." Generative AI and natural language generation (NLG) technologies are used here.

[0729] Sending and displaying responses

[0730] The generated response is sent from the server to the user's terminal, which then displays it to the user. The user can then review the response and continue interacting with the agent. Specific data communication methods include RESTful APIs using the HTTP / HTTPS protocol.

[0731] Examples of specific cases and prompt statements

[0732] For example, consider a scenario where a user is searching for information about a specific product in a shopping mall within the metaverse. When the user asks an agent, "Tell me about the new game console," the server provides personalized information based on the user's past behavior history, preferences, and even the results of emotion recognition. In this case, the agent might respond, "The new model is a popular choice among the latest game consoles. It features realistic graphics and fast loading times. Also, your smile shows you're enjoying it, so please consider it."

[0733] Example of a prompt:

[0734] "Hello, agent. Could you tell me about the latest smartphones?"

[0735] "Tell me about the new game console."

[0736] As described above, this system performs natural language processing on user questions, generating personalized information and emotion-based responses to achieve sophisticated and natural dialogue. By combining user behavior history with emotion recognition, it becomes possible to provide even more user-centric interactions.

[0737] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0738] Step 1:

[0739] User Authentication

[0740] The device accesses the metaverse platform.

[0741] Input: User ID and password

[0742] Output: Login screen displayed

[0743] The terminal displays a login screen to the user, who then enters their user ID and password.

[0744] Specific operation: Display ID and password input fields using an HTML form.

[0745] The terminal sends the entered authentication information to the server.

[0746] Input: User input information

[0747] Output: Sending authentication information (HTTP request)

[0748] The server compares the received authentication information with the database to verify its accuracy.

[0749] Input: Authentication information

[0750] Output: Authentication result (success or failure)

[0751] Specific operation: Execute a database query and search for matching data.

[0752] Upon successful authentication, the server generates a session ID and retrieves the user's profile information.

[0753] Input: Authentication result

[0754] Output: Session ID and user profile

[0755] Specific operation: Generate a session ID using the session management system, and retrieve user information from the database.

[0756] Step 2:

[0757] Generation of autonomous agents

[0758] The server accesses the user's profile information and past activity history.

[0759] Input: User profile information, activity history

[0760] Output: Reference result

[0761] Specific operation: Execute database queries and retrieve relevant data.

[0762] The server generates autonomous agents using generative AI (e.g., GPT-4).

[0763] Input: Reference Result

[0764] Output: Autonomous agent

[0765] Specific operation: User information is passed to the generative AI API to generate an agent.

[0766] The information about the generated agents is stored in the database.

[0767] Input: Autonomous agent

[0768] Output: Agent logs in the database

[0769] Specific action: Execute a data insertion query to save agent information.

[0770] Step 3:

[0771] Natural Language Processing (NLP)

[0772] The user asks an agent a question within the metaverse.

[0773] Example: "Hello, could you tell me about the latest smartphones?"

[0774] Input: User's question

[0775] Output: Question data

[0776] The terminal sends this input to the server.

[0777] Input: User's question

[0778] Output: HTTP request (sends query data to the server)

[0779] Specific action: Send question data in JSON format.

[0780] The server uses a natural language processing engine (such as NLTK or spaCy) to analyze user input and extract keywords and intent.

[0781] Input: User's question data

[0782] Output: Analysis results (keywords, intent, etc.)

[0783] Specific operation: Calls the text analysis module to perform grammatical and semantic analysis.

[0784] Step 4:

[0785] Personalization of information

[0786] Based on the extracted keywords and intent, the server identifies the category of information the user is seeking.

[0787] Input: Analysis results

[0788] Output: Information Category

[0789] Specific operation: Uses a keyword matching algorithm.

[0790] The server further references the user's past behavior history and preferences to personalize the information.

[0791] Input: Information category, user behavior history, preferences

[0792] Output: Personalized information

[0793] Specific action: Execute the user profiling algorithm.

[0794] As a concrete example, users who have previously purchased a smartphone from a specific brand will be given priority in receiving information about the latest models from that brand.

[0795] Step 5:

[0796] emotion recognition

[0797] The server passes the user's input data to the emotion engine, which then analyzes the user's emotions.

[0798] Input: User's text, voice, and facial expression data

[0799] Output: Emotion analysis results

[0800] Specific actions: Input data into an emotion recognition model and identify emotions.

[0801] For example, if the system detects that the user is excited, it prepares a response that corresponds to that emotion.

[0802] Input: Sentiment analysis results

[0803] Output: Emotion-based response

[0804] Specific action: Apply an emotion filtering algorithm.

[0805] Step 6:

[0806] Response generation

[0807] Based on personalized information and sentiment recognition results, the server generates natural language responses for the agent.

[0808] Input: Personalized information, sentiment analysis results

[0809] Output: Natural language response

[0810] Specific operation: Utilizes the natural language generation (NLG) function of a generative AI.

[0811] For example, it can generate a response like, "Hello! For the latest smartphone, I recommend the latest model from BrandX. It has excellent camera performance and is reasonably priced."

[0812] Step 7:

[0813] Sending and displaying responses

[0814] The server sends the generated response to the terminal.

[0815] Input: Natural language response

[0816] Output: HTTP response (response data sent to the terminal)

[0817] Specific action: Send response data in JSON format.

[0818] The terminal displays this response to the user.

[0819] Input: Response data

[0820] Output: Screen display

[0821] Specific operation: Display the response content on the screen using HTML or GUI.

[0822] The user can review this response and continue interacting with the agent.

[0823] Through each of the steps described above, this system enables sophisticated information provision and natural dialogue tailored to the individual needs and emotions of the user.

[0824] (Application Example 2)

[0825] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0826] Conventional interactive agent systems in virtual stores lacked the ability to recognize user emotions, resulting in insufficient personalized information delivery. Consequently, they failed to adequately provide the information users sought, limiting the shopping experience. This invention aims to solve these problems by providing a system that analyzes user emotions and delivers more user-centric interactions.

[0827] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for user authentication, means for generating an autonomous agent, means for analyzing user input using natural language processing, means for personalizing information based on the analysis results, means for analyzing the user's emotions, means for generating an agent response based on the personalized information and the results of the emotion analysis, and means for transmitting the generated response to the user terminal. This makes it possible to provide personalized information that takes the user's emotions into consideration.

[0828] User authentication is the process of verifying whether a user attempting to access a system has the legitimate authority.

[0829] An "autonomous agent" is a program equipped with artificial intelligence that can generate natural language responses in response to user input and actions, and continue the conversation.

[0830] "Natural language processing" is a technology that enables computers to understand, analyze, and generate human language.

[0831] "Personalizing information" means selecting and providing the most relevant information to a user based on their past behavior history and preferences.

[0832] "Analyzing user emotions" refers to the technology of identifying the emotions a user is experiencing based on their statements, facial expressions, voice, and other factors.

[0833] "Generating agent responses" is the process of providing information to users in a natural conversational format based on analyzed and personalized information.

[0834] "Generative artificial intelligence" refers to artificial intelligence technology that has the ability to generate new information and content based on diverse data.

[0835] "User terminal" refers to devices used by the user, such as computers, smartphones, tablets, and head-mounted displays.

[0836] This invention provides an interactive agent system for improving the shopping experience in virtual stores. The system includes user authentication, autonomous agent generation, natural language processing, information personalization, sentiment recognition, and response generation and transmission as processes.

[0837] First, when a user accesses the virtual store, a login screen appears on the terminal. The user enters their ID and password and sends them to the server. The server checks the database and verifies the authentication information. If authentication is successful, the server generates a session ID and retrieves the user's profile information.

[0838] Next, the server uses generative artificial intelligence to generate autonomous agents. Based on the user's profile information and past behavioral history, it generates agents with dialogue scenarios tailored to the user's preferences and interests. The agent information is stored in a database.

[0839] When a user asks an agent a question in the metaverse, they might input something like, "Tell me about the new game console." This input is sent from the terminal to the server. The server uses a natural language processing engine (e.g., OpenAI GPT-3) to analyze this input and extract keywords and intent.

[0840] Based on the analysis results, the server identifies the category of information the user is seeking. Furthermore, it personalizes the information by referring to the user's past behavior history and preferences. In this process, if the user prefers a particular brand, it will prioritize providing information about the latest products from that brand.

[0841] The server passes user input data (text, voice, facial expressions) to the emotion engine, which analyzes the user's emotions. For example, if the system detects that the user is excited, it prepares a response appropriate to that emotion.

[0842] Based on emotion recognition and personalized information, the server generates responses for the agent. For example, a response such as, "Among the latest game consoles, there is a popular one that features fast loading times and realistic graphics," might be generated.

[0843] The generated response is sent from the server to the user's terminal and displayed on the terminal. The user can then review this response and continue interacting with the agent.

[0844] Specific example

[0845] For example, if a user asks "Tell me about the new game console" in a virtual store, the system will provide appropriate information based on the user's preferences and browsing history, responding with something like, "It features realistic graphics and fast loading times."

[0846] Example of a prompt

[0847] "Analyze the following user input and extract key information and intent: I want to buy a new gaming console."

[0848] "Create a virtual shopping assistant with the following characteristics: User is interested in gaming and preferred brands are brand A and brand B."

[0849] The specific hardware used in this system includes smartphones, tablets, and head-mounted displays as user terminals, while the server requires a high-performance processor and a large-capacity database. For software, OpenAI GPT-3 is used for natural language processing, and third-party APIs are used for emotion recognition.

[0850] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0851] Step 1:

[0852] A user accesses a virtual store, and a login screen appears on their terminal. The user enters their user ID and password and sends them to the server. The server compares the entered authentication information with its database, and if the information is valid, it generates a session ID and retrieves the user's profile information.

[0853] Input: User ID, Password

[0854] Output: Authentication result, session ID, user profile information

[0855] Step 2:

[0856] The server uses generative artificial intelligence to generate autonomous agents based on the user's profile information and past behavioral history. The generated agents have conversational scenarios tailored to the user's preferences and interests, and this information is stored in a database.

[0857] Input: User profile information, past activity history

[0858] Output: Autonomous agent, dialogue scenario

[0859] Step 3:

[0860] The user enters a question within the virtual store. For example, an input such as "Tell me about the new game console" is sent from the terminal to the server. The server uses a natural language processing engine (e.g., OpenAI GPT-3) to analyze this input and extract keywords and intent.

[0861] Input: User's question (text)

[0862] Output: Keywords, intent

[0863] Step 4:

[0864] The server analyzes user input to identify the categories of information the user is seeking. Furthermore, it personalizes the information by referencing the user's past behavior and preferences. In this process, it prioritizes and organizes information that the user is likely to be interested in, such as the latest product information for a specific brand.

[0865] Input: Keywords, intent, past behavior history, preferences

[0866] Output: Personalized information

[0867] Step 5:

[0868] The server passes user input data (text, voice, facial expressions) to an emotion engine, which analyzes the user's emotions. For example, it can determine whether the user is excited based on the tone of their text or voice.

[0869] Input: User input data (text, voice, facial expressions)

[0870] Output: User's emotional state

[0871] Step 6:

[0872] The server generates agent responses based on personalized information and sentiment analysis results. This can produce responses such as, "As a modern gaming console, it features fast loading times and realistic graphics." These responses are adjusted according to the user's emotional state.

[0873] Input: Personalized information, user's emotional state

[0874] Output: Agent response (text)

[0875] Step 7:

[0876] The generated response is sent from the server to the user's terminal and displayed on the terminal. The user can then review this response and continue interacting with the agent.

[0877] Input: Agent's response (text)

[0878] Output: Display on the user terminal

[0879] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0880] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0881] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0882] [Third Embodiment]

[0883] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0884] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0885] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0886] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0887] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0888] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0889] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0890] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0891] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0892] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0893] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0894] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0895] This invention is a system in which an autonomous agent in the metaverse interacts with a user in natural language and provides information based on the user's actions and requests. This invention is realized through the processes of user authentication, agent generation, natural language processing, information personalization, response generation, and response transmission.

[0896] An example of this system is described in detail below.

[0897] User Authentication

[0898] When a user accesses the metaverse platform, they are first presented with a login screen. The user enters their user ID and password and submits them to the server. The server verifies this authentication information against its database, and if the credentials are correct, it grants the user access. Upon successful authentication, the server generates a session ID and retrieves the user's profile information.

[0899] Generation of autonomous agents

[0900] The server references the user's profile information and past behavior history, and generates an autonomous agent using generative AI. This agent holds conversational scenarios based on the user's preferences and interests. Information about the generated agent is stored in a database.

[0901] Natural Language Processing (NLP)

[0902] When a user asks an agent in the metaverse a question such as, "Hello, can you tell me about the latest smartphones?", that input is sent from the device to the server. The server uses a natural language processing engine to analyze the user input and extract keywords and intent.

[0903] Personalization of information

[0904] The server personalizes information based on extracted keywords and intent, referencing the user's past behavior history and preferences. Specifically, if a user has previously preferred purchasing smartphones of a particular brand, the server will prioritize generating information about the latest models of that brand.

[0905] Response generation

[0906] Based on personalized information, the server generates a natural language response from the agent. For example, it might say, "Hello! For the latest smartphone, I recommend the latest model from BrandX. It has excellent camera performance and is reasonably priced."

[0907] Sending and displaying responses

[0908] The generated response is sent from the server to the user's terminal, which then displays it to the user. The user can then review the response and continue interacting with the agent.

[0909] Specific example

[0910] For example, consider a scenario where a user is searching for information about a specific product in a shopping mall within the metaverse. When the user asks an agent, "Tell me about the new game console," the server provides personalized information based on the user's past behavior and preferences. In this case, the agent might respond, "The GameConsoleX is a popular new game console. It features realistic graphics and fast loading times."

[0911] Thus, the system of the present invention appropriately provides the information that users need and realizes an advanced user experience within the metaverse.

[0912] The following describes the processing flow.

[0913] Step 1:

[0914] The user opens the login screen for the metaverse platform using their device. The user enters their user ID and password.

[0915] Step 2:

[0916] The terminal sends the authentication information entered by the user to the server.

[0917] Step 3:

[0918] The server compares the received authentication information with the database, and if they match, it generates a session ID and grants the user permission to log in.

[0919] Step 4:

[0920] The server retrieves user profile information from the database and stores it as session information.

[0921] Step 5:

[0922] The server generates an autonomous agent using generative AI based on the user's profile information and past behavioral history. The information about the generated agent is then stored in a database.

[0923] Step 6:

[0924] The device displays an interface within the metaverse for the user to approach the agent.

[0925] Step 7:

[0926] The user enters "Hello, please tell me about the latest smartphones" into the agent via the interface.

[0927] Step 8:

[0928] The terminal sends the user's input text to the server.

[0929] Step 9:

[0930] The server passes the received text to a natural language processing (NLP) engine, which analyzes keywords and intent (e.g., "smartphone," "latest," "tell me," etc.).

[0931] Step 10:

[0932] Based on the analysis results, the server identifies the category of information the user is looking for (new smartphone).

[0933] Step 11:

[0934] The server uses generative AI to create suitable information by referencing the user's past behavior history and preferences.

[0935] Step 12:

[0936] The server generates natural language responses for the agent based on the generated recommendation information.

[0937] Step 13:

[0938] The server sends the agent's response to the user's terminal.

[0939] Step 14:

[0940] The terminal displays the received agent response in the user interface.

[0941] Step 15:

[0942] The user can review the agent's response and then ask further questions or take further action.

[0943] (Example 1)

[0944] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0945] Traditional metaverse platforms have suffered from insufficient information provision and a poor quality of conversational experience for users. In particular, the lack of features to personalize information based on user behavior history and individual preferences prevented users from quickly obtaining the information they needed. Furthermore, the immaturity of technologies for analyzing natural language input and generating responses that matched user intent limited the naturalness and usefulness of conversations. As a result, users were dissatisfied with their experience within the metaverse, making long-term continued use difficult.

[0946] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0947] In this invention, the server includes means for user authentication, means for generating an autonomous agent, means for analyzing user input using natural language processing, means for personalizing information based on the analysis results, means for generating agent responses based on the personalized information, means for sending the generated responses to the user terminal, means for generating an agent by referring to the user's profile information and behavioral history, means for maintaining a dialogue scenario with the agent, means for generating natural language responses for the agent using a generative AI model, and means for extracting keywords and intentions using a natural language processing engine. This makes it possible to provide personalized information based on the user's individual preferences and past behavioral history. Furthermore, by accurately analyzing user input in natural language dialogue and generating responses based on that analysis, it becomes possible to provide a more natural and useful dialogue experience.

[0948] User authentication is the process by which a user verifies that they are a legitimate user when accessing a system, using authentication information such as a user ID and password.

[0949] An "autonomous agent" is an artificial intelligence agent generated using a generative AI model, which maintains dialogue scenarios based on the user's profile information and behavioral history, and enables dialogue in natural language.

[0950] "Natural language processing" is a technology in which a server analyzes the natural language input by a user and extracts keywords and intent.

[0951] "Personalization" refers to individually optimizing information and responses based on a user's past behavior history and preferences.

[0952] A "generative AI model" is an artificial intelligence model trained on a large dataset, such as GPT-4, which performs natural language generation.

[0953] A "session ID" is an identifier generated by the server to uniquely identify a user's session while they are logged into the system.

[0954] A "natural language processing engine" is software or a library that a server uses to analyze a user's natural language input and extract keywords and intent.

[0955] A "dialogue scenario" is a set of dialogue patterns and rules that an agent maintains in order to facilitate smooth conversations with users.

[0956] "User profile information" refers to data that includes the user's basic attribute information (e.g., name, age, interests, etc.).

[0957] "Behavioral history" refers to data on past actions and preferences that a user has taken within the system.

[0958] This invention is a system in which an autonomous agent in the metaverse interacts with a user in natural language and provides information based on the user's actions and requests. This invention is realized through the following processes: user authentication, agent generation, natural language processing (NLP), information personalization, response generation, and response transmission.

[0959] This system consists of the following hardware and software.

[0960] Hardware: High-performance cloud servers (e.g., AWS EC2, Google Cloud Compute Engine), user devices (e.g., PCs, smartphones)

[0961] Software: Relational database management systems (e.g., MySQL, PostgreSQL), open-source NLP libraries (e.g., SpaCy, NLTK), generative AI models (e.g., GPT-4)

[0962] The specific form for implementing this system is as follows:

[0963] User authentication:

[0964] When a user accesses the metaverse platform, the server first displays a login screen. The user enters their user ID and password and sends them from their terminal to the server. The server verifies this authentication information against a relational database (e.g., MySQL), and if the authentication information is correct, it authenticates successfully and generates a session ID.

[0965] Agent generation:

[0966] The server retrieves the user's profile information and past behavior history from the database. Next, the server generates an autonomous agent using a generative AI model (e.g., GPT-4). This agent maintains conversational scenarios based on the user's preferences and interests, and the generated agent's information is stored in the database.

[0967] Natural Language Processing (NLP):

[0968] When a user enters a question into an agent within the metaverse, the content of that question is sent from the user's device to the server. The server analyzes the received user input using a natural language processing engine (e.g., SpaCy) to extract keywords and the user's intent.

[0969] Personalizing information:

[0970] The server retrieves data on the user's past behavior and preferences from a database, based on keywords and intents extracted by the NLP engine. The server uses this data to optimize and personalize the information for the user.

[0971] Response generation:

[0972] The server generates natural language responses using a generative AI model (e.g., GPT-4) based on personalized information. For example, it might generate a response like, "For the latest smartphone, I recommend the latest model from BrandX. It has excellent camera performance and is reasonably priced."

[0973] Sending and displaying responses:

[0974] The generated response is sent from the server to the user's terminal, which then displays it to the user. The user can then review this response and continue interacting with the agent.

[0975] Specific example:

[0976] Consider a scenario where a user is searching for information about a specific product in a shopping mall within the metaverse. When the user asks an agent, "Tell me about the new game console," the server provides personalized information based on the user's past behavior and preferences. In this case, the agent might respond, "The GameConsoleX is a popular new game console. It features realistic graphics and fast loading times."

[0977] Examples of prompts to input into a generative AI model:

[0978] 1. "Generate a response for when a user is looking for information about the latest smartphones."

[0979] 2. "Generate the dialogue in which the agent explains the new game console to the user."

[0980] This enables the system of the present invention to appropriately provide the information that users need and to realize an advanced user experience within the metaverse.

[0981] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0982] Step 1: User Authentication

[0983] The user accesses the metaverse platform using their device and enters their user ID and password on the login screen. This information is sent from the device to the server. The server verifies the received user ID and password against its database. If the authentication information is correct, the server authenticates successfully, generates a session ID, and retrieves the user's profile information. This retrieved profile information is used in the next step.

[0984] Input: User ID, Password

[0985] Output: Authentication result, session ID, profile information

[0986] Step 2: Agent Generation

[0987] The server retrieves user profile information and past behavioral history from a database. The server uses a generative AI model (e.g., GPT-4) to generate an autonomous agent based on the user's preferences and interests. This agent maintains conversational scenarios, and information about the generated agent is stored in the database.

[0988] Input: Profile information, past activity history

[0989] Output: Generated agent information

[0990] Step 3: Natural Language Processing (NLP)

[0991] The user enters a question into the agent within the metaverse. For example, a question such as "Tell me about the latest smartphones" is sent from the terminal to the server. The server analyzes the received user input using a natural language processing engine (e.g., SpaCy) to extract keywords and the user's intent.

[0992] Input: User's question

[0993] Output: Analyzed keywords, user intent

[0994] Step 4: Personalizing Information

[0995] The server retrieves past user behavior history and preference data from a database based on keywords and intentions extracted by the NLP engine. Using this data, the server generates information optimized for the user. For example, a user who has previously preferred purchasing smartphones of a particular brand will be prioritized in receiving information about the latest models of that brand.

[0996] Input: Analyzed keywords, user intent, behavioral history, preference data

[0997] Output: Personalized information

[0998] Step 5: Generating a response

[0999] The server generates natural language responses using a generative AI model (e.g., GPT-4) based on personalized information. For example, it might generate a response like, "For the latest smartphone, I recommend the latest model from BrandX. It has excellent camera performance and is reasonably priced."

[1000] Input: Personalized information

[1001] Output: Generated natural language response

[1002] Step 6: Sending and displaying the response

[1003] The server sends the generated response to the user's terminal. The terminal displays the received response to the user. The user can then review the response and continue interacting with the agent.

[1004] Input: Generated natural language response

[1005] Output: Response to the user

[1006] These steps enable the provision of highly personalized information to users within the metaverse, thereby improving the user experience.

[1007] (Application Example 1)

[1008] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1009] Modern brick-and-mortar stores need effective and timely ways to provide customers with the product information they need. However, in many cases, staff interaction and existing digital signage alone are insufficient to meet individual customer needs. This often leads to decreased customer satisfaction and missed sales opportunities. Furthermore, there is a lack of personalized information delivery systems that effectively utilize users' past purchase history and profile information. Therefore, there is a need to develop new interactive information delivery systems aimed at improving the customer experience and increasing store sales.

[1010] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1011] In this invention, the server includes means for user authentication, means for generating an autonomous agent, means for analyzing user input using natural language processing, means for personalizing information based on the analysis results, means for generating agent responses based on the personalized information, means for displaying the generated responses on a smart device, means for personalizing the agent using the user's past purchase history and profile information, means for customers to interact with the agent using a smart device in a store, and means for displaying the agent's responses on the smart device's display. This enables customers to obtain information tailored to their individual needs in real time through a smart device in a physical store, improving the customer experience and increasing store sales.

[1012] User authentication is the process of verifying the legitimacy of a user accessing a system.

[1013] An "autonomous agent" is a dialogue system that is generated based on the user's profile and past behavioral history, and automatically responds to the user's questions and requests.

[1014] "Natural language processing" is a technology that analyzes the language and text spoken by a user to understand its meaning and intent.

[1015] "Personalizing information" means tailoring the information provided to each user individually based on their past behavioral history and profile information.

[1016] "Agent response generation" is the process by which an autonomous agent creates a response to a user based on personalized information.

[1017] A "smart device" is a portable electronic device with advanced computing capabilities and connectivity, and it is a device that has an interface with the user.

[1018] "Purchase history" refers to a record of products and services that a user has purchased in the past.

[1019] "Profile information" refers to data about a user's personal information, behavioral patterns, and preferences.

[1020] "Customers interacting with agents using smart devices in-store" means that customers in a physical store use devices such as smartphones or smart glasses to communicate with agents via voice or text.

[1021] A "display" is a device used to visually display information.

[1022] This invention is a system in which customers in a physical store can use a smart device to interact with an autonomous agent in natural language and obtain personalized information.

[1023] Hardware and software to be used

[1024] Hardware: Smart glasses (e.g., Google Glass)

[1025] Software: Server-side: Node.js, Natural Language Processing Engine: Google Cloud Natural Language API, Database: MongoDB, User Interface: React Native

[1026] Generative AI Models: Generative artificial intelligence from OpenAI

[1027] Process Overview

[1028] 1. User Authentication:

[1029] The user wears smart glasses and enters their user ID and password through the login screen. The authentication information is sent to the server, which verifies the legitimacy of the login by comparing it with the database. If successful, the server issues a session ID and retrieves the user's profile information.

[1030] 2. Generation of autonomous agents:

[1031] The server matches the user's profile information with their past purchase history and uses a generative AI model to generate an autonomous agent optimized for the user. This agent maintains conversational scenarios based on the user's preferences and past behavior.

[1032] 3. Natural Language Processing:

[1033] When a user speaks to the agent through smart glasses, the voice input is sent to a server. The server uses the Google Cloud Natural Language API to analyze the user's speech and extract keywords and intent.

[1034] 4. Personalizing information:

[1035] Based on extracted keywords and intent, the server personalizes information by referencing the user's past purchase history and profile information. For example, if a user has previously purchased a specific sporting item, the server will prioritize providing information about new products from that brand.

[1036] 5. Generating the response:

[1037] Based on personalized information, the server generates natural language responses from autonomous agents. The responses are tailored to the user's preferences and may include statements such as, "For the latest running shoes, I recommend the BrandX model. It's lightweight and has excellent cushioning."

[1038] 6. Display of response:

[1039] The generated response is sent from the server to the smart glasses and displayed on the glasses' screen. The user can read this response and continue interacting with the agent.

[1040] Specific example

[1041] For example, if a customer in a physical store asks an agent through smart glasses, "Tell me about the latest cameras," the server will prioritize displaying camera information from their preferred brands based on their past purchase history. In this case, the agent might respond, "For the latest cameras, I recommend BrandX models. They offer high resolution and excellent performance even in low light."

[1042] Example of a prompt

[1043] "User profile: { 'interest': 'technology', 'preferredBrand': 'BrandX'}

[1044] User behavior history: [{ 'storeVisit': '2023-01-01', 'purchasedItems': ['BrandX Phone']}]

[1045] Generate a personalized agent response for the query 'Please tell me about the latest cameras' in Japanese.

[1046] This system allows customers to obtain information tailored to their individual needs in real time via smart devices within physical stores, which is expected to improve the customer experience and increase store sales.

[1047] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1048] Step 1:

[1049] User Authentication

[1050] Input: User ID and password

[1051] Process: The user wears smart glasses and enters their user ID and password on the login screen. This authentication information is sent from the device to the server. The server compares it with the authentication information stored in the database to verify the validity of the authentication.

[1052] Output: If authentication is successful, the server generates a session ID and retrieves the user's profile information. If authentication fails, an error message is returned.

[1053] Step 2:

[1054] Generation of autonomous agents

[1055] Input: User profile information and past purchase history

[1056] Processing: After successful authentication, the server references the user's profile information and past purchase history, and uses a generated AI model to create an autonomous agent suited to the user. The agent maintains conversational scenarios based on the user's preferences and past behavior.

[1057] Output: Data from the generated agent. This data is stored in a database on the server.

[1058] Step 3:

[1059] Natural Language Processing

[1060] Input: User voice input (e.g., "Tell me about the latest cameras.")

[1061] Processing: When a user speaks to the agent through smart glasses, the voice input is sent from the device to the server. The server analyzes the voice using the Google Cloud Natural Language API and extracts keywords and intent.

[1062] Output: Analyzed keywords and intent. This information will be used in the next step.

[1063] Step 4:

[1064] Personalization of information

[1065] Input: Analyzed keywords and intent, past purchase history, user profile information

[1066] Processing: Based on the analyzed keywords and intent, the server personalizes information by referencing past purchase history and user profile information. It prioritizes selecting content based on the user's preferences and interests.

[1067] Output: Personalized information. This information will be used in the next step.

[1068] Step 5:

[1069] Response generation

[1070] Input: Personalized information

[1071] Processing: The server uses a generative AI model to generate a natural language response based on personalized information. For example, it might create a response like, "For the latest camera, we recommend the BrandX model. It offers high resolution and excellent performance even in low light."

[1072] Output: The generated response. This response will be used in the next step.

[1073] Step 6:

[1074] Display of response

[1075] Input: Generated response

[1076] Processing: The generated response is sent from the server to the smart glasses. This response is displayed on the smart glasses' screen.

[1077] Output: The response displayed on the user's smart glasses. The user can review this response and continue interacting with the agent.

[1078] This series of processes allows users to receive individually personalized information from an autonomous agent in real time through smart glasses within a physical store.

[1079] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1080] This invention is a system in which an autonomous agent in the metaverse interacts with a user in natural language, recognizes not only the user's actions and requests but also their emotions, and provides personalized information based on that. This invention is realized through the following processes: user authentication, agent generation, natural language processing, information personalization, emotion recognition, response generation, and response transmission.

[1081] An example of this system is described in detail below.

[1082] User Authentication

[1083] When a user accesses the metaverse platform, they are first presented with a login screen. The user enters their user ID and password and submits them to the server. The server verifies this authentication information against its database, and if the credentials are correct, it grants the user access. Upon successful authentication, the server generates a session ID and retrieves the user's profile information.

[1084] Generation of autonomous agents

[1085] The server references the user's profile information and past behavior history, and generates an autonomous agent using generative AI. This agent holds conversational scenarios based on the user's preferences and interests. Information about the generated agent is stored in a database.

[1086] Natural Language Processing (NLP)

[1087] When a user asks an agent in the metaverse a question such as, "Hello, can you tell me about the latest smartphones?", that input is sent from the device to the server. The server uses a natural language processing engine to analyze the user input and extract keywords and intent.

[1088] Personalization of information

[1089] The server identifies the category of information the user is looking for (new smartphones) based on extracted keywords and intent. Furthermore, it personalizes the information by referring to the user's past behavior history and preferences. Specifically, if the user has previously preferred purchasing smartphones of a particular brand, the server will prioritize generating information about the latest models of that brand.

[1090] emotion recognition

[1091] The server passes user input data (text, voice, facial expressions) to the emotion engine, which analyzes the user's emotions. For example, if the system detects that the user is excited, it prepares a response appropriate to that emotion.

[1092] Response generation

[1093] Based on personalized information and sentiment recognition results, the server generates a natural language response for the agent. For example, it might say, "Hello! For the latest smartphone, I recommend the latest model from BrandX. It has excellent camera performance and is reasonably priced."

[1094] Sending and displaying responses

[1095] The generated response is sent from the server to the user's terminal, which then displays it to the user. The user can then review the response and continue interacting with the agent.

[1096] Specific example

[1097] For example, consider a scenario where a user is searching for information about a specific product in a shopping mall within the metaverse. When the user asks an agent, "Tell me about the new game console," the server provides personalized information based on the user's past behavior history, preferences, and the results of emotion recognition. In this case, the agent might respond, "The GameConsoleX is a popular new game console. It features realistic graphics and fast loading times. Also, your smile shows you're enjoying it, so please consider it."

[1098] Thus, the system of the present invention appropriately provides the information that the user needs, realizing an advanced user experience within the metaverse. By combining this with emotion recognition, even more natural and responsive dialogue becomes possible, providing a more user-centric interaction.

[1099] The following describes the processing flow.

[1100] Step 1:

[1101] The user uses their device to open the login screen for the metaverse platform and enters their user ID and password.

[1102] Step 2:

[1103] The terminal sends the authentication information entered by the user to the server.

[1104] Step 3:

[1105] The server compares the received authentication information with the database, and if they match, it generates a session ID and grants the user permission to log in.

[1106] Step 4:

[1107] The server retrieves user profile information from the database and stores it as session information.

[1108] Step 5:

[1109] The server generates an autonomous agent using generative AI based on the user's profile information and past behavioral history. The information about the generated agent is then stored in a database.

[1110] Step 6:

[1111] The device displays an interface within the metaverse for the user to approach the agent.

[1112] Step 7:

[1113] The user enters "Hello, please tell me about the latest smartphones" into the agent via the interface.

[1114] Step 8:

[1115] The terminal sends the user's input text to the server.

[1116] Step 9:

[1117] The server passes the received text to a natural language processing (NLP) engine, which analyzes keywords and intent (e.g., "smartphone," "latest," "tell me," etc.).

[1118] Step 10:

[1119] Based on the analysis results, the server identifies the category of information the user is looking for (new smartphone).

[1120] Step 11:

[1121] The server uses generative AI to create suitable information by referencing the user's past behavior history and preferences.

[1122] Step 12:

[1123] The server passes the user's input data to the emotion engine, which analyzes the user's emotions. For example, it determines emotions from text, tone of voice, facial expressions, and other factors.

[1124] Step 13:

[1125] The server generates natural language responses for the agent based on the results of emotion recognition and personalized information. For example, if the user is excited, it might prepare a response such as, "I can see you're having fun. This smartphone has a particularly good camera, so you can beautifully record your memories."

[1126] Step 14:

[1127] The server sends the response generated by the agent to the user's terminal.

[1128] Step 15:

[1129] The terminal displays the received agent response in the user interface.

[1130] Step 16:

[1131] The user can review the agent's response and continue with further questions or actions. For example, they might type, "That's great! What other features are there?"

[1132] (Example 2)

[1133] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1134] In modern society, there is a growing need for autonomous agents that can interact with users in natural language to improve the user experience within the metaverse. Conventional systems have the problem of not being able to adequately personalize user input and recognize emotions, and can only provide users with non-interactive and uniform information. This invention solves these problems and provides an advanced information provision system that takes into account the user's behavioral history and emotions.

[1135] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1136] In this invention, the server includes means for user authentication, means for generating an autonomous agent, means for analyzing user input using natural language processing, means for personalizing information based on the analysis results, means for recognizing emotions based on user input data, means for generating agent responses based on the personalized information and emotion recognition results, and means for transmitting the generated responses to the user terminal. This enables the provision of appropriate information and natural dialogue tailored to the user's individual preferences and emotions.

[1137] "User authentication" is the process of verifying the legitimacy of a user using their user ID and password when they access a system.

[1138] An "autonomous agent" is an artificial intelligence agent that is generated based on a user's activity history and profile information, and is capable of engaging in natural language dialogue with the user.

[1139] "Natural language processing" is a technology that enables computers to understand and analyze the language that humans use in everyday life (natural language).

[1140] "Personalizing information" means customizing the information provided based on each user's individual behavioral history and preferences.

[1141] "Recognizing emotions" means analyzing user input data (text, voice, facial expressions, etc.) to identify the user's current emotional state.

[1142] "Generating agent responses" means that an autonomous agent creates appropriate responses expressed in natural language based on emotion recognition and personalized information.

[1143] "Generative artificial intelligence" refers to advanced artificial intelligence that can generate natural language and other forms of data based on input data.

[1144] A "user terminal" refers to an electronic device such as a computer, smartphone, or tablet used by a user.

[1145] This invention relates to a system in which a user interacts with an autonomous agent in natural language on a metaverse platform, and personalized information is provided based on the user's behavioral history and emotions. The invention is carried out according to the following procedure.

[1146] User Authentication

[1147] When a user accesses the metaverse platform, the device displays a login screen. The user enters their user ID and password, which the device sends to the server. The server verifies the received authentication information against its database. If authentication is successful, the server generates a session ID and retrieves the user's profile information. This process utilizes common authentication and database management systems.

[1148] Generation of autonomous agents

[1149] The server references the user's profile information and past behavior history, and generates an autonomous agent using a generative AI (e.g., GPT-4). This agent holds conversational scenarios based on the user's preferences and interests. Information about the generated agent is stored in a database. Here, the generative AI API and various database management tools are used.

[1150] Natural Language Processing (NLP)

[1151] When a user asks an agent in the metaverse a question such as "Hello, can you tell me about the latest smartphones?", that input is sent from the device to the server. The server uses a natural language processing engine (e.g., NLTK or spaCy) to analyze the user input and extract keywords and intent. Specifically, grammatical and semantic analysis are performed.

[1152] Personalization of information

[1153] The server identifies the category of information the user is looking for (e.g., a new smartphone) based on extracted keywords and intent. Furthermore, it personalizes the information by referencing the user's past behavior history and preferences. For example, a user who has previously preferred purchasing smartphones from a specific brand will be prioritized in receiving information about the latest models from that brand. User profiling techniques and data analysis tools are utilized here.

[1154] emotion recognition

[1155] The server passes user input data (text, voice, facial expressions) to an emotion engine (for example, OpenAI's emotion recognition model) to analyze the user's emotions. For example, if the server detects that the user is excited, it prepares a response appropriate to that emotion. Specifically, voice analysis, facial expression recognition, and text analysis are used in an integrated manner.

[1156] Response generation

[1157] Based on personalized information and sentiment recognition results, the server generates natural language responses for the agent. For example, it might generate a response like, "Hello! For the latest smartphone, I recommend the latest model from BrandX. It has excellent camera performance and is reasonably priced." Generative AI and natural language generation (NLG) technologies are used here.

[1158] Sending and displaying responses

[1159] The generated response is sent from the server to the user's terminal, which then displays it to the user. The user can then review the response and continue interacting with the agent. Specific data communication methods include RESTful APIs using the HTTP / HTTPS protocol.

[1160] Examples of specific cases and prompt statements

[1161] For example, consider a scenario where a user is searching for information about a specific product in a shopping mall within the metaverse. When the user asks an agent, "Tell me about the new game console," the server provides personalized information based on the user's past behavior history, preferences, and even the results of emotion recognition. In this case, the agent might respond, "The new model is a popular choice among the latest game consoles. It features realistic graphics and fast loading times. Also, your smile shows you're enjoying it, so please consider it."

[1162] Example of a prompt:

[1163] "Hello, agent. Could you tell me about the latest smartphones?"

[1164] "Tell me about the new game console."

[1165] As described above, this system performs natural language processing on user questions, generating personalized information and emotion-based responses to achieve sophisticated and natural dialogue. By combining user behavior history with emotion recognition, it becomes possible to provide even more user-centric interactions.

[1166] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1167] Step 1:

[1168] User Authentication

[1169] The device accesses the metaverse platform.

[1170] Input: User ID and password

[1171] Output: Login screen displayed

[1172] The terminal displays a login screen to the user, who then enters their user ID and password.

[1173] Specific operation: Display ID and password input fields using an HTML form.

[1174] The terminal sends the entered authentication information to the server.

[1175] Input: User input information

[1176] Output: Sending authentication information (HTTP request)

[1177] The server compares the received authentication information with the database to verify its accuracy.

[1178] Input: Authentication information

[1179] Output: Authentication result (success or failure)

[1180] Specific operation: Execute a database query and search for matching data.

[1181] Upon successful authentication, the server generates a session ID and retrieves the user's profile information.

[1182] Input: Authentication result

[1183] Output: Session ID and user profile

[1184] Specific operation: Generate a session ID using the session management system, and retrieve user information from the database.

[1185] Step 2:

[1186] Generation of autonomous agents

[1187] The server accesses the user's profile information and past activity history.

[1188] Input: User profile information, activity history

[1189] Output: Reference result

[1190] Specific operation: Execute database queries and retrieve relevant data.

[1191] The server generates autonomous agents using generative AI (e.g., GPT-4).

[1192] Input: Reference Result

[1193] Output: Autonomous agent

[1194] Specific operation: User information is passed to the generative AI API to generate an agent.

[1195] The information about the generated agents is stored in the database.

[1196] Input: Autonomous agent

[1197] Output: Agent logs in the database

[1198] Specific action: Execute a data insertion query to save agent information.

[1199] Step 3:

[1200] Natural Language Processing (NLP)

[1201] The user asks an agent a question within the metaverse.

[1202] Example: "Hello, could you tell me about the latest smartphones?"

[1203] Input: User's question

[1204] Output: Question data

[1205] The terminal sends this input to the server.

[1206] Input: User's question

[1207] Output: HTTP request (sends query data to the server)

[1208] Specific action: Send question data in JSON format.

[1209] The server uses a natural language processing engine (such as NLTK or spaCy) to analyze user input and extract keywords and intent.

[1210] Input: User's question data

[1211] Output: Analysis results (keywords, intent, etc.)

[1212] Specific operation: Calls the text analysis module to perform grammatical and semantic analysis.

[1213] Step 4:

[1214] Personalization of information

[1215] Based on the extracted keywords and intent, the server identifies the category of information the user is seeking.

[1216] Input: Analysis results

[1217] Output: Information Category

[1218] Specific operation: Uses a keyword matching algorithm.

[1219] The server further references the user's past behavior history and preferences to personalize the information.

[1220] Input: Information category, user behavior history, preferences

[1221] Output: Personalized information

[1222] Specific action: Execute the user profiling algorithm.

[1223] As a concrete example, users who have previously purchased a smartphone from a specific brand will be given priority in receiving information about the latest models from that brand.

[1224] Step 5:

[1225] emotion recognition

[1226] The server passes the user's input data to the emotion engine, which then analyzes the user's emotions.

[1227] Input: User's text, voice, and facial expression data

[1228] Output: Emotion analysis results

[1229] Specific actions: Input data into an emotion recognition model and identify emotions.

[1230] For example, if the system detects that the user is excited, it prepares a response that corresponds to that emotion.

[1231] Input: Sentiment analysis results

[1232] Output: Emotion-based response

[1233] Specific action: Apply an emotion filtering algorithm.

[1234] Step 6:

[1235] Response generation

[1236] Based on personalized information and sentiment recognition results, the server generates natural language responses for the agent.

[1237] Input: Personalized information, sentiment analysis results

[1238] Output: Natural language response

[1239] Specific operation: Utilizes the natural language generation (NLG) function of a generative AI.

[1240] For example, it can generate a response like, "Hello! For the latest smartphone, I recommend the latest model from BrandX. It has excellent camera performance and is reasonably priced."

[1241] Step 7:

[1242] Sending and displaying responses

[1243] The server sends the generated response to the terminal.

[1244] Input: Natural language response

[1245] Output: HTTP response (response data sent to the terminal)

[1246] Specific action: Send response data in JSON format.

[1247] The terminal displays this response to the user.

[1248] Input: Response data

[1249] Output: Screen display

[1250] Specific operation: Display the response content on the screen using HTML or GUI.

[1251] The user can review this response and continue interacting with the agent.

[1252] Through each of the steps described above, this system enables sophisticated information provision and natural dialogue tailored to the individual needs and emotions of the user.

[1253] (Application Example 2)

[1254] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1255] Conventional interactive agent systems in virtual stores lacked the ability to recognize user emotions, resulting in insufficient personalized information delivery. Consequently, they failed to adequately provide the information users sought, limiting the shopping experience. This invention aims to solve these problems by providing a system that analyzes user emotions and delivers more user-centric interactions.

[1256] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for user authentication, means for generating an autonomous agent, means for analyzing user input using natural language processing, means for personalizing information based on the analysis results, means for analyzing the user's emotions, means for generating an agent response based on the personalized information and the results of the emotion analysis, and means for transmitting the generated response to the user terminal. This makes it possible to provide personalized information that takes the user's emotions into consideration.

[1257] User authentication is the process of verifying whether a user attempting to access a system has the legitimate authority.

[1258] An "autonomous agent" is a program equipped with artificial intelligence that can generate natural language responses in response to user input and actions, and continue the conversation.

[1259] "Natural language processing" is a technology that enables computers to understand, analyze, and generate human language.

[1260] "Personalizing information" means selecting and providing the most relevant information to a user based on their past behavior history and preferences.

[1261] "Analyzing user emotions" refers to the technology of identifying the emotions a user is experiencing based on their statements, facial expressions, voice, and other factors.

[1262] "Generating agent responses" is the process of providing information to users in a natural conversational format based on analyzed and personalized information.

[1263] "Generative artificial intelligence" refers to artificial intelligence technology that has the ability to generate new information and content based on diverse data.

[1264] "User terminal" refers to devices used by the user, such as computers, smartphones, tablets, and head-mounted displays.

[1265] This invention provides an interactive agent system for improving the shopping experience in virtual stores. The system includes user authentication, autonomous agent generation, natural language processing, information personalization, sentiment recognition, and response generation and transmission as processes.

[1266] First, when a user accesses the virtual store, a login screen appears on the terminal. The user enters their ID and password and sends them to the server. The server checks the database and verifies the authentication information. If authentication is successful, the server generates a session ID and retrieves the user's profile information.

[1267] Next, the server uses generative artificial intelligence to generate autonomous agents. Based on the user's profile information and past behavioral history, it generates agents with dialogue scenarios tailored to the user's preferences and interests. The agent information is stored in a database.

[1268] When a user asks an agent a question in the metaverse, they might input something like, "Tell me about the new game console." This input is sent from the terminal to the server. The server uses a natural language processing engine (e.g., OpenAI GPT-3) to analyze this input and extract keywords and intent.

[1269] Based on the analysis results, the server identifies the category of information the user is seeking. Furthermore, it personalizes the information by referring to the user's past behavior history and preferences. In this process, if the user prefers a particular brand, it will prioritize providing information about the latest products from that brand.

[1270] The server passes user input data (text, voice, facial expressions) to the emotion engine, which analyzes the user's emotions. For example, if the system detects that the user is excited, it prepares a response appropriate to that emotion.

[1271] Based on emotion recognition and personalized information, the server generates responses for the agent. For example, a response such as, "Among the latest game consoles, there is a popular one that features fast loading times and realistic graphics," might be generated.

[1272] The generated response is sent from the server to the user's terminal and displayed on the terminal. The user can then review this response and continue interacting with the agent.

[1273] Specific example

[1274] For example, if a user asks "Tell me about the new game console" in a virtual store, the system will provide appropriate information based on the user's preferences and browsing history, responding with something like, "It features realistic graphics and fast loading times."

[1275] Example of a prompt

[1276] "Analyze the following user input and extract key information and intent: I want to buy a new gaming console."

[1277] "Create a virtual shopping assistant with the following characteristics: User is interested in gaming and preferred brands are brand A and brand B."

[1278] The specific hardware used in this system includes smartphones, tablets, and head-mounted displays as user terminals, while the server requires a high-performance processor and a large-capacity database. For software, OpenAI GPT-3 is used for natural language processing, and third-party APIs are used for emotion recognition.

[1279] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1280] Step 1:

[1281] A user accesses a virtual store, and a login screen appears on their terminal. The user enters their user ID and password and sends them to the server. The server compares the entered authentication information with its database, and if the information is valid, it generates a session ID and retrieves the user's profile information.

[1282] Input: User ID, Password

[1283] Output: Authentication result, session ID, user profile information

[1284] Step 2:

[1285] The server uses generative artificial intelligence to generate autonomous agents based on the user's profile information and past behavioral history. The generated agents have conversational scenarios tailored to the user's preferences and interests, and this information is stored in a database.

[1286] Input: User profile information, past activity history

[1287] Output: Autonomous agent, dialogue scenario

[1288] Step 3:

[1289] The user enters a question within the virtual store. For example, an input such as "Tell me about the new game console" is sent from the terminal to the server. The server uses a natural language processing engine (e.g., OpenAI GPT-3) to analyze this input and extract keywords and intent.

[1290] Input: User's question (text)

[1291] Output: Keywords, intent

[1292] Step 4:

[1293] The server analyzes user input to identify the categories of information the user is seeking. Furthermore, it personalizes the information by referencing the user's past behavior and preferences. In this process, it prioritizes and organizes information that the user is likely to be interested in, such as the latest product information for a specific brand.

[1294] Input: Keywords, intent, past behavior history, preferences

[1295] Output: Personalized information

[1296] Step 5:

[1297] The server passes user input data (text, voice, facial expressions) to an emotion engine, which analyzes the user's emotions. For example, it can determine whether the user is excited based on the tone of their text or voice.

[1298] Input: User input data (text, voice, facial expressions)

[1299] Output: User's emotional state

[1300] Step 6:

[1301] The server generates agent responses based on personalized information and sentiment analysis results. This can produce responses such as, "As a modern gaming console, it features fast loading times and realistic graphics." These responses are adjusted according to the user's emotional state.

[1302] Input: Personalized information, user's emotional state

[1303] Output: Agent response (text)

[1304] Step 7:

[1305] The generated response is sent from the server to the user's terminal and displayed on the terminal. The user can then review this response and continue interacting with the agent.

[1306] Input: Agent's response (text)

[1307] Output: Display on the user terminal

[1308] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1309] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1310] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1311] [Fourth Embodiment]

[1312] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1313] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1314] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1315] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1316] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1317] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1318] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1319] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1320] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1321] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1322] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1323] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1324] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1325] This invention is a system in which an autonomous agent in the metaverse interacts with a user in natural language and provides information based on the user's actions and requests. This invention is realized through the processes of user authentication, agent generation, natural language processing, information personalization, response generation, and response transmission.

[1326] An example of this system is described in detail below.

[1327] User Authentication

[1328] When a user accesses the metaverse platform, they are first presented with a login screen. The user enters their user ID and password and submits them to the server. The server verifies this authentication information against its database, and if the credentials are correct, it grants the user access. Upon successful authentication, the server generates a session ID and retrieves the user's profile information.

[1329] Generation of autonomous agents

[1330] The server references the user's profile information and past behavior history, and generates an autonomous agent using generative AI. This agent holds conversational scenarios based on the user's preferences and interests. Information about the generated agent is stored in a database.

[1331] Natural Language Processing (NLP)

[1332] When a user asks an agent in the metaverse a question such as, "Hello, can you tell me about the latest smartphones?", that input is sent from the device to the server. The server uses a natural language processing engine to analyze the user input and extract keywords and intent.

[1333] Personalization of information

[1334] The server personalizes information based on extracted keywords and intent, referencing the user's past behavior history and preferences. Specifically, if a user has previously preferred purchasing smartphones of a particular brand, the server will prioritize generating information about the latest models of that brand.

[1335] Response generation

[1336] Based on personalized information, the server generates a natural language response from the agent. For example, it might say, "Hello! For the latest smartphone, I recommend the latest model from BrandX. It has excellent camera performance and is reasonably priced."

[1337] Sending and displaying responses

[1338] The generated response is sent from the server to the user's terminal, which then displays it to the user. The user can then review the response and continue interacting with the agent.

[1339] Specific example

[1340] For example, consider a scenario where a user is searching for information about a specific product in a shopping mall within the metaverse. When the user asks an agent, "Tell me about the new game console," the server provides personalized information based on the user's past behavior and preferences. In this case, the agent might respond, "The GameConsoleX is a popular new game console. It features realistic graphics and fast loading times."

[1341] Thus, the system of the present invention appropriately provides the information that users need and realizes an advanced user experience within the metaverse.

[1342] The following describes the processing flow.

[1343] Step 1:

[1344] The user opens the login screen for the metaverse platform using their device. The user enters their user ID and password.

[1345] Step 2:

[1346] The terminal sends the authentication information entered by the user to the server.

[1347] Step 3:

[1348] The server compares the received authentication information with the database, and if they match, it generates a session ID and grants the user permission to log in.

[1349] Step 4:

[1350] The server retrieves user profile information from the database and stores it as session information.

[1351] Step 5:

[1352] The server generates an autonomous agent using generative AI based on the user's profile information and past behavioral history. The information about the generated agent is then stored in a database.

[1353] Step 6:

[1354] The device displays an interface within the metaverse for the user to approach the agent.

[1355] Step 7:

[1356] The user enters "Hello, please tell me about the latest smartphones" into the agent via the interface.

[1357] Step 8:

[1358] The terminal sends the user's input text to the server.

[1359] Step 9:

[1360] The server passes the received text to a natural language processing (NLP) engine, which analyzes keywords and intent (e.g., "smartphone," "latest," "tell me," etc.).

[1361] Step 10:

[1362] Based on the analysis results, the server identifies the category of information the user is looking for (new smartphone).

[1363] Step 11:

[1364] The server uses generative AI to create suitable information by referencing the user's past behavior history and preferences.

[1365] Step 12:

[1366] The server generates natural language responses for the agent based on the generated recommendation information.

[1367] Step 13:

[1368] The server sends the agent's response to the user's terminal.

[1369] Step 14:

[1370] The terminal displays the received agent response in the user interface.

[1371] Step 15:

[1372] The user can review the agent's response and then ask further questions or take further action.

[1373] (Example 1)

[1374] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1375] Traditional metaverse platforms have suffered from insufficient information provision and a poor quality of conversational experience for users. In particular, the lack of features to personalize information based on user behavior history and individual preferences prevented users from quickly obtaining the information they needed. Furthermore, the immaturity of technologies for analyzing natural language input and generating responses that matched user intent limited the naturalness and usefulness of conversations. As a result, users were dissatisfied with their experience within the metaverse, making long-term continued use difficult.

[1376] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1377] In this invention, the server includes means for user authentication, means for generating an autonomous agent, means for analyzing user input using natural language processing, means for personalizing information based on the analysis results, means for generating agent responses based on the personalized information, means for sending the generated responses to the user terminal, means for generating an agent by referring to the user's profile information and behavioral history, means for maintaining a dialogue scenario with the agent, means for generating natural language responses for the agent using a generative AI model, and means for extracting keywords and intentions using a natural language processing engine. This makes it possible to provide personalized information based on the user's individual preferences and past behavioral history. Furthermore, by accurately analyzing user input in natural language dialogue and generating responses based on that analysis, it becomes possible to provide a more natural and useful dialogue experience.

[1378] User authentication is the process by which a user verifies that they are a legitimate user when accessing a system, using authentication information such as a user ID and password.

[1379] An "autonomous agent" is an artificial intelligence agent generated using a generative AI model, which maintains dialogue scenarios based on the user's profile information and behavioral history, and enables dialogue in natural language.

[1380] "Natural language processing" is a technology in which a server analyzes the natural language input by a user and extracts keywords and intent.

[1381] "Personalization" refers to individually optimizing information and responses based on a user's past behavior history and preferences.

[1382] A "generative AI model" is an artificial intelligence model trained on a large dataset, such as GPT-4, which performs natural language generation.

[1383] A "session ID" is an identifier generated by the server to uniquely identify a user's session while they are logged into the system.

[1384] A "natural language processing engine" is software or a library that a server uses to analyze a user's natural language input and extract keywords and intent.

[1385] A "dialogue scenario" is a set of dialogue patterns and rules that an agent maintains in order to facilitate smooth conversations with users.

[1386] "User profile information" refers to data that includes the user's basic attribute information (e.g., name, age, interests, etc.).

[1387] "Behavioral history" refers to data on past actions and preferences that a user has taken within the system.

[1388] This invention is a system in which an autonomous agent in the metaverse interacts with a user in natural language and provides information based on the user's actions and requests. This invention is realized through the following processes: user authentication, agent generation, natural language processing (NLP), information personalization, response generation, and response transmission.

[1389] This system consists of the following hardware and software.

[1390] Hardware: High-performance cloud servers (e.g., AWS EC2, Google Cloud Compute Engine), user devices (e.g., PCs, smartphones)

[1391] Software: Relational database management systems (e.g., MySQL, PostgreSQL), open-source NLP libraries (e.g., SpaCy, NLTK), generative AI models (e.g., GPT-4)

[1392] The specific form for implementing this system is as follows:

[1393] User authentication:

[1394] When a user accesses the metaverse platform, the server first displays a login screen. The user enters their user ID and password and sends them from their terminal to the server. The server verifies this authentication information against a relational database (e.g., MySQL), and if the authentication information is correct, it authenticates successfully and generates a session ID.

[1395] Agent generation:

[1396] The server retrieves the user's profile information and past behavior history from the database. Next, the server generates an autonomous agent using a generative AI model (e.g., GPT-4). This agent maintains conversational scenarios based on the user's preferences and interests, and the generated agent's information is stored in the database.

[1397] Natural Language Processing (NLP):

[1398] When a user enters a question into an agent within the metaverse, the content of that question is sent from the user's device to the server. The server analyzes the received user input using a natural language processing engine (e.g., SpaCy) to extract keywords and the user's intent.

[1399] Personalizing information:

[1400] The server retrieves data on the user's past behavior and preferences from a database, based on keywords and intents extracted by the NLP engine. The server uses this data to optimize and personalize the information for the user.

[1401] Response generation:

[1402] The server generates natural language responses using a generative AI model (e.g., GPT-4) based on personalized information. For example, it might generate a response like, "For the latest smartphone, I recommend the latest model from BrandX. It has excellent camera performance and is reasonably priced."

[1403] Sending and displaying responses:

[1404] The generated response is sent from the server to the user's terminal, which then displays it to the user. The user can then review this response and continue interacting with the agent.

[1405] Specific example:

[1406] Consider a scenario where a user is searching for information about a specific product in a shopping mall within the metaverse. When the user asks an agent, "Tell me about the new game console," the server provides personalized information based on the user's past behavior and preferences. In this case, the agent might respond, "The GameConsoleX is a popular new game console. It features realistic graphics and fast loading times."

[1407] Examples of prompts to input into a generative AI model:

[1408] 1. "Generate a response for when a user is looking for information about the latest smartphones."

[1409] 2. "Generate the dialogue in which the agent explains the new game console to the user."

[1410] This enables the system of the present invention to appropriately provide the information that users need and to realize an advanced user experience within the metaverse.

[1411] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1412] Step 1: User Authentication

[1413] The user accesses the metaverse platform using their device and enters their user ID and password on the login screen. This information is sent from the device to the server. The server verifies the received user ID and password against its database. If the authentication information is correct, the server authenticates successfully, generates a session ID, and retrieves the user's profile information. This retrieved profile information is used in the next step.

[1414] Input: User ID, Password

[1415] Output: Authentication result, session ID, profile information

[1416] Step 2: Agent Generation

[1417] The server retrieves user profile information and past behavioral history from a database. The server uses a generative AI model (e.g., GPT-4) to generate an autonomous agent based on the user's preferences and interests. This agent maintains conversational scenarios, and information about the generated agent is stored in the database.

[1418] Input: Profile information, past activity history

[1419] Output: Generated agent information

[1420] Step 3: Natural Language Processing (NLP)

[1421] The user enters a question into the agent within the metaverse. For example, a question such as "Tell me about the latest smartphones" is sent from the terminal to the server. The server analyzes the received user input using a natural language processing engine (e.g., SpaCy) to extract keywords and the user's intent.

[1422] Input: User's question

[1423] Output: Analyzed keywords, user intent

[1424] Step 4: Personalizing Information

[1425] The server retrieves past user behavior history and preference data from a database based on keywords and intentions extracted by the NLP engine. Using this data, the server generates information optimized for the user. For example, a user who has previously preferred purchasing smartphones of a particular brand will be prioritized in receiving information about the latest models of that brand.

[1426] Input: Analyzed keywords, user intent, behavioral history, preference data

[1427] Output: Personalized information

[1428] Step 5: Generating a response

[1429] The server generates natural language responses using a generative AI model (e.g., GPT-4) based on personalized information. For example, it might generate a response like, "For the latest smartphone, I recommend the latest model from BrandX. It has excellent camera performance and is reasonably priced."

[1430] Input: Personalized information

[1431] Output: Generated natural language response

[1432] Step 6: Sending and displaying the response

[1433] The server sends the generated response to the user's terminal. The terminal displays the received response to the user. The user can then review the response and continue interacting with the agent.

[1434] Input: Generated natural language response

[1435] Output: Response to the user

[1436] These steps enable the provision of highly personalized information to users within the metaverse, thereby improving the user experience.

[1437] (Application Example 1)

[1438] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1439] Modern brick-and-mortar stores need effective and timely ways to provide customers with the product information they need. However, in many cases, staff interaction and existing digital signage alone are insufficient to meet individual customer needs. This often leads to decreased customer satisfaction and missed sales opportunities. Furthermore, there is a lack of personalized information delivery systems that effectively utilize users' past purchase history and profile information. Therefore, there is a need to develop new interactive information delivery systems aimed at improving the customer experience and increasing store sales.

[1440] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1441] In this invention, the server includes means for user authentication, means for generating an autonomous agent, means for analyzing user input using natural language processing, means for personalizing information based on the analysis results, means for generating agent responses based on the personalized information, means for displaying the generated responses on a smart device, means for personalizing the agent using the user's past purchase history and profile information, means for customers to interact with the agent using a smart device in a store, and means for displaying the agent's responses on the smart device's display. This enables customers to obtain information tailored to their individual needs in real time through a smart device in a physical store, improving the customer experience and increasing store sales.

[1442] User authentication is the process of verifying the legitimacy of a user accessing a system.

[1443] An "autonomous agent" is a dialogue system that is generated based on the user's profile and past behavioral history, and automatically responds to the user's questions and requests.

[1444] "Natural language processing" is a technology that analyzes the language and text spoken by a user to understand its meaning and intent.

[1445] "Personalizing information" means tailoring the information provided to each user individually based on their past behavioral history and profile information.

[1446] "Agent response generation" is the process by which an autonomous agent creates a response to a user based on personalized information.

[1447] A "smart device" is a portable electronic device with advanced computing capabilities and connectivity, and it is a device that has an interface with the user.

[1448] "Purchase history" refers to a record of products and services that a user has purchased in the past.

[1449] "Profile information" refers to data about a user's personal information, behavioral patterns, and preferences.

[1450] "Customers interacting with agents using smart devices in-store" means that customers in a physical store use devices such as smartphones or smart glasses to communicate with agents via voice or text.

[1451] A "display" is a device used to visually display information.

[1452] This invention is a system in which customers in a physical store can use a smart device to interact with an autonomous agent in natural language and obtain personalized information.

[1453] Hardware and software to be used

[1454] Hardware: Smart glasses (e.g., Google Glass)

[1455] Software: Server-side: Node.js, Natural Language Processing Engine: Google Cloud Natural Language API, Database: MongoDB, User Interface: React Native

[1456] Generative AI Models: Generative artificial intelligence from OpenAI

[1457] Process Overview

[1458] 1. User Authentication:

[1459] The user wears smart glasses and enters their user ID and password through the login screen. The authentication information is sent to the server, which verifies the legitimacy of the login by comparing it with the database. If successful, the server issues a session ID and retrieves the user's profile information.

[1460] 2. Generation of autonomous agents:

[1461] The server matches the user's profile information with their past purchase history and uses a generative AI model to generate an autonomous agent optimized for the user. This agent maintains conversational scenarios based on the user's preferences and past behavior.

[1462] 3. Natural Language Processing:

[1463] When a user speaks to the agent through smart glasses, the voice input is sent to a server. The server uses the Google Cloud Natural Language API to analyze the user's speech and extract keywords and intent.

[1464] 4. Personalizing information:

[1465] Based on extracted keywords and intent, the server personalizes information by referencing the user's past purchase history and profile information. For example, if a user has previously purchased a specific sporting item, the server will prioritize providing information about new products from that brand.

[1466] 5. Generating the response:

[1467] Based on personalized information, the server generates natural language responses from autonomous agents. The responses are tailored to the user's preferences and may include statements such as, "For the latest running shoes, I recommend the BrandX model. It's lightweight and has excellent cushioning."

[1468] 6. Display of response:

[1469] The generated response is sent from the server to the smart glasses and displayed on the glasses' screen. The user can read this response and continue interacting with the agent.

[1470] Specific example

[1471] For example, if a customer in a physical store asks an agent through smart glasses, "Tell me about the latest cameras," the server will prioritize displaying camera information from their preferred brands based on their past purchase history. In this case, the agent might respond, "For the latest cameras, I recommend BrandX models. They offer high resolution and excellent performance even in low light."

[1472] Example of a prompt

[1473] "User profile: { 'interest': 'technology', 'preferredBrand': 'BrandX'}

[1474] User behavior history: [{ 'storeVisit': '2023-01-01', 'purchasedItems': ['BrandX Phone']}]

[1475] Generate a personalized agent response for the query 'Please tell me about the latest cameras' in Japanese.

[1476] This system allows customers to obtain information tailored to their individual needs in real time via smart devices within physical stores, which is expected to improve the customer experience and increase store sales.

[1477] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1478] Step 1:

[1479] User Authentication

[1480] Input: User ID and password

[1481] Process: The user wears smart glasses and enters their user ID and password on the login screen. This authentication information is sent from the device to the server. The server compares it with the authentication information stored in the database to verify the validity of the authentication.

[1482] Output: If authentication is successful, the server generates a session ID and retrieves the user's profile information. If authentication fails, an error message is returned.

[1483] Step 2:

[1484] Generation of autonomous agents

[1485] Input: User profile information and past purchase history

[1486] Processing: After successful authentication, the server references the user's profile information and past purchase history, and uses a generated AI model to create an autonomous agent suited to the user. The agent maintains conversational scenarios based on the user's preferences and past behavior.

[1487] Output: Data from the generated agent. This data is stored in a database on the server.

[1488] Step 3:

[1489] Natural Language Processing

[1490] Input: User voice input (e.g., "Tell me about the latest cameras.")

[1491] Processing: When a user speaks to the agent through smart glasses, the voice input is sent from the device to the server. The server analyzes the voice using the Google Cloud Natural Language API and extracts keywords and intent.

[1492] Output: Analyzed keywords and intent. This information will be used in the next step.

[1493] Step 4:

[1494] Personalization of information

[1495] Input: Analyzed keywords and intent, past purchase history, user profile information

[1496] Processing: Based on the analyzed keywords and intent, the server personalizes information by referencing past purchase history and user profile information. It prioritizes selecting content based on the user's preferences and interests.

[1497] Output: Personalized information. This information will be used in the next step.

[1498] Step 5:

[1499] Response generation

[1500] Input: Personalized information

[1501] Processing: The server uses a generative AI model to generate a natural language response based on personalized information. For example, it might create a response like, "For the latest camera, we recommend the BrandX model. It offers high resolution and excellent performance even in low light."

[1502] Output: The generated response. This response will be used in the next step.

[1503] Step 6:

[1504] Display of response

[1505] Input: Generated response

[1506] Processing: The generated response is sent from the server to the smart glasses. This response is displayed on the smart glasses' screen.

[1507] Output: The response displayed on the user's smart glasses. The user can review this response and continue interacting with the agent.

[1508] This series of processes allows users to receive individually personalized information from an autonomous agent in real time through smart glasses within a physical store.

[1509] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1510] This invention is a system in which an autonomous agent in the metaverse interacts with a user in natural language, recognizes not only the user's actions and requests but also their emotions, and provides personalized information based on that. This invention is realized through the following processes: user authentication, agent generation, natural language processing, information personalization, emotion recognition, response generation, and response transmission.

[1511] An example of this system is described in detail below.

[1512] User Authentication

[1513] When a user accesses the metaverse platform, they are first presented with a login screen. The user enters their user ID and password and submits them to the server. The server verifies this authentication information against its database, and if the credentials are correct, it grants the user access. Upon successful authentication, the server generates a session ID and retrieves the user's profile information.

[1514] Generation of autonomous agents

[1515] The server references the user's profile information and past behavior history, and generates an autonomous agent using generative AI. This agent holds conversational scenarios based on the user's preferences and interests. Information about the generated agent is stored in a database.

[1516] Natural Language Processing (NLP)

[1517] When a user asks an agent in the metaverse a question such as, "Hello, can you tell me about the latest smartphones?", that input is sent from the device to the server. The server uses a natural language processing engine to analyze the user input and extract keywords and intent.

[1518] Personalization of information

[1519] The server identifies the category of information the user is looking for (new smartphones) based on extracted keywords and intent. Furthermore, it personalizes the information by referring to the user's past behavior history and preferences. Specifically, if the user has previously preferred purchasing smartphones of a particular brand, the server will prioritize generating information about the latest models of that brand.

[1520] emotion recognition

[1521] The server passes user input data (text, voice, facial expressions) to the emotion engine, which analyzes the user's emotions. For example, if the system detects that the user is excited, it prepares a response appropriate to that emotion.

[1522] Response generation

[1523] Based on personalized information and sentiment recognition results, the server generates a natural language response for the agent. For example, it might say, "Hello! For the latest smartphone, I recommend the latest model from BrandX. It has excellent camera performance and is reasonably priced."

[1524] Sending and displaying responses

[1525] The generated response is sent from the server to the user's terminal, which then displays it to the user. The user can then review the response and continue interacting with the agent.

[1526] Specific example

[1527] For example, consider a scenario where a user is searching for information about a specific product in a shopping mall within the metaverse. When the user asks an agent, "Tell me about the new game console," the server provides personalized information based on the user's past behavior history, preferences, and the results of emotion recognition. In this case, the agent might respond, "The GameConsoleX is a popular new game console. It features realistic graphics and fast loading times. Also, your smile shows you're enjoying it, so please consider it."

[1528] Thus, the system of the present invention appropriately provides the information that the user needs, realizing an advanced user experience within the metaverse. By combining this with emotion recognition, even more natural and responsive dialogue becomes possible, providing a more user-centric interaction.

[1529] The following describes the processing flow.

[1530] Step 1:

[1531] The user uses their device to open the login screen for the metaverse platform and enters their user ID and password.

[1532] Step 2:

[1533] The terminal sends the authentication information entered by the user to the server.

[1534] Step 3:

[1535] The server compares the received authentication information with the database, and if they match, it generates a session ID and grants the user permission to log in.

[1536] Step 4:

[1537] The server retrieves user profile information from the database and stores it as session information.

[1538] Step 5:

[1539] The server generates an autonomous agent using generative AI based on the user's profile information and past behavioral history. The information about the generated agent is then stored in a database.

[1540] Step 6:

[1541] The device displays an interface within the metaverse for the user to approach the agent.

[1542] Step 7:

[1543] The user enters "Hello, please tell me about the latest smartphones" into the agent via the interface.

[1544] Step 8:

[1545] The terminal sends the user's input text to the server.

[1546] Step 9:

[1547] The server passes the received text to a natural language processing (NLP) engine, which analyzes keywords and intent (e.g., "smartphone," "latest," "tell me," etc.).

[1548] Step 10:

[1549] Based on the analysis results, the server identifies the category of information the user is looking for (new smartphone).

[1550] Step 11:

[1551] The server uses generative AI to create suitable information by referencing the user's past behavior history and preferences.

[1552] Step 12:

[1553] The server passes the user's input data to the emotion engine, which analyzes the user's emotions. For example, it determines emotions from text, tone of voice, facial expressions, and other factors.

[1554] Step 13:

[1555] The server generates natural language responses for the agent based on the results of emotion recognition and personalized information. For example, if the user is excited, it might prepare a response such as, "I can see you're having fun. This smartphone has a particularly good camera, so you can beautifully record your memories."

[1556] Step 14:

[1557] The server sends the response generated by the agent to the user's terminal.

[1558] Step 15:

[1559] The terminal displays the received agent response in the user interface.

[1560] Step 16:

[1561] The user can review the agent's response and continue with further questions or actions. For example, they might type, "That's great! What other features are there?"

[1562] (Example 2)

[1563] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1564] In modern society, there is a growing need for autonomous agents that can interact with users in natural language to improve the user experience within the metaverse. Conventional systems have the problem of not being able to adequately personalize user input and recognize emotions, and can only provide users with non-interactive and uniform information. This invention solves these problems and provides an advanced information provision system that takes into account the user's behavioral history and emotions.

[1565] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1566] In this invention, the server includes means for user authentication, means for generating an autonomous agent, means for analyzing user input using natural language processing, means for personalizing information based on the analysis results, means for recognizing emotions based on user input data, means for generating agent responses based on the personalized information and emotion recognition results, and means for transmitting the generated responses to the user terminal. This enables the provision of appropriate information and natural dialogue tailored to the user's individual preferences and emotions.

[1567] "User authentication" is the process of verifying the legitimacy of a user using their user ID and password when they access a system.

[1568] An "autonomous agent" is an artificial intelligence agent that is generated based on a user's activity history and profile information, and is capable of engaging in natural language dialogue with the user.

[1569] "Natural language processing" is a technology that enables computers to understand and analyze the language that humans use in everyday life (natural language).

[1570] "Personalizing information" means customizing the information provided based on each user's individual behavioral history and preferences.

[1571] "Recognizing emotions" means analyzing user input data (text, voice, facial expressions, etc.) to identify the user's current emotional state.

[1572] "Generating agent responses" means that an autonomous agent creates appropriate responses expressed in natural language based on emotion recognition and personalized information.

[1573] "Generative artificial intelligence" refers to advanced artificial intelligence that can generate natural language and other forms of data based on input data.

[1574] A "user terminal" refers to an electronic device such as a computer, smartphone, or tablet used by a user.

[1575] This invention relates to a system in which a user interacts with an autonomous agent in natural language on a metaverse platform, and personalized information is provided based on the user's behavioral history and emotions. The invention is carried out according to the following procedure.

[1576] User Authentication

[1577] When a user accesses the metaverse platform, the device displays a login screen. The user enters their user ID and password, which the device sends to the server. The server verifies the received authentication information against its database. If authentication is successful, the server generates a session ID and retrieves the user's profile information. This process utilizes common authentication and database management systems.

[1578] Generation of autonomous agents

[1579] The server references the user's profile information and past behavior history, and generates an autonomous agent using a generative AI (e.g., GPT-4). This agent holds conversational scenarios based on the user's preferences and interests. Information about the generated agent is stored in a database. Here, the generative AI API and various database management tools are used.

[1580] Natural Language Processing (NLP)

[1581] When a user asks an agent in the metaverse a question such as "Hello, can you tell me about the latest smartphones?", that input is sent from the device to the server. The server uses a natural language processing engine (e.g., NLTK or spaCy) to analyze the user input and extract keywords and intent. Specifically, grammatical and semantic analysis are performed.

[1582] Personalization of information

[1583] The server identifies the category of information the user is looking for (e.g., a new smartphone) based on extracted keywords and intent. Furthermore, it personalizes the information by referencing the user's past behavior history and preferences. For example, a user who has previously preferred purchasing smartphones from a specific brand will be prioritized in receiving information about the latest models from that brand. User profiling techniques and data analysis tools are utilized here.

[1584] emotion recognition

[1585] The server passes user input data (text, voice, facial expressions) to an emotion engine (for example, OpenAI's emotion recognition model) to analyze the user's emotions. For example, if the server detects that the user is excited, it prepares a response appropriate to that emotion. Specifically, voice analysis, facial expression recognition, and text analysis are used in an integrated manner.

[1586] Response generation

[1587] Based on personalized information and sentiment recognition results, the server generates natural language responses for the agent. For example, it might generate a response like, "Hello! For the latest smartphone, I recommend the latest model from BrandX. It has excellent camera performance and is reasonably priced." Generative AI and natural language generation (NLG) technologies are used here.

[1588] Sending and displaying responses

[1589] The generated response is sent from the server to the user's terminal, which then displays it to the user. The user can then review the response and continue interacting with the agent. Specific data communication methods include RESTful APIs using the HTTP / HTTPS protocol.

[1590] Examples of specific cases and prompt statements

[1591] For example, consider a scenario where a user is searching for information about a specific product in a shopping mall within the metaverse. When the user asks an agent, "Tell me about the new game console," the server provides personalized information based on the user's past behavior history, preferences, and even the results of emotion recognition. In this case, the agent might respond, "The new model is a popular choice among the latest game consoles. It features realistic graphics and fast loading times. Also, your smile shows you're enjoying it, so please consider it."

[1592] Example of a prompt:

[1593] "Hello, agent. Could you tell me about the latest smartphones?"

[1594] "Tell me about the new game console."

[1595] As described above, this system performs natural language processing on user questions, generating personalized information and emotion-based responses to achieve sophisticated and natural dialogue. By combining user behavior history with emotion recognition, it becomes possible to provide even more user-centric interactions.

[1596] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1597] Step 1:

[1598] User Authentication

[1599] The device accesses the metaverse platform.

[1600] Input: User ID and password

[1601] Output: Login screen displayed

[1602] The terminal displays a login screen to the user, who then enters their user ID and password.

[1603] Specific operation: Display ID and password input fields using an HTML form.

[1604] The terminal sends the entered authentication information to the server.

[1605] Input: User input information

[1606] Output: Sending authentication information (HTTP request)

[1607] The server compares the received authentication information with the database to verify its accuracy.

[1608] Input: Authentication information

[1609] Output: Authentication result (success or failure)

[1610] Specific operation: Execute a database query and search for matching data.

[1611] Upon successful authentication, the server generates a session ID and retrieves the user's profile information.

[1612] Input: Authentication result

[1613] Output: Session ID and user profile

[1614] Specific operation: Generate a session ID using the session management system, and retrieve user information from the database.

[1615] Step 2:

[1616] Generation of autonomous agents

[1617] The server accesses the user's profile information and past activity history.

[1618] Input: User profile information, activity history

[1619] Output: Reference result

[1620] Specific operation: Execute database queries and retrieve relevant data.

[1621] The server generates autonomous agents using generative AI (e.g., GPT-4).

[1622] Input: Reference Result

[1623] Output: Autonomous agent

[1624] Specific operation: User information is passed to the generative AI API to generate an agent.

[1625] The information about the generated agents is stored in the database.

[1626] Input: Autonomous agent

[1627] Output: Agent logs in the database

[1628] Specific action: Execute a data insertion query to save agent information.

[1629] Step 3:

[1630] Natural Language Processing (NLP)

[1631] The user asks an agent a question within the metaverse.

[1632] Example: "Hello, could you tell me about the latest smartphones?"

[1633] Input: User's question

[1634] Output: Question data

[1635] The terminal sends this input to the server.

[1636] Input: User's question

[1637] Output: HTTP request (sends query data to the server)

[1638] Specific action: Send question data in JSON format.

[1639] The server uses a natural language processing engine (such as NLTK or spaCy) to analyze user input and extract keywords and intent.

[1640] Input: User's question data

[1641] Output: Analysis results (keywords, intent, etc.)

[1642] Specific operation: Calls the text analysis module to perform grammatical and semantic analysis.

[1643] Step 4:

[1644] Personalization of information

[1645] Based on the extracted keywords and intent, the server identifies the category of information the user is seeking.

[1646] Input: Analysis results

[1647] Output: Information Category

[1648] Specific operation: Uses a keyword matching algorithm.

[1649] The server further references the user's past behavior history and preferences to personalize the information.

[1650] Input: Information category, user behavior history, preferences

[1651] Output: Personalized information

[1652] Specific action: Execute the user profiling algorithm.

[1653] As a concrete example, users who have previously purchased a smartphone from a specific brand will be given priority in receiving information about the latest models from that brand.

[1654] Step 5:

[1655] emotion recognition

[1656] The server passes the user's input data to the emotion engine, which then analyzes the user's emotions.

[1657] Input: User's text, voice, and facial expression data

[1658] Output: Emotion analysis results

[1659] Specific actions: Input data into an emotion recognition model and identify emotions.

[1660] For example, if the system detects that the user is excited, it prepares a response that corresponds to that emotion.

[1661] Input: Sentiment analysis results

[1662] Output: Emotion-based response

[1663] Specific action: Apply an emotion filtering algorithm.

[1664] Step 6:

[1665] Response generation

[1666] Based on personalized information and sentiment recognition results, the server generates natural language responses for the agent.

[1667] Input: Personalized information, sentiment analysis results

[1668] Output: Natural language response

[1669] Specific operation: Utilizes the natural language generation (NLG) function of a generative AI.

[1670] For example, it can generate a response like, "Hello! For the latest smartphone, I recommend the latest model from BrandX. It has excellent camera performance and is reasonably priced."

[1671] Step 7:

[1672] Sending and displaying responses

[1673] The server sends the generated response to the terminal.

[1674] Input: Natural language response

[1675] Output: HTTP response (response data sent to the terminal)

[1676] Specific action: Send response data in JSON format.

[1677] The terminal displays this response to the user.

[1678] Input: Response data

[1679] Output: Screen display

[1680] Specific operation: Display the response content on the screen using HTML or GUI.

[1681] The user can review this response and continue interacting with the agent.

[1682] Through each of the steps described above, this system enables sophisticated information provision and natural dialogue tailored to the individual needs and emotions of the user.

[1683] (Application Example 2)

[1684] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1685] Conventional interactive agent systems in virtual stores lacked the ability to recognize user emotions, resulting in insufficient personalized information delivery. Consequently, they failed to adequately provide the information users sought, limiting the shopping experience. This invention aims to solve these problems by providing a system that analyzes user emotions and delivers more user-centric interactions.

[1686] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for user authentication, means for generating an autonomous agent, means for analyzing user input using natural language processing, means for personalizing information based on the analysis results, means for analyzing the user's emotions, means for generating an agent response based on the personalized information and the results of the emotion analysis, and means for transmitting the generated response to the user terminal. This makes it possible to provide personalized information that takes the user's emotions into consideration.

[1687] User authentication is the process of verifying whether a user attempting to access a system has the legitimate authority.

[1688] An "autonomous agent" is a program equipped with artificial intelligence that can generate natural language responses in response to user input and actions, and continue the conversation.

[1689] "Natural language processing" is a technology that enables computers to understand, analyze, and generate human language.

[1690] "Personalizing information" means selecting and providing the most relevant information to a user based on their past behavior history and preferences.

[1691] "Analyzing user emotions" refers to the technology of identifying the emotions a user is experiencing based on their statements, facial expressions, voice, and other factors.

[1692] "Generating agent responses" is the process of providing information to users in a natural conversational format based on analyzed and personalized information.

[1693] "Generative artificial intelligence" refers to artificial intelligence technology that has the ability to generate new information and content based on diverse data.

[1694] "User terminal" refers to devices used by the user, such as computers, smartphones, tablets, and head-mounted displays.

[1695] This invention provides an interactive agent system for improving the shopping experience in virtual stores. The system includes user authentication, autonomous agent generation, natural language processing, information personalization, sentiment recognition, and response generation and transmission as processes.

[1696] First, when a user accesses the virtual store, a login screen appears on the terminal. The user enters their ID and password and sends them to the server. The server checks the database and verifies the authentication information. If authentication is successful, the server generates a session ID and retrieves the user's profile information.

[1697] Next, the server uses generative artificial intelligence to generate autonomous agents. Based on the user's profile information and past behavioral history, it generates agents with dialogue scenarios tailored to the user's preferences and interests. The agent information is stored in a database.

[1698] When a user asks an agent a question in the metaverse, they might input something like, "Tell me about the new game console." This input is sent from the terminal to the server. The server uses a natural language processing engine (e.g., OpenAI GPT-3) to analyze this input and extract keywords and intent.

[1699] Based on the analysis results, the server identifies the category of information the user is seeking. Furthermore, it personalizes the information by referring to the user's past behavior history and preferences. In this process, if the user prefers a particular brand, it will prioritize providing information about the latest products from that brand.

[1700] The server passes user input data (text, voice, facial expressions) to the emotion engine, which analyzes the user's emotions. For example, if the system detects that the user is excited, it prepares a response appropriate to that emotion.

[1701] Based on emotion recognition and personalized information, the server generates responses for the agent. For example, a response such as, "Among the latest game consoles, there is a popular one that features fast loading times and realistic graphics," might be generated.

[1702] The generated response is sent from the server to the user's terminal and displayed on the terminal. The user can then review this response and continue interacting with the agent.

[1703] Specific example

[1704] For example, if a user asks "Tell me about the new game console" in a virtual store, the system will provide appropriate information based on the user's preferences and browsing history, responding with something like, "It features realistic graphics and fast loading times."

[1705] Example of a prompt

[1706] "Analyze the following user input and extract key information and intent: I want to buy a new gaming console."

[1707] "Create a virtual shopping assistant with the following characteristics: User is interested in gaming and preferred brands are brand A and brand B."

[1708] The specific hardware used in this system includes smartphones, tablets, and head-mounted displays as user terminals, while the server requires a high-performance processor and a large-capacity database. For software, OpenAI GPT-3 is used for natural language processing, and third-party APIs are used for emotion recognition.

[1709] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1710] Step 1:

[1711] A user accesses a virtual store, and a login screen appears on their terminal. The user enters their user ID and password and sends them to the server. The server compares the entered authentication information with its database, and if the information is valid, it generates a session ID and retrieves the user's profile information.

[1712] Input: User ID, Password

[1713] Output: Authentication result, session ID, user profile information

[1714] Step 2:

[1715] The server uses generative artificial intelligence to generate autonomous agents based on the user's profile information and past behavioral history. The generated agents have conversational scenarios tailored to the user's preferences and interests, and this information is stored in a database.

[1716] Input: User profile information, past activity history

[1717] Output: Autonomous agent, dialogue scenario

[1718] Step 3:

[1719] The user enters a question within the virtual store. For example, an input such as "Tell me about the new game console" is sent from the terminal to the server. The server uses a natural language processing engine (e.g., OpenAI GPT-3) to analyze this input and extract keywords and intent.

[1720] Input: User's question (text)

[1721] Output: Keywords, intent

[1722] Step 4:

[1723] The server analyzes user input to identify the categories of information the user is seeking. Furthermore, it personalizes the information by referencing the user's past behavior and preferences. In this process, it prioritizes and organizes information that the user is likely to be interested in, such as the latest product information for a specific brand.

[1724] Input: Keywords, intent, past behavior history, preferences

[1725] Output: Personalized information

[1726] Step 5:

[1727] The server passes user input data (text, voice, facial expressions) to an emotion engine, which analyzes the user's emotions. For example, it can determine whether the user is excited based on the tone of their text or voice.

[1728] Input: User input data (text, voice, facial expressions)

[1729] Output: User's emotional state

[1730] Step 6:

[1731] The server generates agent responses based on personalized information and sentiment analysis results. This can produce responses such as, "As a modern gaming console, it features fast loading times and realistic graphics." These responses are adjusted according to the user's emotional state.

[1732] Input: Personalized information, user's emotional state

[1733] Output: Agent response (text)

[1734] Step 7:

[1735] The generated response is sent from the server to the user's terminal and displayed on the terminal. The user can then review this response and continue interacting with the agent.

[1736] Input: Agent's response (text)

[1737] Output: Display on the user terminal

[1738] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1739] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1740] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1741] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1742] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1743] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1744] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1745] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1746] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1747] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1748] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1749] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1750] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1751] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1752] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1753] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1754] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1755] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1756] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1757] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1758] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[1759] The following is further disclosed regarding the embodiments described above.

[1760] (Claim 1)

[1761] Means of performing user authentication,

[1762] A means for generating an autonomous agent,

[1763] A means of analyzing user input using natural language processing,

[1764] A means of personalizing information based on analysis results,

[1765] A means of generating an agent's response based on personalized information,

[1766] A system including means for sending the generated response to a user terminal.

[1767] (Claim 2)

[1768] The system according to claim 1, further comprising means for recording the user's behavior history and using it for agent generation and response personalization.

[1769] (Claim 3)

[1770] The system according to claim 1, comprising means for generating an autonomous agent using generative artificial intelligence.

[1771] "Example 1"

[1772] (Claim 1)

[1773] Means of performing user authentication,

[1774] A means for generating an autonomous agent,

[1775] A means of analyzing user input using natural language processing,

[1776] A means of personalizing information based on analysis results,

[1777] A means of generating an agent's response based on personalized information,

[1778] A means for sending the generated response to the user terminal,

[1779] A means of generating agents by referring to user profile information and behavioral history,

[1780] A means of maintaining the dialogue scenario with the agent,

[1781] A means for generating natural language responses from an agent using a generative AI model,

[1782] A system that includes means for extracting keywords and intent using a natural language processing engine.

[1783] (Claim 2)

[1784] The system according to claim 1, further comprising means for recording the user's behavior history and using it for agent generation and response personalization.

[1785] (Claim 3)

[1786] The system according to claim 1, comprising means for generating an autonomous agent using generative artificial intelligence.

[1787] "Application Example 1"

[1788] (Claim 1)

[1789] Means of performing user authentication,

[1790] A means for generating an autonomous agent,

[1791] A means of analyzing user input using natural language processing,

[1792] A means of personalizing information based on analysis results,

[1793] A means of generating an agent's response based on personalized information,

[1794] A means for displaying the response generated on a smart device,

[1795] A means of personalizing agents using the user's past purchase history and profile information,

[1796] A means for customers to interact with agents using smart devices within a store,

[1797] A means of displaying the agent's response on the smart device's display,

[1798] A system that includes this.

[1799] (Claim 2)

[1800] The system according to claim 1, further comprising means for recording the user's behavior history and using it for agent generation and response personalization.

[1801] (Claim 3)

[1802] The system according to claim 1, comprising means for generating an autonomous agent using generative artificial intelligence.

[1803] "Example 2 of combining an emotion engine"

[1804] (Claim 1)

[1805] Means of performing user authentication,

[1806] A means for generating an autonomous agent,

[1807] A means of analyzing user input using natural language processing,

[1808] A means of personalizing information based on analysis results,

[1809] A means of recognizing emotions based on user input data,

[1810] A means for generating an agent's response based on personalized information and the results of emotion recognition,

[1811] A system including means for sending the generated response to a user terminal.

[1812] (Claim 2)

[1813] The system according to claim 1, further comprising means for recording the user's behavior history and using it for agent generation and response personalization.

[1814] (Claim 3)

[1815] The system according to claim 1, comprising means for generating an autonomous agent using generative artificial intelligence.

[1816] "Application example 2 when combining with an emotional engine"

[1817] (Claim 1)

[1818] Means of performing user authentication,

[1819] A means for generating an autonomous agent,

[1820] A means of analyzing user input using natural language processing,

[1821] A means of personalizing information based on analysis results,

[1822] A means of analyzing user emotions,

[1823] A means for generating an agent's response based on personalized information and the results of sentiment analysis,

[1824] A system including means for sending the generated response to a user terminal.

[1825] (Claim 2)

[1826] The system according to claim 1, further comprising means for recording the user's behavior history and using it for agent generation and response personalization.

[1827] (Claim 3)

[1828] The system according to claim 1, comprising means for generating an autonomous agent using generative artificial intelligence. [Explanation of Symbols]

[1829] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Means of performing user authentication, A means for generating an autonomous agent, A means of analyzing user input using natural language processing, A means of personalizing information based on analysis results, A means of generating an agent's response based on personalized information, A system including means for sending the generated response to a user terminal.

2. The system according to claim 1, further comprising means for recording the user's behavior history and using it for agent generation and response personalization.

3. The system according to claim 1, comprising means for generating an autonomous agent using generative artificial intelligence.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A