System

The virtual environment system addresses the limitations of traditional social media for children by enabling real-time chat, avatar movement, and personalized educational content through generation AI, thereby enhancing user interaction and learning.

JP2025071073APending Publication Date: 2025-05-02SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024184691
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-20
Filing Date
2024-10-21
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

Traditional social media applications for children lack interactive and educational content, failing to maximize user engagement and learning effects due to uniform game and activity offerings that do not consider individual user attributes and learning information.

Method used

A virtual environment system that enables real-time text or voice chat, avatar movement, and emotional expression, combined with a generation AI that provides personalized educational games and activities tailored to user interests and learning needs.

Benefits of technology

The system enhances user interaction and learning by providing rich, personalized content that maximizes engagement and educational effectiveness within a virtual environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025071073000001_ABST
    Figure 2025071073000001_ABST
Patent Text Reader

Abstract

To provide a system.SOLUTION: A system provides an information sharing application service that enables interaction between a first user, and a second user or a virtual character, within a virtual environment. The system receives a text chat or voice chat input from the first user and presents the received input to the second user, thereby achieves interaction via text chat or voice chat, and also gives movement to the first user's avatar in response to the first user's operation to cause it to express at least one of a greeting and an emotion, inputs information about the first user to generative AI, and provides at least one of educational games, activities, and entertainment to the first user based on a response from the generative AI.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including a description and related instruction sentence regarding the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2022-180282 A Summary of the Invention [Problem to be solved by the invention]

[0004] In conventional SNS applications, children often interact with friends and virtual characters in a restricted environment, and educational content was limited. In addition, uniform games and activities were provided without considering the attributes and learning information of individual users, which meant that users' interests and learning effects could not be maximized. [Means for solving the problem]

[0005] The present invention provides a means to solve the above problems by promoting children's interaction in a virtual environment and providing educational games and activities. Specifically, the following means are adopted:

[0006] 1. Means for enabling users to communicate through text and voice chat: Users can communicate in real time with friends and virtual characters within the virtual environment through text and voice chat.

[0007] 2. A way to give the user's avatar movement to greet or express emotions: By giving movement to their avatars, users can communicate more realistically with other users and AI characters, enabling richer interactions through greetings and the expression of emotions.

[0008] 3. How generative AI can be used to provide educational games and activities: By using generative AI, we provide educational games and activities that take into account the user's attributes and learning information. By providing content that matches the user's interests and learning effectiveness, we provide more effective learning and fun.

[0009] Through these measures, children will be able to enjoy content that incorporates educational elements while engaging in richer interactions.

[0010] "In-virtual environment" refers to a computer-generated imaginary space or environment in which users can interact with friends and virtual characters.

[0011] "Educational games and activities" refers to games and activities that incorporate a learning or educational element. These content are intended to provide users with educational information or skills.

[0012] "Generative AI" is a type of generative AI that refers to a model trained to process natural language. Generative AI can generate responses to user input and provide individually customized content by taking into account the user's attributes and learning information. [Brief description of the drawings]

[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Diagram 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. FIG. [Diagram 3] FIG. 11 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Diagram 5] FIG. 13 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 13 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 13 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] 4 is a sequence diagram showing a process flow of the data processing system according to the first embodiment. FIG. [Figure 12] 11 is a sequence diagram showing a process flow of the data processing system in application example 1. FIG. [Figure 13] FIG. 11 is a sequence diagram showing the flow of processing of the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 11 is a sequence diagram showing the flow of processing in the data processing system in application example 2 when combined with an emotion engine. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0014] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a signed processor (hereinafter simply referred to as a "processor") may be one arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be one type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), etc.

[0017] In the following embodiments, a signed RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by the processor.

[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0019] In the following embodiments, a communication I / F (Interface) with a code is an interface including a communication processor and an antenna. The communication I / F controls communication between multiple computers. An example of a communication standard applied to the communication I / F is a wireless communication standard including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. In addition, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed by connecting them with "and / or." In addition, the same concept as "A and / or B" is also applied to the expression "A and / or B."

[0021] [First embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0022] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0023] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).

[0024] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0025] The reception device 38 includes a touch panel 38A and a microphone 38B, and receives user input. The touch panel 38A detects contact with a pointer (e.g., a pen or a finger) to receive user input by the touch of the pointer. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (e.g., voice and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs voice according to instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a Complementary Metal-Oxide-Semiconductor (CMOS) image sensor or a Charge Coupled Device (CCD) image sensor.

[0027] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54.

[0028] FIG. 2 shows an example of main functions of the data processing device 12 and the smart device 14.

[0029] As shown in Fig. 2, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32. The specific process program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific process program 56 from the storage 32, and executes the read specific process program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific process program 56 executed on the RAM 30.

[0030] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores a reception output program 60. The reception output program 60 is used together with the specific processing program 56 by the data processing system 10. The processor 46 reads out the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0032] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0033] An embodiment for implementing the present invention includes the following elements.

[0034] 1. Server: The server receives the user's login information and performs authentication. It also manages the user's avatar information and data required for game activities, and provides appropriate content in response to the user's requests. It also processes input to the generative AI and information shared with other users.

[0035] 2. Terminal: The terminal allows the user to log in to an information sharing application such as a social networking application service (hereinafter simply referred to as the "application") and perform operations and communication within the virtual environment. The terminal transmits the user's text chat and voice chat input to the server, receives and displays the generative AI's responses, and transmits the user's avatar movements and game activity selections to the server.

[0036] 3. User: Users log into the application and interact with friends and virtual characters in a virtual environment, communicating through text and voice chat, animating their avatars to greet others, express emotions, and select educational games and activities to share with other users.

[0037] 4. Generative AI: A generative AI is a generative AI that generates responses based on user input. It receives information from the user's text chat and avatar movements, as well as information about the choices and suggestions of other users and AI characters, and generates appropriate responses. Generative AI runs on the server and processes input and output from the server.

[0038] All of these elements work together to allow users to interact with friends and virtual characters in a virtual environment and enjoy educational games and activities. The server manages user information and provides appropriate content. The device accepts user operations, communicates with the server, and displays the generative AI's responses. The generative AI generates responses based on the user's input and sends them to the device via the server.

[0039] The process flow will be explained below. Step 1: User logs into the application Step 2: The device sends the user's login information to the server Step 3: The server receives the user's login information and performs authentication. Step 4: The server sends the authentication result to the terminal. Step 5: The user selects an avatar to enter the virtual environment. Step 6: The device sends the user's avatar information to the server. Step 7: The server receives the user's avatar information and displays the avatar in the virtual environment. Step 8: The user interacts with friends and virtual characters through text and voice chat. Step 9: The device sends the user's chat input to the server Step 10: The server receives the user's chat input and passes it to the generative AI. Step 11: Generative AI generates a response based on the user's chat input Step 12: The server receives the response from the generative AI and sends it to the device. Step 13: The device displays the generative AI's response Step 14: The user gives the avatar movement to greet or express emotion. Step 15: The device sends the user's avatar movements to the server. Step 16: The server receives the user's avatar movements and moves the avatar in the virtual environment. Step 17: User selects educational games and activities Step 18: The device sends the user's selection to the server Step 19: The server serves games and activities based on the user's selection. Step 20: User experiences the game or activity Step 21: The user plays a game or activity with other users or AI characters. Step 22: The device sends the choices and suggestions of other users and AI characters to the server. Step 23: The server uses the choices and suggestions of other users and AI characters as inputs to the generative AI. Step 24: The generative AI generates a response based on the choices and suggestions of other users and the AI ​​character. Step 25: The server receives the response from the generative AI and sends it to the terminal. Step 26: The device displays the generative AI's response. Step 27: The user shares the results of the game or activity with other users. Step 28: The device sends the user's sharing contacts and chat contents to the server. Step 29: The server uses the user's shared contacts and chat contents as input to the generative AI. Step 30: Generative AI generates responses based on who the user is sharing with and what the chat is about. Step 31: The server receives the response from the generative AI and sends it to the terminal. Step 32: The device displays the generative AI's response. Step 33: Use information from other users who are registered as friends by the user Step 34: User sends invitation notification and points to friends Step 35: The device sends the user's friend registration information to the server. Step 36: The server receives the user's friend registration information and uses it as input for the generative AI. Step 37: The generative AI generates a response based on the user's friend registration information. Step 38: The server receives the response from the generative AI and sends it to the terminal. Step 39: The device displays the generative AI's response. Step 40: The user logs out and exits the application

[0040] Example 1 Next, a description will be given of Example 1. In the following description, the data processing device 12 is referred to as a "server" and the smart device 14 is referred to as a "terminal."

[0041] Conventional information sharing applications have limited means for users to effectively interact in a virtual environment, making it difficult to obtain satisfactory results, especially in real-time communication and interactive experiences using generative AI. For example, there is a lack of systems that reflect user operations and chat content in real time and provide consistent responses. In addition, the quality of responses and accuracy of suggestions by generative AI have also been issues. In order to solve these problems, we provide the present invention.

[0042] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0043] In this invention, the server includes a means for authenticating the user's login information, a means for acquiring and managing the user's avatar information and game activity information, a means for receiving the user's text chat and voice chat input and sharing it with other users, and a means for generating responses based on the user's input using a generative AI model and sending the responses to the terminal. This allows the user to effectively interact with other users and virtual characters in real time within the virtual environment. Furthermore, the generative AI provides high-quality responses, improving the user's experience.

[0044] "User" refers to an individual who utilizes the system to interact and perform activities within the virtual environment. A "server" refers to a computing device that authenticates user login information, accesses a database to manage information, and responds to requests from users. "Terminal" refers to a device that allows a user to access the system and perform operations and send and receive information. Examples include smartphones and personal computers. "Login Information" refers to authentication information, such as a username and password, that a User enters to be authenticated to a System. "Authentication" refers to the process of verifying that a user can legitimately use a system based on the login information entered by the user. "Avatar" refers to a virtual character that a user uses to represent themselves within a virtual environment. "Game Activity Information" refers to data regarding activity within games and other virtual environments in which a User participates. "Text chat" refers to a means by which a user communicates with other users by inputting text. "Voice chat" refers to a means by which users communicate with other users using voice. A "generative AI model" refers to an artificial intelligence model that generates responses or suggestions based on user input. A "prompt sentence" refers to an input sentence that is provided to a generative AI model to generate some kind of response. "Feedback" refers to the evaluation or opinion given by the user regarding the response of the system or the generative AI model.

[0045] This invention is an information sharing application system that allows users to smoothly interact with other users and virtual characters in a virtual environment. This system is realized by the cooperation and operation of a server, a terminal, a user, and a generative AI model.

[0046] System Overview

[0047] server The server receives the user's login information and performs authentication. For example, an encryption library such as OpenSSL (registered trademark) is used for authentication. The server also stores the user's avatar information and past game activity information in a database (e.g., MySQL (registered trademark)) and manages this information. When the user sends a request, the server retrieves the corresponding data from the database and provides it to the user's terminal.

[0048] In addition, the server receives user input information and forwards it to the generative AI. Generative AI is usually implemented using machine learning libraries such as TENSORFLOW (registered trademark) and PyTorch (registered trademark). The server also manages shared information with other users. For example, if a user is chatting with a friend, the server receives the message and forwards it to the appropriate person.

[0049] Terminal The terminal provides an interface for users to access and log in to the application. It is responsible for sending text chat and voice chat information entered by the user to the server. It also transmits the movements of the avatar selected by the user and game activity to the server, enabling real-time operation.

[0050] The device receives the response from the server and displays it to the user. In particular, the response from the generative AI includes replies to the user's chat and suggestions for avatar movements. For example, in the case of a web application, this process realizes real-time communication on the browser using JavaScript (registered trademark) or WebSocket.

[0051] User Users log into the application using a terminal and interact with other users and virtual characters in the virtual environment. Users can communicate through text and voice chat, and can give their avatars animations to express greetings and emotions.

[0052] In addition, users can select educational games and activities and share them with other users. For example, if a user selects "Math Quiz," quiz questions will be provided by the generative AI, allowing users to compete or collaborate with other users to advance their learning.

[0053] Generative AI Model The generative AI model runs on a server and is responsible for generating responses based on user input. It receives information about the user's text chat, avatar movements, and other user and AI character choices and suggestions, and generates answers accordingly.

[0054] For example, Transformer-based models are often used as generative AI models, and when you input a prompt like "Please suggest that the user start a quiz," the AI ​​generates a response like "Let's ask, How about starting a quiz?"

[0055] Examples Examples of prompt statements 1. If the user wants their avatar to say "Hello!": Prompt: "Have the user say 'Hello!' to their avatar." Generative AI response: "Hello!"

[0056] 2. If the user wants to participate in an educational quiz: Prompt: "The user wants to participate in an educational math quiz." Generative AI response: "Great choice! Now, first question: What is 5 + 7?"

[0057] These elements work together to create a system that allows users to enjoy a wide variety of experiences within a virtual environment.

[0058] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0059] Program processing flow Step 1: User login process explanation User: Launches the application and enters username and password on the login screen. Terminal: Receives the user's login information and sends it to the server. Server: The received login information is compared against a database (e.g. MySQL (registered trademark)) and authentication is performed. Input is the user name and password, and output is the authentication result (success / failure).

[0060] operation The user enters "john_doe" and "password123" into the form and clicks the login button. The terminal encrypts this information and sends it to the server. The server checks the password against the records in its database, and if authentication is successful, it sends a "login successful" message to the terminal.

[0061] Step 2: Get your avatar and game activity information explanation Server: After the user successfully logs in, retrieve the avatar information and past game activity information from the database. Input is the user ID, and output is the retrieved avatar information and activity history. Terminal: displays information received from the server. Input is avatar information and activity information from the server, and output is the display of that information.

[0062] operation The server runs a database query to retrieve avatar information and past game history for user "john_doe". The device displays the acquired avatar information (e.g., "character wearing a blue hat") and past game history (e.g., "played math quiz on October 1st") on the screen.

[0063] Step 3: User interaction and communication explanation User: Controls an avatar and communicates with other users via text chat and voice chat. Terminal: Sends text and voice data entered by the user to the server. Input is text chat and voice data, and output is the result sent to the server. Server: Shares the received data with other users. Input is the received text or voice data, and output is the shared data.

[0064] operation A user types "Hello everyone!" in text chat and presses send. The terminal sends the text to the server. The server sends a notification to other logged-in users saying "john_doe: Hello everyone!"

[0065] Step 4: Generate prompts for generative AI and obtain responses explanation Server: Sends text chat content and avatar movement information to the generative AI. Input is text chat and avatar movement information, output is prompt text and generated responses. Generative AI: Processes the received prompt and generates an appropriate response. The input is the prompt and the output is the generated response. Server: Receives the response from the generative AI, performs further processing if necessary, and sends the response to the terminal. The input is the response from the generative AI, and the output is a notification to the terminal.

[0066] operation The server sends a prompt to the generative AI saying, "The user said 'Hello, everyone!' Please generate an appropriate response." The generative AI generates the response, "Hello john_doe! How are you?" The server receives the response from the generative AI and sends it to the terminal. The device will display the generative AI's response on the chat screen.

[0067] Step 5: Displaying the response and feeding it back to the server explanation Terminal: Receives the generative AI's response and displays it to the user. At the same time, sends the user's reaction and evaluation to the server. The input is the response from the generative AI, and the output is the display and feedback data for the user. User: Checks the generative AI's response and provides feedback if necessary. The input is the response from the generative AI, and the output is evaluation and feedback. Server: Receives user feedback and uses it to improve the generative AI model. The input is feedback data, and the output is feedback to the generative AI model.

[0068] operation The device will display the message "Hello john_doe! How are you?" on the chat screen. The user rates the response as "very good" and sends the feedback to the server. The server stores the feedback and later uses it as training data for the generative AI.

[0069] (Application example 1) Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0070] Conventional educational entertainment and activities in virtual environments have lacked user interaction and personalization, resulting in low learning effectiveness and satisfaction. In addition, there were also insufficient methods for effectively utilizing responses generated in real time and mechanisms for deepening interactions between users.

[0071] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0072] In this invention, the server includes a means for enabling a user to communicate through text messaging or voice messaging, a means for giving a movement to a user's avatar to greet or express emotions, a means for providing educational entertainment or activities using a generation AI, a means for presenting educational entertainment or activities to a user and displaying the results through a visual output device or an audio output device, a means for acquiring a response by the generation AI in real time and presenting the response on a user interface, a means for providing educational entertainment or activities based on an answer obtained by inputting user attribute information and learning information into the generation AI, a data management means for analyzing a user's behavioral data and dialogue history and generating a more personalized response, a means for using the selection or suggestion of other users or AI agents based on the answer from the generation AI, a means for sharing the results of educational entertainment or activities between users and making suggestions to the user based on the answer from the generation AI, a sentence generation means for generating a prompt sentence and inputting it into the generation AI, and an evaluation means for evaluating and improving the response generated based on the prompt sentence. This allows a user to learn while enjoying personalized educational entertainment or activities, and to deepen interactions with other users and AI agents.

[0073] A "virtual environment" is a digital space created using computer technology in which users can have new experiences that are different from the real world. An "information sharing application" is software that enables multiple users to exchange and share information, including data, messages, media, etc., over the Internet. "Text messaging" means a means by which users communicate in real time by sending messages in text form. "Voice messaging" is a means by which users can send messages in the form of voice and communicate in real time. An "avatar" is a digital character that represents a user within a virtual space and is a means for expressing the user's movements and emotions. "Generative AI" is an artificial intelligence technique that generates text, images, etc. based on user input. "Educational entertainment" is games and activities that are designed to entertain users while also providing educational elements. A "visual output device" is a device, such as a digital display or a head-mounted display, that a user uses to obtain visual information. An "audio output device" is a device such as a speaker or headphones that a user uses to obtain audio information. A "user interface" is the arrangement and design of screens and input devices that allow a user to interact with and operate an application. "Attribute information" is data that indicates individual characteristics of a user, such as age, sex, and interests. "Learning information" is data related to the user's past learning history and knowledge level. "Behavioral data" is a record of what operations a user performed within the system. An "interaction history" is a record of messages and responses that have been exchanged between a user and a system in the past. An "AI agent" is a character that uses artificial intelligence and is designed to interact with users in a virtual space. A "prompt" is text that contains instructions or questions that are entered into the generation AI. A "sentence generation means" is a means for creating a prompt sentence to be supplied to the generation AI. An "evaluation means" is a means for evaluating the responses created by the generative AI to find areas for improvement.

[0074] The present invention provides a system for users to experience educational entertainment and activities in a virtual environment and to interact with other users and AI agents in real time. Specific embodiments for implementing the present invention are described below.

[0075] System Configuration The system mainly consists of three elements: the server, the terminal, and the user. The role of each element is explained below.

[0076] 1. Server Authentication means: The server receives the user's login information and performs authentication. For this purpose, a database (DB) system such as MySQL (registered trademark) or PostgreSQL (registered trademark) is used. Data management means: The server manages user avatar information, game activity data, and other user information. This is done using a cloud storage service (e.g., AWS (registered trademark) S3). Generative AI: The generative AI engine generates responses in real time based on user input. ChatGPT (registered trademark) and the like are used as generative AI engines. Communication method: The server uses WebSocket or HTTP protocols to manage communication with the user terminal.

[0077] 2. Terminal User input acceptance means: The terminal accepts the user's text and voice input and sends it to the server, using voice recognition software (e.g., Google® Speech-to-Text API) and a text input interface. Display means: The device displays the generated AI's responses and the results of game activities. This can be done using smart glasses or a head-mounted display. Movement control means: The terminal controls the movement of the avatar based on the user's operation. This is achieved by using motion capture technology.

[0078] 3. Users Input means: Users interact with other users and AI agents in the virtual environment through voice and text input. Avatar Control: Users control their own avatars to express actions and emotions. Activity Participation: Users participate in educational, entertainment, and activities.

[0079] Example of processing flow A specific example of the processing flow of the entire system is shown below.

[0080] 1. User Login The user enters login information from the terminal and transmits it to the server. The server uses the authentication means to authenticate the user.

[0081] 2. Activity selection The user selects educational entertainment and activities through the terminal's interface. The server receives the user's selection and provides the appropriate data.

[0082] 3. Generative AI response generation The user enters a question via text or voice. The server sends prompts to the generation AI, which generates responses in real time. Example response: "The user said, 'Do you want to proceed to the next question?' Generate an appropriate response."

[0083] 4. Display in the interface The server sends the generated response to the terminal, which presents it to the user through its visual and / or audio output devices.

[0084] In this manner, the present invention allows users to enjoy interactive educational entertainment and activities within a virtual environment.

[0085] The flow of the specific process in the application example 1 will be described with reference to FIG.

[0086] Step 1: The server receives the user's login information and performs authentication. The user ID and password entered by the user from the terminal are sent to the server. The server retrieves the corresponding user information from the database and checks whether it matches. The input is the user's login information, and the output is the authentication result (success or failure).

[0087] Step 2: A user selects an educational entertainment or activity using the interface of the terminal. The terminal transmits information about the activity selected by the user to the server. The server retrieves data corresponding to the selected activity from a database and transmits it to the terminal. The input is the information about the activity selected by the user, and the output is the data corresponding to the activity.

[0088] Step 3: The user inputs a question via text or voice. The text or voice data entered by the user into the terminal is sent to the server. The server converts the voice data into text (voice recognition) and sends it to the generation AI as a prompt. The input is the user's question (text, voice), and the output is the prompt.

[0089] Step 4: The server sends a prompt to the generation AI, which generates a response. The generation AI generates a text response based on the prompt and returns it to the server. The server receives the response and adjusts the response content, taking into account the user's attribute information and behavioral data. The input is the prompt and the user's attribute information, and the output is the generated response.

[0090] Step 5: The server sends the generated response to the terminal, which presents it to the user. The terminal displays the response message on a visual output device and optionally plays it through an audio output device. The input is the generated response and the output is the response presented in visual and audio output.

[0091] Step 6: The user confirms the generated response and decides on the next action: the user enters a further question or selects the next activity. The inputs are the user's confirmation and the next action, and the output is the information to proceed to the next step.

[0092] Step 7: The server provides an interaction mechanism for users to share the results of their educational activities. Users perform operations on their terminals to share their results, and the server manages the data. The input is the user's shared data, and the output is the shared information provided to other users.

[0093] The above is a specific processing flow of the program of the system that realizes the application example. This processing flow allows users to enjoy educational entertainment and activities in real time.

[0094] Furthermore, an emotion engine that estimates the emotion of the user may be combined. That is, the identification processing unit 290 may estimate the emotion of the user using the emotion identification model 59, and perform identification processing using the emotion of the user.

[0095] An embodiment for implementing the present invention includes the following elements.

[0096] 1. Server: The server realizes a system that combines an emotion engine that recognizes the user's emotions. The server analyzes the contents of the user's text chat and voice chat and the characteristics of the voice, and estimates the user's emotional state. The server also adjusts the difficulty and content of games and activities based on the user's emotional state estimated by the emotion engine.

[0097] 2. Terminal: The terminal is where the user logs into the application and operates and communicates within the virtual environment. The terminal transmits the user's text and voice chat input to the server, and the server provides games and activities based on the user's emotional state estimated by the emotion engine. The terminal also controls the user's avatar movement and interaction with other users.

[0098] 3. User: Users log into the application and interact with friends and virtual characters in a virtual environment. They communicate through text and voice chat, experience games and activities based on their emotional state estimated by an emotion engine, and interact with other users by giving their avatars animations to greet others and express emotions.

[0099] The above elements work together to realize a system that recognizes the user's emotions and provides games and activities accordingly. The server uses the emotion engine to estimate the user's emotional state and provide appropriate content. The terminal accepts the user's operations, communicates with the server, and displays responses via the emotion engine. The user can interact with friends and virtual characters in the virtual environment and enjoy content that matches their emotions.

[0100] The process flow will be explained below. Step 1: User logs into the application Step 2: The device sends the user's login information to the server Step 3: The server receives the user's login information and performs authentication. Step 4: The server sends the authentication result to the terminal. Step 5: The user selects an avatar to enter the virtual environment. Step 6: The device sends the user's avatar information to the server. Step 7: The server receives the user's avatar information and displays the avatar in the virtual environment. Step 8: The user interacts with friends and virtual characters through text and voice chat. Step 9: The device sends the user's chat input to the server Step 10: The server receives the user's chat input and passes it as input to the emotion engine. Step 11: The emotion engine analyzes the user's chat input and infers their emotional state Step 12: The server receives the results from the emotion engine and selects the appropriate game or activity. Step 13: The server sends the selected game or activity to the device. Step 14: The device displays the selected game or activity for the user to experience. Step 15: The user gives the avatar movement to greet or express emotion. Step 16: The terminal transmits the user's avatar movements to the server. Step 17: The server receives the user's avatar movements and moves the avatar in the virtual environment. Step 18: The user plays a game or activity with other users or AI characters. Step 19: The device sends the choices and suggestions of other users and AI characters to the server. Step 20: The server uses the choices and suggestions of other users and AI characters as inputs to the emotion engine. Step 21: The emotion engine generates responses based on the choices and suggestions of other users and AI characters. Step 22: The server receives the emotion engine's response and sends it to the terminal. Step 23: The device displays the emotion engine's response. Step 24: The user shares the results of the game or activity with other users. Step 25: The device sends the user's sharing contacts and chat contents to the server. Step 26: The server uses the user's shared contacts and chat contents as input to the emotion engine. Step 27: The emotion engine generates a response based on who the user is sharing with and what the chat is about. Step 28: The server receives the emotion engine's response and sends it to the terminal. Step 29: The device displays the emotion engine's response. Step 30: Use information about other users who are registered as friends by the user Step 31: User sends invitation notification and points to friends Step 32: The device sends the user's friend registration information to the server. Step 33: The server receives the user's friend registration information and uses it as input to the emotion engine. Step 34: The emotion engine generates a response based on the user's friend information. Step 35: The server receives the emotion engine's response and sends it to the terminal. Step 36: The device displays the emotion engine's response. Step 37: The user logs out and exits the application

[0101] Example 2 Next, a description will be given of Example 2. In the following description, the data processing device 12 is referred to as a "server" and the smart device 14 is referred to as a "terminal."

[0102] Conventional communication systems and information sharing applications in virtual environments have the problem that it is difficult to adjust content in real time according to the user's emotional state. In addition, there is a lack of systems that can properly analyze the user's emotional state and provide games and activities based on it. This results in a uniform user experience and a lack of personalized support tailored to each individual's emotional state.

[0103] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for estimating the emotional state of the user by utilizing an emotion analysis engine, a means for adjusting the content of the game or activity based on the estimated emotional state, and a means for converting voice data into text data. This makes it possible to adjust the content in real time according to the emotional state of the user, and to provide a personalized experience.

[0104] A "virtual environment" is a computer-generated digital space that a user can interact with. A "communication application" is software that enables users to interact with other users and virtual characters through text and / or voice chat. "Text chat" is a means by which users exchange messages in real time through text input. "Voice chat" is a means by which users communicate in real time through voice input. "Character" refers to an avatar controlled by a user within a virtual environment, or a virtual entity controlled by a computer. An "emotion analysis engine" is an algorithm or software used to infer a user's emotional state from text or voice input. An "emotional state" refers to the psychological state a user feels at a particular point in time and can be classified as positive, negative, neutral, etc. "Games and Activities" are interactive entertainment or learning activities that users enjoy within a virtual environment. "Means for converting voice data into text data" refers to technology or software used to convert a user's voice input into text form. "Personalization" means optimizing systems and content to suit the characteristics and circumstances of each individual user. "Real-time" refers to responding immediately to user input and changes in the environment and processing without delay.

[0105] The following system configuration will be described as an embodiment of the present invention. The main components are a server, a terminal, and a user's operation method.

[0106] The server runs on a cloud service and is equipped with multiple function engines, including a sentiment analysis engine and a conversion engine. Specifically, the server uses a Natural Language Processing (NLP) API for sentiment analysis and a speech recognition API for converting voice data into text data. This makes it possible to analyze the user's emotional state and generate and provide interactive content in real time according to that state.

[0107] A terminal is a device that provides an interface for users to access applications and perform activities and communications within the virtual environment. Examples of terminals include PCs, smartphones, and tablets. The terminal has the function of sending the user's text chat and voice chat input to the server. The terminal also displays content to the user based on the emotional state received from the server. Furthermore, the terminal accepts user operations in real time and performs the actions and emotional expressions of the virtual character.

[0108] Users log into the application through their terminals and interact with other users and virtual characters in the virtual environment. The user's text and voice chat inputs are analyzed by an emotion analysis engine, and appropriate games and activities are provided based on the user's emotional state. For example, when a user says "hello," a character responds with a smile.

[0109] Examples

[0110] Example 1: A user interacts with a friend through text chat. 1. A user logs into the application from a terminal and starts a text chat. He types, "What's the weather like today?" 2. The device sends this text to the server. 3. The server uses a sentiment analysis engine to analyze "What's the weather like today?" and determines the emotional state as "neutral." 4. The server generates weather-related interactive activities as additional content and sends them to the terminal. 5. The terminal displays this information to the user, who can then enjoy further weather information or other activities.

[0111] Example 2: When a user interacts with a virtual character through voice chat 1. The user uses voice chat to talk to a virtual character, asking, "How are you?" 2. The terminal transmits the voice data to the server. 3. The server uses a voice recognition engine to convert the voice data into text data such as "How are you?" 4. The server passes the converted text data to a sentiment analysis engine and determines the emotional state as "positive." 5. The server generates an activity to respond cheerfully based on this positive emotional state. 6. The device displays this information to the user and the virtual character responds with something like "I'm fine! How about you?"

[0112] Examples of prompt statements "If the emotion engine analyzes the user's emotional state and estimates it to be 'happy,' please provide specific ideas for what kind of content (games or activities) should be provided."

[0113] From the above system configuration and specific examples, it can be understood that the emotion recognition and real-time content provision method according to the present invention can be clearly implemented.

[0114] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0115] Step 1: A user logs into the application from a terminal. A username and password are required as input. If the login is successful, the user profile and usage history are output from the server and displayed on the terminal. Specifically, the user enters his / her authentication information on the login screen and clicks the "Login" button.

[0116] Step 2: A user initiates a text or voice chat. As input, the text message or voice data is sent to the server. The server receives this data and prepares it for processing. For example, if a user sends a text message to a friend saying "How are you doing?", the device sends this message to the server.

[0117] Step 3: When the server receives the voice data, it converts the voice data into text data using Amazon (registered trademark) Transcribe. The server receives the voice data as input and generates converted text data as output. Specifically, the voice data of "How are you doing lately?" is converted into the text data of "How are you doing lately?"

[0118] Step 4: The server sends the received or converted text data to the Google Cloud Natural Language API for sentiment analysis. Text data is provided as input, and a sentiment score and a judgment result of the sentiment state are returned as output. For example, the text "How are you doing lately?" is analyzed to detect the sentiment state "neutral."

[0119] Step 5: The server generates appropriate content based on the results of the emotion analysis. It uses the emotional state determination as input and generates games and activities as output. Specifically, relaxing content is selected and generated for a "neutral" emotional state.

[0120] Step 6: The server sends the generated content to the terminal. As input, the server provides the generated content and sends information to the terminal. As output, the terminal displays this content to the user. For example, a puzzle game for relaxation is displayed on the terminal.

[0121] Step 7: The user uses the provided content. User operation data and additional feedback data are generated and sent to the server. Operation information and feedback from the user are received as input, and this is sent to the server as output. Specifically, when the user inputs his / her impressions while solving the puzzle, the data is sent to the server.

[0122] Step 8: The server again inputs the received feedback into the sentiment analysis engine and re-adjusts the content if necessary. Using the feedback data as input, the server provides the user with updated content or a re-analysis of the emotional state as output. For example, based on the feedback "This puzzle is a bit difficult", new content with a slightly lower difficulty level is generated.

[0123] (Application example 2) Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0124] Conventional applications for exchanging information in virtual environments lacked the ability to grasp the user's emotional state in real time and recommend and play optimal content. This made it difficult to provide content that was in tune with the user's emotions, limiting the user experience. Furthermore, there was no system that integrated educational elements with the provision of content based on emotions. A system that can solve these problems is needed.

[0125] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0126] In this invention, the server includes a means for receiving text chat or voice chat input by a user and presenting the received input to other users to realize exchanges by text chat or voice chat, a means for giving a movement to the user's avatar in response to the user's operation and having it express at least one of a greeting and an emotion, a means for inputting user information to a generation AI and providing the user with at least one of an educational game and an activity based on a response from the generation AI, and a means for analyzing the user's emotions in real time and recommending and playing appropriate entertainment (music or video content) based on the emotions. This makes it possible to provide content that is in line with the user's emotions, thereby providing a richer user experience.

[0127] An "information exchange application" is software that enables a user to communicate with friends and virtual characters within a virtual environment through text chat and voice chat. An "avatar" is a digital character that represents a user within a virtual environment and expresses the user's actions and emotions. "Generative AI" is artificial intelligence that automatically recommends educational games, activities, and content based on the user's input information and emotional state. "Emotion analysis" is a technology that analyzes the content of a user's text chat or voice chat and infers the user's emotional state. "Content recommendation" is a function that suggests optimal music and video content based on the user's emotional state and attribute information. "Educational games and activities" are interactive games and activities designed to learn or enhance knowledge. "Real-time" refers to immediate processing or response in the current time. A "virtual environment" is a computer-generated digital space with which a user interacts.

[0128] The embodiment for implementing the invention includes the following elements.

[0129] server The server has a built-in emotion engine that recognizes the user's emotions. Specifically, the emotion engine analyzes and estimates the user's emotional state from the contents of text and voice chat. The server utilizes a generative AI model to recommend and play optimal music and video content based on the user's emotional state. The server also adjusts the difficulty and content of educational games and activities according to the user's emotional state. The system includes the following software: emotion_engine(emotion engine) content_recommender (content recommendation system) text_analysis audio_recognition

[0130] Terminal The terminal is the device through which the user logs into the application and performs operations and communication within the virtual environment. The terminal receives the user's voice and text input and transmits it to the server. It also provides the user with content and activities based on the emotional state estimated by the emotion engine. Terminals include smartphones and smart glasses. The terminal has the following functions: Function to give movement to avatars Text and voice chat functionality Content playback function

[0131] User Users log into the application and interact with friends and virtual characters in a virtual environment. They communicate through text and voice chat, and enjoy music and video content based on their emotional state estimated by the emotion engine. They also participate in educational games and activities, experiencing adjustments in difficulty and content according to their emotional state. An example of a specific user scenario is as follows:

[0132] 1. The user launches the app and speaks, "I'm feeling a bit sad today." 2. The app analyzes the voice and the emotion engine determines that the user is feeling "sad." 3. The content recommendation system recommends soothing music and inspiring videos that suit the emotion of "sad." 4. The recommended content will be played automatically.

[0133] Examples of prompt statements User input: "I'm feeling a bit sad today." Analyzed emotion: "Sadness" Suitable content: "Healing music" "Inspirational videos" Suggested prompt: "If the user is feeling sad, please recommend soothing music that will soothe the soul. Also, please recommend an inspiring video."

[0134] In this way, content that is in tune with the user's emotions can be provided, providing a rich user experience.

[0135] The flow of the specific process in the application example 2 will be described with reference to FIG.

[0136] Step 1: The device receives the user's text chat or voice chat input. When the user speaks "I'm feeling a bit sad today," the device receives this voice data. The input data is voice data. The device then uses voice recognition software to convert this voice data into text data. The converted text data is output.

[0137] Step 2: The terminal transmits the converted text data to the server. The input here is text data, and the output to the server is also text data. Specifically, the terminal transmits the text data to the server via the network.

[0138] Step 3: The server inputs the received text data into an emotion engine. The emotion engine analyzes the user's emotional state from the text data. The input here is the text data, and the emotion analysis engine estimates the emotional state. The estimated emotional state, e.g., "sadness," is output.

[0139] Step 4: The server inputs the estimated emotional state into a content recommendation system, which recommends optimal music and video content based on the emotional state. The input is the emotional state, and the content recommendation system uses this data to select recommended content. The selected recommended content is output.

[0140] Step 5: The server transmits the recommended content to the terminal. The input is the recommended content data, and the output is also the recommended content data. The server transmits this to the terminal via the network.

[0141] Step 6: The terminal presents the received content to the user. Specifically, it uses the terminal's media player function to play music and videos. The input is the recommended content data, and the output is the content to be played that is presented to the user. The terminal plays the content using a screen display and speakers.

[0142] In this way, a system is realized that provides content in real time according to the user's emotions.

[0143] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires a voice indicating a user input for the result of the specific processing. The control unit 46A transmits the voice data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0144] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0145] In the above embodiment, an example was given in which the specific process was performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0146] [Second embodiment] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0147] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0148] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).

[0149] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0150] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs the voice according to instructions from the processor 46.

[0151] Camera 42 is a small digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (e.g., an imaging range defined by an angle of view equivalent to the width of the field of vision of an average healthy person).

[0152] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54. The transmission and reception of various types of information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is performed in a secure state.

[0153] Fig. 4 shows an example of main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0154] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32, and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0155] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0156] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0157] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal".

[0158] An embodiment for implementing the present invention includes the following elements.

[0159] 1. Server: The server receives the user's login information and performs authentication. It also manages the user's avatar information and data required for game activities, and provides appropriate content in response to the user's requests. It also processes input to the generative AI and information shared with other users.

[0160] 2. Terminal: The terminal is where the user logs into the application and operates and communicates within the virtual environment. It transmits the user's text and voice chat input to the server, receives and displays the generative AI's responses, and transmits the user's avatar movements and game activity selections to the server.

[0161] 3. User: Users log into the application and interact with friends and virtual characters in a virtual environment, communicating through text and voice chat, animating their avatars to greet others, express emotions, and select educational games and activities to share with other users.

[0162] 4. Generative AI: A generative AI is a generative AI that generates responses based on user input. It receives information from the user's text chat and avatar movements, as well as information about the choices and suggestions of other users and AI characters, and generates appropriate responses. Generative AI runs on the server and processes input and output from the server.

[0163] All of these elements work together to allow users to interact with friends and virtual characters in a virtual environment and enjoy educational games and activities. The server manages user information and provides appropriate content. The device accepts user operations, communicates with the server, and displays the generative AI's responses. The generative AI generates responses based on the user's input and sends them to the device via the server.

[0164] The process flow will be explained below. Step 1: User logs into the application Step 2: The device sends the user's login information to the server Step 3: The server receives the user's login information and performs authentication. Step 4: The server sends the authentication result to the terminal. Step 5: The user selects an avatar to enter the virtual environment. Step 6: The device sends the user's avatar information to the server. Step 7: The server receives the user's avatar information and displays the avatar in the virtual environment. Step 8: The user interacts with friends and virtual characters through text and voice chat. Step 9: The device sends the user's chat input to the server Step 10: The server receives the user's chat input and passes it to the generative AI. Step 11: Generative AI generates a response based on the user's chat input Step 12: The server receives the response from the generative AI and sends it to the device. Step 13: The device displays the generative AI's response Step 14: The user gives the avatar movement to greet or express emotion. Step 15: The device sends the user's avatar movements to the server. Step 16: The server receives the user's avatar movements and moves the avatar in the virtual environment. Step 17: User selects educational games and activities Step 18: The device sends the user's selection to the server Step 19: The server serves games and activities based on the user's selection. Step 20: User experiences the game or activity Step 21: The user plays a game or activity with other users or AI characters. Step 22: The device sends the choices and suggestions of other users and AI characters to the server. Step 23: The server uses the choices and suggestions of other users and AI characters as inputs to the generative AI. Step 24: The generative AI generates a response based on the choices and suggestions of other users and the AI ​​character. Step 25: The server receives the response from the generative AI and sends it to the terminal. Step 26: The device displays the generative AI's response. Step 27: The user shares the results of the game or activity with other users. Step 28: The device sends the user's sharing contacts and chat contents to the server. Step 29: The server uses the user's shared contacts and chat contents as input to the generative AI. Step 30: Generative AI generates responses based on who the user is sharing with and what the chat is about. Step 31: The server receives the response from the generative AI and sends it to the terminal. Step 32: The device displays the generative AI's response. Step 33: Use information from other users who are registered as friends by the user Step 34: User sends invitation notification and points to friends Step 35: The device sends the user's friend registration information to the server. Step 36: The server receives the user's friend registration information and uses it as input for the generative AI. Step 37: The generative AI generates a response based on the user's friend registration information. Step 38: The server receives the response from the generative AI and sends it to the terminal. Step 39: The device displays the generative AI's response. Step 40: The user logs out and exits the application

[0165] Example 1 Next, a description will be given of Example 1. In the following description, the data processing device 12 is referred to as a "server" and the smart glasses 214 are referred to as a "terminal".

[0166] Conventional information sharing applications have limited means for users to effectively interact in a virtual environment, making it difficult to obtain satisfactory results, especially in real-time communication and interactive experiences using generative AI. For example, there is a lack of systems that reflect user operations and chat content in real time and provide consistent responses. In addition, the quality of responses and accuracy of suggestions by generative AI have also been issues. In order to solve these problems, we provide the present invention.

[0167] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0168] In this invention, the server includes a means for authenticating the user's login information, a means for acquiring and managing the user's avatar information and game activity information, a means for receiving the user's text chat and voice chat input and sharing it with other users, and a means for generating responses based on the user's input using a generative AI model and sending the responses to the terminal. This allows the user to effectively interact with other users and virtual characters in real time within the virtual environment. Furthermore, the generative AI provides high-quality responses, improving the user's experience.

[0169] "User" refers to an individual who utilizes the system to interact and perform activities within the virtual environment. A "server" refers to a computing device that authenticates user login information, accesses a database to manage information, and responds to requests from users. "Terminal" refers to a device that allows a user to access the system and perform operations and send and receive information. Examples include smartphones and personal computers. "Login Information" refers to authentication information, such as a username and password, that a User enters to be authenticated to a System. "Authentication" refers to the process of verifying that a user can legitimately use a system based on the login information entered by the user. "Avatar" refers to a virtual character that a user uses to represent themselves within a virtual environment. "Game Activity Information" refers to data regarding activity within games and other virtual environments in which a User participates. "Text chat" refers to a means by which a user communicates with other users by inputting text. "Voice chat" refers to a means by which users communicate with other users using voice. A "generative AI model" refers to an artificial intelligence model that generates responses or suggestions based on user input. A "prompt sentence" refers to an input sentence that is provided to a generative AI model to generate some kind of response. "Feedback" refers to the evaluation or opinion given by the user regarding the response of the system or the generative AI model.

[0170] This invention is an information sharing application system that allows users to smoothly interact with other users and virtual characters in a virtual environment. This system is realized by the cooperation and operation of a server, a terminal, a user, and a generative AI model.

[0171] System Overview

[0172] server The server receives the user's login information and performs authentication. For example, an encryption library such as OpenSSL (registered trademark) is used for authentication. The server also stores the user's avatar information and past game activity information in a database (e.g., MySQL (registered trademark)) and manages this information. When the user sends a request, the server retrieves the corresponding data from the database and provides it to the user's terminal.

[0173] In addition, the server receives user input information and forwards it to the generative AI. Generative AI is usually implemented using machine learning libraries such as TensorFlow (registered trademark) and PyTorch (registered trademark). The server also manages shared information with other users. For example, if a user is chatting with a friend, the server receives the message and forwards it to the appropriate person.

[0174] Terminal The terminal provides an interface for users to access and log in to the application. It is responsible for sending text chat and voice chat information entered by the user to the server. It also transmits the movements of the avatar selected by the user and game activity to the server, enabling real-time operation.

[0175] The device receives the response from the server and displays it to the user. In particular, the response from the generative AI includes replies to the user's chat and suggestions for avatar movements. For example, in the case of a web application, this process realizes real-time communication on the browser using JavaScript (registered trademark) or WebSocket.

[0176] User Users log into the application using a terminal and interact with other users and virtual characters in the virtual environment. Users can communicate through text and voice chat, and can give their avatars animations to express greetings and emotions.

[0177] In addition, users can select educational games and activities and share them with other users. For example, if a user selects "Math Quiz," quiz questions will be provided by the generative AI, allowing users to compete or collaborate with other users to advance their learning.

[0178] Generative AI Model The generative AI model runs on a server and is responsible for generating responses based on user input. It receives information about the user's text chat, avatar movements, and other user and AI character choices and suggestions, and generates answers accordingly.

[0179] For example, Transformer-based models are often used as generative AI models, and when you input a prompt like "Please suggest that the user start a quiz," the AI ​​generates a response like "Let's ask, How about starting a quiz?"

[0180] Examples Examples of prompt statements 1. If the user wants their avatar to say "Hello!": Prompt: "Have the user say 'Hello!' to their avatar." Generative AI response: "Hello!"

[0181] 2. If the user wants to participate in an educational quiz: Prompt: "The user wants to participate in an educational math quiz." Generative AI response: "Great choice! Now, first question: What is 5 + 7?"

[0182] These elements work together to create a system that allows users to enjoy a wide variety of experiences within a virtual environment.

[0183] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0184] Program processing flow Step 1: User login process explanation User: Launches the application and enters username and password on the login screen. Terminal: Receives the user's login information and sends it to the server. Server: The received login information is compared against a database (e.g. MySQL (registered trademark)) and authentication is performed. Input is the user name and password, and output is the authentication result (success / failure).

[0185] operation The user enters "john_doe" and "password123" into the form and clicks the login button. The terminal encrypts this information and sends it to the server. The server checks the password against the records in its database, and if authentication is successful, it sends a "login successful" message to the terminal.

[0186] Step 2: Get your avatar and game activity information explanation Server: After the user successfully logs in, retrieve the avatar information and past game activity information from the database. Input is the user ID, and output is the retrieved avatar information and activity history. Terminal: displays information received from the server. Input is avatar information and activity information from the server, and output is the display of that information.

[0187] operation The server runs a database query to retrieve avatar information and past game history for user "john_doe". The device displays the acquired avatar information (e.g., "character wearing a blue hat") and past game history (e.g., "played math quiz on October 1st") on the screen.

[0188] Step 3: User interaction and communication explanation User: Controls an avatar and communicates with other users via text chat and voice chat. Terminal: Sends text and voice data entered by the user to the server. Input is text chat and voice data, and output is the result sent to the server. Server: Shares the received data with other users. Input is the received text or voice data, and output is the shared data.

[0189] operation A user types "Hello everyone!" in text chat and presses send. The terminal sends the text to the server. The server sends a notification to other logged-in users saying "john_doe: Hello everyone!"

[0190] Step 4: Generate prompts for generative AI and obtain responses explanation Server: Sends text chat content and avatar movement information to the generative AI. Input is text chat and avatar movement information, output is prompt text and generated responses. Generative AI: Processes the received prompt and generates an appropriate response. The input is the prompt and the output is the generated response. Server: Receives the response from the generative AI, performs further processing if necessary, and sends the response to the terminal. The input is the response from the generative AI, and the output is a notification to the terminal.

[0191] operation The server sends a prompt to the generative AI saying, "The user said 'Hello, everyone!' Please generate an appropriate response." The generative AI generates the response, "Hello john_doe! How are you?" The server receives the response from the generative AI and sends it to the terminal. The device will display the generative AI's response on the chat screen.

[0192] Step 5: Displaying the response and feeding it back to the server explanation Terminal: Receives the generative AI's response and displays it to the user. At the same time, sends the user's reaction and evaluation to the server. The input is the response from the generative AI, and the output is the display and feedback data for the user. User: Checks the generative AI's response and provides feedback if necessary. The input is the response from the generative AI, and the output is evaluation and feedback. Server: Receives user feedback and uses it to improve the generative AI model. The input is feedback data, and the output is feedback to the generative AI model.

[0193] operation The device will display the message "Hello john_doe! How are you?" on the chat screen. The user rates the response as "very good" and sends the feedback to the server. The server stores the feedback and later uses it as training data for the generative AI.

[0194] (Application example 1) Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal".

[0195] Conventional educational entertainment and activities in virtual environments have lacked user interaction and personalization, resulting in low learning effectiveness and satisfaction. In addition, there were also insufficient methods for effectively utilizing responses generated in real time and mechanisms for deepening interactions between users.

[0196] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0197] In this invention, the server includes a means for enabling a user to communicate through text messaging or voice messaging, a means for giving a movement to a user's avatar to greet or express emotions, a means for providing educational entertainment or activities using a generation AI, a means for presenting educational entertainment or activities to a user and displaying the results through a visual output device or an audio output device, a means for acquiring a response by the generation AI in real time and presenting the response on a user interface, a means for providing educational entertainment or activities based on an answer obtained by inputting user attribute information and learning information into the generation AI, a data management means for analyzing a user's behavioral data and dialogue history and generating a more personalized response, a means for using the selection or suggestion of other users or AI agents based on the answer from the generation AI, a means for sharing the results of educational entertainment or activities between users and making suggestions to the user based on the answer from the generation AI, a sentence generation means for generating a prompt sentence and inputting it into the generation AI, and an evaluation means for evaluating and improving the response generated based on the prompt sentence. This allows a user to learn while enjoying personalized educational entertainment or activities, and to deepen interactions with other users and AI agents.

[0198] A "virtual environment" is a digital space created using computer technology in which users can have new experiences that are different from the real world. An "information sharing application" is software that enables multiple users to exchange and share information, including data, messages, media, etc., over the Internet. "Text messaging" means a means by which users communicate in real time by sending messages in text form. "Voice messaging" is a means by which users can send messages in the form of voice and communicate in real time. An "avatar" is a digital character that represents a user within a virtual space and is a means for expressing the user's movements and emotions. "Generative AI" is an artificial intelligence technique that generates text, images, etc. based on user input. "Educational entertainment" is games and activities that are designed to entertain users while also providing educational elements. A "visual output device" is a device, such as a digital display or a head-mounted display, that a user uses to obtain visual information. An "audio output device" is a device such as a speaker or headphones that a user uses to obtain audio information. A "user interface" is the arrangement and design of screens and input devices that allow a user to interact with and operate an application. "Attribute information" is data that indicates individual characteristics of a user, such as age, sex, and interests. "Learning information" is data related to the user's past learning history and knowledge level. "Behavioral data" is a record of what operations a user performed within the system. An "interaction history" is a record of messages and responses that have been exchanged between a user and a system in the past. An "AI agent" is a character that uses artificial intelligence and is designed to interact with users in a virtual space. A "prompt" is text that contains instructions or questions that are entered into the generation AI. A "sentence generation means" is a means for creating a prompt sentence to be supplied to the generation AI. An "evaluation means" is a means for evaluating the responses created by the generative AI to find areas for improvement.

[0199] The present invention provides a system for users to experience educational entertainment and activities in a virtual environment and to interact with other users and AI agents in real time. Specific embodiments for implementing the present invention are described below.

[0200] System Configuration The system mainly consists of three elements: the server, the terminal, and the user. The role of each element is explained below.

[0201] 1. Server Authentication means: The server receives the user's login information and performs authentication. For this purpose, a database (DB) system such as MySQL (registered trademark) or PostgreSQL (registered trademark) is used. Data management means: The server manages user avatar information, game activity data, and other user information. This is done using a cloud storage service (e.g., AWS S3 (registered trademark)). Generative AI: The generative AI engine generates responses in real time based on user input. ChatGPT (registered trademark) and the like are used as generative AI engines. Communication method: The server uses WebSocket or HTTP protocols to manage communication with the user terminal.

[0202] 2. Terminal User input acceptance means: The terminal accepts the user's text and voice input and sends it to the server, using voice recognition software (e.g., Google® Speech-to-Text API) and a text input interface. Display means: The device displays the generated AI's responses and the results of game activities. This can be done using smart glasses or a head-mounted display. Movement control means: The terminal controls the movement of the avatar based on the user's operation. This is achieved by using motion capture technology.

[0203] 3. Users Input means: Users interact with other users and AI agents in the virtual environment through voice and text input. Avatar Control: Users control their own avatars to express actions and emotions. Activity Participation: Users participate in educational, entertainment, and activities.

[0204] Example of processing flow A specific example of the processing flow of the entire system is shown below.

[0205] 1. User Login The user enters login information from the terminal and transmits it to the server. The server uses the authentication means to authenticate the user.

[0206] 2. Activity selection The user selects educational entertainment and activities through the terminal's interface. The server receives the user's selection and provides the appropriate data.

[0207] 3. Generative AI response generation The user enters a question via text or voice. The server sends prompts to the generation AI, which generates responses in real time. Example response: "The user said, 'Do you want to proceed to the next question?' Generate an appropriate response."

[0208] 4. Display in the interface The server sends the generated response to the terminal, which presents it to the user through its visual and / or audio output devices.

[0209] In this manner, the present invention allows users to enjoy interactive educational entertainment and activities within a virtual environment.

[0210] The flow of the specific process in the application example 1 will be described with reference to FIG.

[0211] Step 1: The server receives the user's login information and performs authentication. The user ID and password entered by the user from the terminal are sent to the server. The server retrieves the corresponding user information from the database and checks whether it matches. The input is the user's login information, and the output is the authentication result (success or failure).

[0212] Step 2: A user selects an educational entertainment or activity using the interface of the terminal. The terminal transmits information about the activity selected by the user to the server. The server retrieves data corresponding to the selected activity from a database and transmits it to the terminal. The input is the information about the activity selected by the user, and the output is the data corresponding to the activity.

[0213] Step 3: The user inputs a question via text or voice. The text or voice data entered by the user into the terminal is sent to the server. The server converts the voice data into text (voice recognition) and sends it to the generation AI as a prompt. The input is the user's question (text, voice), and the output is the prompt.

[0214] Step 4: The server sends a prompt to the generation AI, which generates a response. The generation AI generates a text response based on the prompt and returns it to the server. The server receives the response and adjusts the response content, taking into account the user's attribute information and behavioral data. The input is the prompt and the user's attribute information, and the output is the generated response.

[0215] Step 5: The server sends the generated response to the terminal, which presents it to the user. The terminal displays the response message on a visual output device and optionally plays it through an audio output device. The input is the generated response and the output is the response presented in visual and audio output.

[0216] Step 6: The user confirms the generated response and decides on the next action: the user enters a further question or selects the next activity. The inputs are the user's confirmation and the next action, and the output is the information to proceed to the next step.

[0217] Step 7: The server provides an interaction mechanism for users to share the results of their educational activities. Users perform operations on their terminals to share their results, and the server manages the data. The input is the user's shared data, and the output is the shared information provided to other users.

[0218] The above is a specific processing flow of the program of the system that realizes the application example. This processing flow allows users to enjoy educational entertainment and activities in real time.

[0219] In addition, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0220] An embodiment for implementing the present invention includes the following elements.

[0221] 1. Server: The server realizes a system that combines an emotion engine that recognizes the user's emotions. The server analyzes the contents of the user's text chat and voice chat and the characteristics of the voice, and estimates the user's emotional state. The server also adjusts the difficulty and content of games and activities based on the user's emotional state estimated by the emotion engine.

[0222] 2. Terminal: The terminal is where the user logs into the application and operates and communicates within the virtual environment. The terminal transmits the user's text and voice chat input to the server, and the server provides games and activities based on the user's emotional state estimated by the emotion engine. The terminal also controls the user's avatar movement and interaction with other users.

[0223] 3. User: Users log into the application and interact with friends and virtual characters in a virtual environment. They communicate through text and voice chat, experience games and activities based on their emotional state estimated by an emotion engine, and interact with other users by giving their avatars animations to greet others and express emotions.

[0224] The above elements work together to realize a system that recognizes the user's emotions and provides games and activities accordingly. The server uses the emotion engine to estimate the user's emotional state and provide appropriate content. The terminal accepts the user's operations, communicates with the server, and displays responses via the emotion engine. The user can interact with friends and virtual characters in the virtual environment and enjoy content that matches their emotions.

[0225] The process flow will be explained below. Step 1: User logs into the application Step 2: The device sends the user's login information to the server Step 3: The server receives the user's login information and performs authentication. Step 4: The server sends the authentication result to the terminal. Step 5: The user selects an avatar to enter the virtual environment. Step 6: The device sends the user's avatar information to the server. Step 7: The server receives the user's avatar information and displays the avatar in the virtual environment. Step 8: The user interacts with friends and virtual characters through text and voice chat. Step 9: The device sends the user's chat input to the server Step 10: The server receives the user's chat input and passes it as input to the emotion engine. Step 11: The emotion engine analyzes the user's chat input and infers their emotional state Step 12: The server receives the results from the emotion engine and selects the appropriate game or activity. Step 13: The server sends the selected game or activity to the device. Step 14: The device displays the selected game or activity for the user to experience. Step 15: The user gives the avatar movement to greet or express emotion. Step 16: The terminal transmits the user's avatar movements to the server. Step 17: The server receives the user's avatar movements and moves the avatar in the virtual environment. Step 18: The user plays a game or activity with other users or AI characters. Step 19: The device sends the choices and suggestions of other users and AI characters to the server. Step 20: The server uses the choices and suggestions of other users and AI characters as inputs to the emotion engine. Step 21: The emotion engine generates responses based on the choices and suggestions of other users and AI characters. Step 22: The server receives the emotion engine's response and sends it to the terminal. Step 23: The device displays the emotion engine's response. Step 24: The user shares the results of the game or activity with other users. Step 25: The device sends the user's sharing contacts and chat contents to the server. Step 26: The server uses the user's shared contacts and chat contents as input to the emotion engine. Step 27: The emotion engine generates a response based on who the user is sharing with and what the chat is about. Step 28: The server receives the emotion engine's response and sends it to the terminal. Step 29: The device displays the emotion engine's response. Step 30: Use information about other users who are registered as friends by the user Step 31: User sends invitation notification and points to friends Step 32: The device sends the user's friend registration information to the server. Step 33: The server receives the user's friend registration information and uses it as input to the emotion engine. Step 34: The emotion engine generates a response based on the user's friend information. Step 35: The server receives the emotion engine's response and sends it to the terminal. Step 36: The device displays the emotion engine's response. Step 37: The user logs out and exits the application

[0226] Example 2 Next, a description will be given of Example 2. In the following description, the data processing device 12 is referred to as a "server" and the smart glasses 214 are referred to as a "terminal".

[0227] Conventional communication systems and information sharing applications in virtual environments have the problem that it is difficult to adjust content in real time according to the user's emotional state. In addition, there is a lack of systems that can properly analyze the user's emotional state and provide games and activities based on it. This results in a uniform user experience and a lack of personalized support tailored to each individual's emotional state.

[0228] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for estimating the emotional state of the user by utilizing an emotion analysis engine, a means for adjusting the content of the game or activity based on the estimated emotional state, and a means for converting voice data into text data. This makes it possible to adjust the content in real time according to the emotional state of the user, and to provide a personalized experience.

[0229] A "virtual environment" is a computer-generated digital space that a user can interact with. A "communication application" is software that enables users to interact with other users and virtual characters through text and / or voice chat. "Text chat" is a means by which users exchange messages in real time through text input. "Voice chat" is a means by which users communicate in real time through voice input. "Character" refers to an avatar controlled by a user within a virtual environment, or a virtual entity controlled by a computer. An "emotion analysis engine" is an algorithm or software used to infer a user's emotional state from text or voice input. An "emotional state" refers to the psychological state a user feels at a particular point in time and can be classified as positive, negative, neutral, etc. "Games and Activities" are interactive entertainment or learning activities that users enjoy within a virtual environment. "Means for converting voice data into text data" refers to technology or software used to convert a user's voice input into text form. "Personalization" means optimizing systems and content to suit the characteristics and circumstances of each individual user. "Real-time" refers to responding immediately to user input and changes in the environment and processing without delay.

[0230] The following system configuration will be described as an embodiment of the present invention. The main components are a server, a terminal, and a user's operation method.

[0231] The server runs on a cloud service and is equipped with multiple function engines, including a sentiment analysis engine and a conversion engine. Specifically, the server uses a Natural Language Processing (NLP) API for sentiment analysis and a speech recognition API for converting voice data into text data. This makes it possible to analyze the user's emotional state and generate and provide interactive content in real time according to that state.

[0232] A terminal is a device that provides an interface for users to access applications and perform activities and communications within the virtual environment. Examples of terminals include PCs, smartphones, and tablets. The terminal has the function of sending the user's text chat and voice chat input to the server. The terminal also displays content to the user based on the emotional state received from the server. Furthermore, the terminal accepts user operations in real time and performs the actions and emotional expressions of the virtual character.

[0233] Users log into the application through their terminals and interact with other users and virtual characters in the virtual environment. The user's text and voice chat inputs are analyzed by an emotion analysis engine, and appropriate games and activities are provided based on the user's emotional state. For example, when a user says "hello," a character responds with a smile.

[0234] Examples

[0235] Example 1: A user interacts with a friend through text chat. 1. A user logs into the application from a terminal and starts a text chat. He types, "What's the weather like today?" 2. The device sends this text to the server. 3. The server uses a sentiment analysis engine to analyze "What's the weather like today?" and determines the emotional state as "neutral." 4. The server generates weather-related interactive activities as additional content and sends them to the terminal. 5. The terminal displays this information to the user, who can then enjoy further weather information or other activities.

[0236] Example 2: When a user interacts with a virtual character through voice chat 1. The user uses voice chat to talk to a virtual character, asking, "How are you?" 2. The terminal transmits the voice data to the server. 3. The server uses a voice recognition engine to convert the voice data into text data such as "How are you?" 4. The server passes the converted text data to a sentiment analysis engine and determines the emotional state as "positive." 5. The server generates an activity to respond cheerfully based on this positive emotional state. 6. The device displays this information to the user and the virtual character responds with something like "I'm fine! How about you?"

[0237] Examples of prompt statements "If the emotion engine analyzes the user's emotional state and estimates it to be 'happy,' please provide specific ideas for what kind of content (games or activities) should be provided."

[0238] From the above system configuration and specific examples, it can be understood that the emotion recognition and real-time content provision method according to the present invention can be clearly implemented.

[0239] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0240] Step 1: A user logs into the application from a terminal. A username and password are required as input. If the login is successful, the user profile and usage history are output from the server and displayed on the terminal. Specifically, the user enters his / her authentication information on the login screen and clicks the "Login" button.

[0241] Step 2: A user initiates a text or voice chat. As input, the text message or voice data is sent to the server. The server receives this data and prepares it for processing. For example, if a user sends a text message to a friend saying "How are you doing?", the device sends this message to the server.

[0242] Step 3: When the server receives the voice data, it converts the voice data into text data using Amazon (registered trademark) Transcribe. The server receives the voice data as input and generates converted text data as output. Specifically, the voice data of "How are you doing lately?" is converted into the text data of "How are you doing lately?"

[0243] Step 4: The server sends the received or converted text data to the Google Cloud Natural Language API for sentiment analysis. Text data is provided as input, and a sentiment score and a judgment result of the sentiment state are returned as output. For example, the text "How are you doing lately?" is analyzed to detect the sentiment state "neutral."

[0244] Step 5: The server generates appropriate content based on the results of the emotion analysis. It uses the emotional state determination as input and generates games and activities as output. Specifically, relaxing content is selected and generated for a "neutral" emotional state.

[0245] Step 6: The server sends the generated content to the terminal. As input, the server provides the generated content and sends information to the terminal. As output, the terminal displays this content to the user. For example, a puzzle game for relaxation is displayed on the terminal.

[0246] Step 7: The user uses the provided content. User operation data and additional feedback data are generated and sent to the server. Operation information and feedback from the user are received as input, and this is sent to the server as output. Specifically, when the user inputs his / her impressions while solving the puzzle, the data is sent to the server.

[0247] Step 8: The server again inputs the received feedback into the sentiment analysis engine and re-adjusts the content if necessary. Using the feedback data as input, the server provides the user with updated content or a re-analysis of the emotional state as output. For example, based on the feedback "This puzzle is a bit difficult", new content with a slightly lower difficulty level is generated.

[0248] (Application example 2) Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal".

[0249] Conventional applications for exchanging information in virtual environments lacked the ability to grasp the user's emotional state in real time and recommend and play optimal content. This made it difficult to provide content that was in tune with the user's emotions, limiting the user experience. Furthermore, there was no system that integrated educational elements with the provision of content based on emotions. A system that can solve these problems is needed.

[0250] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0251] In this invention, the server includes a means for receiving text chat or voice chat input by a user and presenting the received input to other users to realize exchanges by text chat or voice chat, a means for giving a movement to the user's avatar in response to the user's operation and having it express at least one of a greeting and an emotion, a means for inputting user information to a generation AI and providing the user with at least one of an educational game and an activity based on a response from the generation AI, and a means for analyzing the user's emotions in real time and recommending and playing appropriate entertainment (music or video content) based on the emotions. This makes it possible to provide content that is in line with the user's emotions, thereby providing a richer user experience.

[0252] An "information exchange application" is software that enables a user to communicate with friends and virtual characters within a virtual environment through text chat and voice chat. An "avatar" is a digital character that represents a user within a virtual environment and expresses the user's actions and emotions. "Generative AI" is artificial intelligence that automatically recommends educational games, activities, and content based on the user's input information and emotional state. "Emotion analysis" is a technology that analyzes the content of a user's text chat or voice chat and infers the user's emotional state. "Content recommendation" is a function that suggests optimal music and video content based on the user's emotional state and attribute information. "Educational games and activities" are interactive games and activities designed to learn or enhance knowledge. "Real-time" refers to immediate processing or response in the current time. A "virtual environment" is a computer-generated digital space with which a user interacts.

[0253] The embodiment for implementing the invention includes the following elements.

[0254] server The server has a built-in emotion engine that recognizes the user's emotions. Specifically, the emotion engine analyzes and estimates the user's emotional state from the contents of text and voice chat. The server utilizes a generative AI model to recommend and play optimal music and video content based on the user's emotional state. The server also adjusts the difficulty and content of educational games and activities according to the user's emotional state. The system includes the following software: emotion_engine(emotion engine) content_recommender (content recommendation system) text_analysis audio_recognition

[0255] Terminal The terminal is the device through which the user logs into the application and performs operations and communication within the virtual environment. The terminal receives the user's voice and text input and transmits it to the server. It also provides the user with content and activities based on the emotional state estimated by the emotion engine. Terminals include smartphones and smart glasses. The terminal has the following functions: Function to give movement to avatars Text and voice chat functionality Content playback function

[0256] User Users log into the application and interact with friends and virtual characters in a virtual environment. They communicate through text and voice chat, and enjoy music and video content based on their emotional state estimated by the emotion engine. They also participate in educational games and activities, experiencing adjustments in difficulty and content according to their emotional state. An example of a specific user scenario is as follows:

[0257] 1. The user launches the app and speaks, "I'm feeling a bit sad today." 2. The app analyzes the voice and the emotion engine determines that the user is feeling "sad." 3. The content recommendation system recommends soothing music and inspiring videos that suit the emotion of "sad." 4. The recommended content will be played automatically.

[0258] Examples of prompt statements User input: "I'm feeling a bit sad today." Analyzed emotion: "Sadness" Suitable content: "Healing music" "Inspirational videos" Suggested prompt: "If the user is feeling sad, please recommend soothing music that will soothe the soul. Also, please recommend an inspiring video."

[0259] In this way, content that is in tune with the user's emotions can be provided, providing a rich user experience.

[0260] The flow of the specific process in the application example 2 will be described with reference to FIG.

[0261] Step 1: The device receives the user's text chat or voice chat input. When the user speaks "I'm feeling a bit sad today," the device receives this voice data. The input data is voice data. The device then uses voice recognition software to convert this voice data into text data. The converted text data is output.

[0262] Step 2: The terminal transmits the converted text data to the server. The input here is text data, and the output to the server is also text data. Specifically, the terminal transmits the text data to the server via the network.

[0263] Step 3: The server inputs the received text data into an emotion engine. The emotion engine analyzes the user's emotional state from the text data. The input here is the text data, and the emotion analysis engine estimates the emotional state. The estimated emotional state, e.g., "sadness," is output.

[0264] Step 4: The server inputs the estimated emotional state into a content recommendation system, which recommends optimal music and video content based on the emotional state. The input is the emotional state, and the content recommendation system uses this data to select recommended content. The selected recommended content is output.

[0265] Step 5: The server transmits the recommended content to the terminal. The input is the recommended content data, and the output is also the recommended content data. The server transmits this to the terminal via the network.

[0266] Step 6: The terminal presents the received content to the user. Specifically, it uses the terminal's media player function to play music and videos. The input is the recommended content data, and the output is the content to be played that is presented to the user. The terminal plays the content using a screen display and speakers.

[0267] In this way, a system is realized that provides content in real time according to the user's emotions.

[0268] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires a voice indicating a user input for the result of the specific processing. The control unit 46A transmits the voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0269] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) Internet Search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0270] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0271] [Third embodiment] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0272] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0273] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).

[0274] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0275] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs the voice according to instructions from the processor 46.

[0276] Camera 42 is a small digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (e.g., an imaging range defined by an angle of view equivalent to the width of the field of vision of an average healthy person).

[0277] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54. The transmission and reception of various types of information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is performed in a secure state.

[0278] Fig. 6 shows an example of main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0279] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32, and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0280] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0281] In the headset type terminal 314, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0282] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server", and the headset type terminal 314 will be referred to as the "terminal".

[0283] An embodiment for implementing the present invention includes the following elements.

[0284] 1. Server: The server receives the user's login information and performs authentication. It also manages the user's avatar information and data required for game activities, and provides appropriate content in response to the user's requests. It also processes input to the generative AI and information shared with other users.

[0285] 2. Terminal: The terminal is where the user logs into the application and operates and communicates within the virtual environment. It transmits the user's text and voice chat input to the server, receives and displays the generative AI's responses, and transmits the user's avatar movements and game activity selections to the server.

[0286] 3. User: Users log into the application and interact with friends and virtual characters in a virtual environment, communicating through text and voice chat, animating their avatars to greet others, express emotions, and select educational games and activities to share with other users.

[0287] 4. Generative AI: A generative AI is a generative AI that generates responses based on user input. It receives information from the user's text chat and avatar movements, as well as information about the choices and suggestions of other users and AI characters, and generates appropriate responses. Generative AI runs on the server and processes input and output from the server.

[0288] All of these elements work together to allow users to interact with friends and virtual characters in a virtual environment and enjoy educational games and activities. The server manages user information and provides appropriate content. The device accepts user operations, communicates with the server, and displays the generative AI's responses. The generative AI generates responses based on the user's input and sends them to the device via the server.

[0289] The process flow will be explained below. Step 1: User logs into the application Step 2: The device sends the user's login information to the server Step 3: The server receives the user's login information and performs authentication. Step 4: The server sends the authentication result to the terminal. Step 5: The user selects an avatar to enter the virtual environment. Step 6: The device sends the user's avatar information to the server. Step 7: The server receives the user's avatar information and displays the avatar in the virtual environment. Step 8: The user interacts with friends and virtual characters through text and voice chat. Step 9: The device sends the user's chat input to the server Step 10: The server receives the user's chat input and passes it to the generative AI. Step 11: Generative AI generates a response based on the user's chat input Step 12: The server receives the response from the generative AI and sends it to the device. Step 13: The device displays the generative AI's response Step 14: The user gives the avatar movement to greet or express emotion. Step 15: The device sends the user's avatar movements to the server. Step 16: The server receives the user's avatar movements and moves the avatar in the virtual environment. Step 17: User selects educational games and activities Step 18: The device sends the user's selection to the server Step 19: The server serves games and activities based on the user's selection. Step 20: User experiences the game or activity Step 21: The user plays a game or activity with other users or AI characters. Step 22: The device sends the choices and suggestions of other users and AI characters to the server. Step 23: The server uses the choices and suggestions of other users and AI characters as inputs to the generative AI. Step 24: The generative AI generates a response based on the choices and suggestions of other users and the AI ​​character. Step 25: The server receives the response from the generative AI and sends it to the terminal. Step 26: The device displays the generative AI's response. Step 27: The user shares the results of the game or activity with other users. Step 28: The device sends the user's sharing contacts and chat contents to the server. Step 29: The server uses the user's shared contacts and chat contents as input to the generative AI. Step 30: Generative AI generates responses based on who the user is sharing with and what the chat is about. Step 31: The server receives the response from the generative AI and sends it to the terminal. Step 32: The device displays the generative AI's response. Step 33: Use information from other users who are registered as friends by the user Step 34: User sends invitation notification and points to friends Step 35: The device sends the user's friend registration information to the server. Step 36: The server receives the user's friend registration information and uses it as input for the generative AI. Step 37: The generative AI generates a response based on the user's friend registration information. Step 38: The server receives the response from the generative AI and sends it to the terminal. Step 39: The device displays the generative AI's response. Step 40: The user logs out and exits the application

[0290] Example 1 Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal".

[0291] Conventional information sharing applications have limited means for users to effectively interact in a virtual environment, making it difficult to obtain satisfactory results, especially in real-time communication and interactive experiences using generative AI. For example, there is a lack of systems that reflect user operations and chat content in real time and provide consistent responses. In addition, the quality of responses and accuracy of suggestions by generative AI have also been issues. In order to solve these problems, we provide the present invention.

[0292] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0293] In this invention, the server includes a means for authenticating the user's login information, a means for acquiring and managing the user's avatar information and game activity information, a means for receiving the user's text chat and voice chat input and sharing it with other users, and a means for generating responses based on the user's input using a generative AI model and sending the responses to the terminal. This allows the user to effectively interact with other users and virtual characters in real time within the virtual environment. Furthermore, the generative AI provides high-quality responses, improving the user's experience.

[0294] "User" refers to an individual who utilizes the system to interact and perform activities within the virtual environment. A "server" refers to a computing device that authenticates user login information, accesses a database to manage information, and responds to requests from users. "Terminal" refers to a device that allows a user to access the system and perform operations and send and receive information. Examples include smartphones and personal computers. "Login Information" refers to authentication information, such as a username and password, that a User enters to be authenticated to a System. "Authentication" refers to the process of verifying that a user can legitimately use a system based on the login information entered by the user. "Avatar" refers to a virtual character that a user uses to represent themselves within a virtual environment. "Game Activity Information" refers to data regarding activity within games and other virtual environments in which a User participates. "Text chat" refers to a means by which a user communicates with other users by inputting text. "Voice chat" refers to a means by which users communicate with other users using voice. A "generative AI model" refers to an artificial intelligence model that generates responses or suggestions based on user input. A "prompt sentence" refers to an input sentence that is provided to a generative AI model to generate some kind of response. "Feedback" refers to the evaluation or opinion given by the user regarding the response of the system or the generative AI model.

[0295] This invention is an information sharing application system that allows users to smoothly interact with other users and virtual characters in a virtual environment. This system is realized by the cooperation and operation of a server, a terminal, a user, and a generative AI model.

[0296] System Overview

[0297] server The server receives the user's login information and performs authentication. For example, an encryption library such as OpenSSL (registered trademark) is used for authentication. The server also stores the user's avatar information and past game activity information in a database (e.g., MySQL (registered trademark)) and manages this information. When the user sends a request, the server retrieves the corresponding data from the database and provides it to the user's terminal.

[0298] In addition, the server receives user input information and forwards it to the generative AI. Generative AI is usually implemented using machine learning libraries such as TensorFlow (registered trademark) and PyTorch (registered trademark). The server also manages shared information with other users. For example, if a user is chatting with a friend, the server receives the message and forwards it to the appropriate person.

[0299] Terminal The terminal provides an interface for users to access and log in to the application. It is responsible for sending text chat and voice chat information entered by the user to the server. It also transmits the movements of the avatar selected by the user and game activity to the server, enabling real-time operation.

[0300] The device receives the response from the server and displays it to the user. In particular, the response from the generative AI includes replies to the user's chat and suggestions for avatar movements. For example, in the case of a web application, this process realizes real-time communication on the browser using JavaScript (registered trademark) or WebSocket.

[0301] User Users log into the application using a terminal and interact with other users and virtual characters in the virtual environment. Users can communicate through text and voice chat, and can give their avatars animations to express greetings and emotions.

[0302] In addition, users can select educational games and activities and share them with other users. For example, if a user selects "Math Quiz," quiz questions will be provided by the generative AI, allowing users to compete or collaborate with other users to advance their learning.

[0303] Generative AI Model The generative AI model runs on a server and is responsible for generating responses based on user input. It receives information about the user's text chat, avatar movements, and other user and AI character choices and suggestions, and generates answers accordingly.

[0304] For example, Transformer-based models are often used as generative AI models, and when you input a prompt like "Please suggest that the user start a quiz," the AI ​​generates a response like "Let's ask, How about starting a quiz?"

[0305] Examples Examples of prompt statements 1. If the user wants their avatar to say "Hello!": Prompt: "Have the user say 'Hello!' to their avatar." Generative AI response: "Hello!"

[0306] 2. If the user wants to participate in an educational quiz: Prompt: "The user wants to participate in an educational math quiz." Generative AI response: "Great choice! Now, first question: What is 5 + 7?"

[0307] These elements work together to create a system that allows users to enjoy a wide variety of experiences within a virtual environment.

[0308] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0309] Program processing flow Step 1: User login process explanation User: Launches the application and enters username and password on the login screen. Terminal: Receives the user's login information and sends it to the server. Server: The received login information is compared against a database (e.g. MySQL (registered trademark)) and authentication is performed. Input is the user name and password, and output is the authentication result (success / failure).

[0310] operation The user enters "john_doe" and "password123" into the form and clicks the login button. The terminal encrypts this information and sends it to the server. The server checks the password against the records in its database, and if authentication is successful, it sends a "login successful" message to the terminal.

[0311] Step 2: Get your avatar and game activity information explanation Server: After the user successfully logs in, retrieve the avatar information and past game activity information from the database. Input is the user ID, and output is the retrieved avatar information and activity history. Terminal: displays information received from the server. Input is avatar information and activity information from the server, and output is the display of that information.

[0312] operation The server runs a database query to retrieve avatar information and past game history for user "john_doe". The device displays the acquired avatar information (e.g., "character wearing a blue hat") and past game history (e.g., "played math quiz on October 1st") on the screen.

[0313] Step 3: User interaction and communication explanation User: Controls an avatar and communicates with other users via text chat and voice chat. Terminal: Sends text and voice data entered by the user to the server. Input is text chat and voice data, and output is the result sent to the server. Server: Shares the received data with other users. Input is the received text or voice data, and output is the shared data.

[0314] operation A user types "Hello everyone!" in text chat and presses send. The terminal sends the text to the server. The server sends a notification to other logged-in users saying "john_doe: Hello everyone!"

[0315] Step 4: Generate prompts for generative AI and obtain responses explanation Server: Sends text chat content and avatar movement information to the generative AI. Input is text chat and avatar movement information, output is prompt text and generated responses. Generative AI: Processes the received prompt and generates an appropriate response. The input is the prompt and the output is the generated response. Server: Receives the response from the generative AI, performs further processing if necessary, and sends the response to the terminal. The input is the response from the generative AI, and the output is a notification to the terminal.

[0316] operation The server sends a prompt to the generative AI saying, "The user said 'Hello, everyone!' Please generate an appropriate response." The generative AI generates the response, "Hello john_doe! How are you?" The server receives the response from the generative AI and sends it to the terminal. The device will display the generative AI's response on the chat screen.

[0317] Step 5: Displaying the response and feeding it back to the server explanation Terminal: Receives the generative AI's response and displays it to the user. At the same time, sends the user's reaction and evaluation to the server. The input is the response from the generative AI, and the output is the display and feedback data for the user. User: Checks the generative AI's response and provides feedback if necessary. The input is the response from the generative AI, and the output is evaluation and feedback. Server: Receives user feedback and uses it to improve the generative AI model. The input is feedback data, and the output is feedback to the generative AI model.

[0318] operation The device will display the message "Hello john_doe! How are you?" on the chat screen. The user rates the response as "very good" and sends the feedback to the server. The server stores the feedback and later uses it as training data for the generative AI.

[0319] (Application example 1) Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0320] Conventional educational entertainment and activities in virtual environments have lacked user interaction and personalization, resulting in low learning effectiveness and satisfaction. In addition, there were also insufficient methods for effectively utilizing responses generated in real time and mechanisms for deepening interactions between users.

[0321] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0322] In this invention, the server includes a means for enabling a user to communicate through text messaging or voice messaging, a means for giving a movement to a user's avatar to greet or express emotions, a means for providing educational entertainment or activities using a generation AI, a means for presenting educational entertainment or activities to a user and displaying the results through a visual output device or an audio output device, a means for acquiring a response by the generation AI in real time and presenting the response on a user interface, a means for providing educational entertainment or activities based on an answer obtained by inputting user attribute information and learning information into the generation AI, a data management means for analyzing a user's behavioral data and dialogue history and generating a more personalized response, a means for using the selection or suggestion of other users or AI agents based on the answer from the generation AI, a means for sharing the results of educational entertainment or activities between users and making suggestions to the user based on the answer from the generation AI, a sentence generation means for generating a prompt sentence and inputting it into the generation AI, and an evaluation means for evaluating and improving the response generated based on the prompt sentence. This allows a user to learn while enjoying personalized educational entertainment or activities, and to deepen interactions with other users and AI agents.

[0323] A "virtual environment" is a digital space created using computer technology in which users can have new experiences that are different from the real world. An "information sharing application" is software that enables multiple users to exchange and share information, including data, messages, media, etc., over the Internet. "Text messaging" means a means by which users communicate in real time by sending messages in text form. "Voice messaging" is a means by which users can send messages in the form of voice and communicate in real time. An "avatar" is a digital character that represents a user within a virtual space and is a means for expressing the user's movements and emotions. "Generative AI" is an artificial intelligence technique that generates text, images, etc. based on user input. "Educational entertainment" is games and activities that are designed to entertain users while also providing educational elements. A "visual output device" is a device, such as a digital display or a head-mounted display, that a user uses to obtain visual information. An "audio output device" is a device such as a speaker or headphones that a user uses to obtain audio information. A "user interface" is the arrangement and design of screens and input devices that allow a user to interact with and operate an application. "Attribute information" is data that indicates individual characteristics of a user, such as age, sex, and interests. "Learning information" is data related to the user's past learning history and knowledge level. "Behavioral data" is a record of what operations a user performed within the system. An "interaction history" is a record of messages and responses that have been exchanged between a user and a system in the past. An "AI agent" is a character that uses artificial intelligence and is designed to interact with users in a virtual space. A "prompt" is text that contains instructions or questions that are entered into the generation AI. A "sentence generation means" is a means for creating a prompt sentence to be supplied to the generation AI. An "evaluation means" is a means for evaluating the responses created by the generative AI to find areas for improvement.

[0324] The present invention provides a system for users to experience educational entertainment and activities in a virtual environment and to interact with other users and AI agents in real time. Specific embodiments for implementing the present invention are described below.

[0325] System Configuration The system mainly consists of three elements: the server, the terminal, and the user. The role of each element is explained below.

[0326] 1. Server Authentication means: The server receives the user's login information and performs authentication. For this purpose, a database (DB) system such as MySQL (registered trademark) or PostgreSQL (registered trademark) is used. Data management means: The server manages user avatar information, game activity data, and other user information. This is done using a cloud storage service (e.g., AWS (registered trademark) S3). Generative AI: The generative AI engine generates responses in real time based on user input. ChatGPT (registered trademark) and the like are used as generative AI engines. Communication method: The server uses WebSocket or HTTP protocols to manage communication with the user terminal.

[0327] 2. Terminal User input acceptance means: The terminal accepts the user's text and voice input and sends it to the server, using voice recognition software (e.g., Google® Speech-to-Text API) and a text input interface. Display means: The device displays the generated AI's responses and the results of game activities. This can be done using smart glasses or a head-mounted display. Movement control means: The terminal controls the movement of the avatar based on the user's operation. This is achieved by using motion capture technology.

[0328] 3. Users Input means: Users interact with other users and AI agents in the virtual environment through voice and text input. Avatar Control: Users control their own avatars to express actions and emotions. Activity Participation: Users participate in educational, entertainment, and activities.

[0329] Example of processing flow A specific example of the processing flow of the entire system is shown below.

[0330] 1. User Login The user enters login information from the terminal and transmits it to the server. The server uses the authentication means to authenticate the user.

[0331] 2. Activity selection The user selects educational entertainment and activities through the terminal's interface. The server receives the user's selection and provides the appropriate data.

[0332] 3. Generative AI response generation The user enters a question via text or voice. The server sends prompts to the generation AI, which generates responses in real time. Example response: "The user said, 'Do you want to proceed to the next question?' Generate an appropriate response."

[0333] 4. Display in the interface The server sends the generated response to the terminal, which presents it to the user through its visual and / or audio output devices.

[0334] In this manner, the present invention allows users to enjoy interactive educational entertainment and activities within a virtual environment.

[0335] The flow of the specific process in the application example 1 will be described with reference to FIG.

[0336] Step 1: The server receives the user's login information and performs authentication. The user ID and password entered by the user from the terminal are sent to the server. The server retrieves the corresponding user information from the database and checks whether it matches. The input is the user's login information, and the output is the authentication result (success or failure).

[0337] Step 2: A user selects an educational entertainment or activity using the interface of the terminal. The terminal transmits information about the activity selected by the user to the server. The server retrieves data corresponding to the selected activity from a database and transmits it to the terminal. The input is the information about the activity selected by the user, and the output is the data corresponding to the activity.

[0338] Step 3: The user inputs a question via text or voice. The text or voice data entered by the user into the terminal is sent to the server. The server converts the voice data into text (voice recognition) and sends it to the generation AI as a prompt. The input is the user's question (text, voice), and the output is the prompt.

[0339] Step 4: The server sends a prompt to the generation AI, which generates a response. The generation AI generates a text response based on the prompt and returns it to the server. The server receives the response and adjusts the response content, taking into account the user's attribute information and behavioral data. The input is the prompt and the user's attribute information, and the output is the generated response.

[0340] Step 5: The server sends the generated response to the terminal, which presents it to the user. The terminal displays the response message on a visual output device and optionally plays it through an audio output device. The input is the generated response and the output is the response presented in visual and audio output.

[0341] Step 6: The user confirms the generated response and decides on the next action: the user enters a further question or selects the next activity. The inputs are the user's confirmation and the next action, and the output is the information to proceed to the next step.

[0342] Step 7: The server provides an interaction mechanism for users to share the results of their educational activities. Users perform operations on their terminals to share their results, and the server manages the data. The input is the user's shared data, and the output is the shared information provided to other users.

[0343] The above is a specific processing flow of the program of the system that realizes the application example. This processing flow allows users to enjoy educational entertainment and activities in real time.

[0344] In addition, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0345] An embodiment for implementing the present invention includes the following elements.

[0346] 1. Server: The server realizes a system that combines an emotion engine that recognizes the user's emotions. The server analyzes the contents of the user's text chat and voice chat and the characteristics of the voice, and estimates the user's emotional state. The server also adjusts the difficulty and content of games and activities based on the user's emotional state estimated by the emotion engine.

[0347] 2. Terminal: The terminal is where the user logs into the application and operates and communicates within the virtual environment. The terminal transmits the user's text and voice chat input to the server, and the server provides games and activities based on the user's emotional state estimated by the emotion engine. The terminal also controls the user's avatar movement and interaction with other users.

[0348] 3. User: Users log into the application and interact with friends and virtual characters in a virtual environment. They communicate through text and voice chat, experience games and activities based on their emotional state estimated by an emotion engine, and interact with other users by giving their avatars animations to greet others and express emotions.

[0349] The above elements work together to realize a system that recognizes the user's emotions and provides games and activities accordingly. The server uses the emotion engine to estimate the user's emotional state and provide appropriate content. The terminal accepts the user's operations, communicates with the server, and displays responses via the emotion engine. The user can interact with friends and virtual characters in the virtual environment and enjoy content that matches their emotions.

[0350] The process flow will be explained below. Step 1: User logs into the application Step 2: The device sends the user's login information to the server Step 3: The server receives the user's login information and performs authentication. Step 4: The server sends the authentication result to the terminal. Step 5: The user selects an avatar to enter the virtual environment. Step 6: The device sends the user's avatar information to the server. Step 7: The server receives the user's avatar information and displays the avatar in the virtual environment. Step 8: The user interacts with friends and virtual characters through text and voice chat. Step 9: The device sends the user's chat input to the server Step 10: The server receives the user's chat input and passes it as input to the emotion engine. Step 11: The emotion engine analyzes the user's chat input and infers their emotional state Step 12: The server receives the results from the emotion engine and selects the appropriate game or activity. Step 13: The server sends the selected game or activity to the device. Step 14: The device displays the selected game or activity for the user to experience. Step 15: The user gives the avatar movement to greet or express emotion. Step 16: The terminal transmits the user's avatar movements to the server. Step 17: The server receives the user's avatar movements and moves the avatar in the virtual environment. Step 18: The user plays a game or activity with other users or AI characters. Step 19: The device sends the choices and suggestions of other users and AI characters to the server. Step 20: The server uses the choices and suggestions of other users and AI characters as inputs to the emotion engine. Step 21: The emotion engine generates responses based on the choices and suggestions of other users and AI characters. Step 22: The server receives the emotion engine's response and sends it to the terminal. Step 23: The device displays the emotion engine's response. Step 24: The user shares the results of the game or activity with other users. Step 25: The device sends the user's sharing contacts and chat contents to the server. Step 26: The server uses the user's shared contacts and chat contents as input to the emotion engine. Step 27: The emotion engine generates a response based on who the user is sharing with and what the chat is about. Step 28: The server receives the emotion engine's response and sends it to the terminal. Step 29: The device displays the emotion engine's response. Step 30: Use information about other users who are registered as friends by the user Step 31: User sends invitation notification and points to friends Step 32: The device sends the user's friend registration information to the server. Step 33: The server receives the user's friend registration information and uses it as input to the emotion engine. Step 34: The emotion engine generates a response based on the user's friend information. Step 35: The server receives the emotion engine's response and sends it to the terminal. Step 36: The device displays the emotion engine's response. Step 37: The user logs out and exits the application

[0351] Example 2 Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal".

[0352] Conventional communication systems and information sharing applications in virtual environments have the problem that it is difficult to adjust content in real time according to the user's emotional state. In addition, there is a lack of systems that can properly analyze the user's emotional state and provide games and activities based on it. This results in a uniform user experience and a lack of personalized support tailored to each individual's emotional state.

[0353] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for estimating the emotional state of the user by utilizing an emotion analysis engine, a means for adjusting the content of the game or activity based on the estimated emotional state, and a means for converting voice data into text data. This makes it possible to adjust the content in real time according to the emotional state of the user, and to provide a personalized experience.

[0354] A "virtual environment" is a computer-generated digital space that a user can interact with. A "communication application" is software that enables users to interact with other users and virtual characters through text and / or voice chat. "Text chat" is a means by which users exchange messages in real time through text input. "Voice chat" is a means by which users communicate in real time through voice input. "Character" refers to an avatar controlled by a user within a virtual environment, or a virtual entity controlled by a computer. An "emotion analysis engine" is an algorithm or software used to infer a user's emotional state from text or voice input. An "emotional state" refers to the psychological state a user feels at a particular point in time and can be classified as positive, negative, neutral, etc. "Games and Activities" are interactive entertainment or learning activities that users enjoy within a virtual environment. "Means for converting voice data into text data" refers to technology or software used to convert a user's voice input into text form. "Personalization" means optimizing systems and content to suit the characteristics and circumstances of each individual user. "Real-time" refers to responding immediately to user input and changes in the environment and processing without delay.

[0355] The following system configuration will be described as an embodiment of the present invention. The main components are a server, a terminal, and a user's operation method.

[0356] The server runs on a cloud service and is equipped with multiple function engines, including a sentiment analysis engine and a conversion engine. Specifically, the server uses a Natural Language Processing (NLP) API for sentiment analysis and a speech recognition API for converting voice data into text data. This makes it possible to analyze the user's emotional state and generate and provide interactive content in real time according to that state.

[0357] A terminal is a device that provides an interface for users to access applications and perform activities and communications within the virtual environment. Examples of terminals include PCs, smartphones, and tablets. The terminal has the function of sending the user's text chat and voice chat input to the server. The terminal also displays content to the user based on the emotional state received from the server. Furthermore, the terminal accepts user operations in real time and performs the actions and emotional expressions of the virtual character.

[0358] Users log into the application through their terminals and interact with other users and virtual characters in the virtual environment. The user's text and voice chat inputs are analyzed by an emotion analysis engine, and appropriate games and activities are provided based on the user's emotional state. For example, when a user says "hello," a character responds with a smile.

[0359] Examples

[0360] Example 1: A user interacts with a friend through text chat. 1. A user logs into the application from a terminal and starts a text chat. He types, "What's the weather like today?" 2. The device sends this text to the server. 3. The server uses a sentiment analysis engine to analyze "What's the weather like today?" and determines the emotional state as "neutral." 4. The server generates weather-related interactive activities as additional content and sends them to the terminal. 5. The terminal displays this information to the user, who can then enjoy further weather information or other activities.

[0361] Example 2: When a user interacts with a virtual character through voice chat 1. The user uses voice chat to talk to a virtual character, asking, "How are you?" 2. The terminal transmits the voice data to the server. 3. The server uses a voice recognition engine to convert the voice data into text data such as "How are you?" 4. The server passes the converted text data to a sentiment analysis engine and determines the emotional state as "positive." 5. The server generates an activity to respond cheerfully based on this positive emotional state. 6. The device displays this information to the user and the virtual character responds with something like "I'm fine! How about you?"

[0362] Examples of prompt statements "If the emotion engine analyzes the user's emotional state and estimates it to be 'happy,' please provide specific ideas for what kind of content (games or activities) should be provided."

[0363] From the above system configuration and specific examples, it can be understood that the emotion recognition and real-time content provision method according to the present invention can be clearly implemented.

[0364] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0365] Step 1: A user logs into the application from a terminal. A username and password are required as input. If the login is successful, the user profile and usage history are output from the server and displayed on the terminal. Specifically, the user enters his / her authentication information on the login screen and clicks the "Login" button.

[0366] Step 2: A user initiates a text or voice chat. As input, the text message or voice data is sent to the server. The server receives this data and prepares it for processing. For example, if a user sends a text message to a friend saying "How are you doing?", the device sends this message to the server.

[0367] Step 3: When the server receives the voice data, it converts the voice data into text data using Amazon (registered trademark) Transcribe. The server receives the voice data as input and generates converted text data as output. Specifically, the voice data of "How are you doing lately?" is converted into the text data of "How are you doing lately?"

[0368] Step 4: The server sends the received or converted text data to the Google Cloud Natural Language API for sentiment analysis. Text data is provided as input, and a sentiment score and a judgment result of the sentiment state are returned as output. For example, the text "How are you doing lately?" is analyzed to detect the sentiment state "neutral."

[0369] Step 5: The server generates appropriate content based on the results of the emotion analysis. It uses the emotional state determination as input and generates games and activities as output. Specifically, relaxing content is selected and generated for a "neutral" emotional state.

[0370] Step 6: The server sends the generated content to the terminal. As input, the server provides the generated content and sends information to the terminal. As output, the terminal displays this content to the user. For example, a puzzle game for relaxation is displayed on the terminal.

[0371] Step 7: The user uses the provided content. User operation data and additional feedback data are generated and sent to the server. Operation information and feedback from the user are received as input, and this is sent to the server as output. Specifically, when the user inputs his / her impressions while solving the puzzle, the data is sent to the server.

[0372] Step 8: The server again inputs the received feedback into the sentiment analysis engine and re-adjusts the content if necessary. Using the feedback data as input, the server provides the user with updated content or a re-analysis of the emotional state as output. For example, based on the feedback "This puzzle is a bit difficult", new content with a slightly lower difficulty level is generated.

[0373] (Application example 2) Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server", and the headset type terminal 314 will be referred to as a "terminal".

[0374] Conventional applications for exchanging information in virtual environments lacked the ability to grasp the user's emotional state in real time and recommend and play optimal content. This made it difficult to provide content that was in tune with the user's emotions, limiting the user experience. Furthermore, there was no system that integrated educational elements with the provision of content based on emotions. A system that can solve these problems is needed.

[0375] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0376] In this invention, the server includes a means for receiving text chat or voice chat input by a user and presenting the received input to other users to realize exchanges by text chat or voice chat, a means for giving a movement to the user's avatar in response to the user's operation and having it express at least one of a greeting and an emotion, a means for inputting user information to a generation AI and providing the user with at least one of an educational game and an activity based on a response from the generation AI, and a means for analyzing the user's emotions in real time and recommending and playing appropriate entertainment (music or video content) based on the emotions. This makes it possible to provide content that is in line with the user's emotions, thereby providing a richer user experience.

[0377] An "information exchange application" is software that enables a user to communicate with friends and virtual characters within a virtual environment through text chat and voice chat. An "avatar" is a digital character that represents a user within a virtual environment and expresses the user's actions and emotions. "Generative AI" is artificial intelligence that automatically recommends educational games, activities, and content based on the user's input information and emotional state. "Emotion analysis" is a technology that analyzes the content of a user's text chat or voice chat and infers the user's emotional state. "Content recommendation" is a function that suggests optimal music and video content based on the user's emotional state and attribute information. "Educational games and activities" are interactive games and activities designed to learn or enhance knowledge. "Real-time" refers to immediate processing or response in the current time. A "virtual environment" is a computer-generated digital space with which a user interacts.

[0378] The embodiment for implementing the invention includes the following elements.

[0379] server The server has a built-in emotion engine that recognizes the user's emotions. Specifically, the emotion engine analyzes and estimates the user's emotional state from the contents of text and voice chat. The server utilizes a generative AI model to recommend and play optimal music and video content based on the user's emotional state. The server also adjusts the difficulty and content of educational games and activities according to the user's emotional state. The system includes the following software: emotion_engine(emotion engine) content_recommender (content recommendation system) text_analysis audio_recognition

[0380] Terminal The terminal is the device through which the user logs into the application and performs operations and communication within the virtual environment. The terminal receives the user's voice and text input and transmits it to the server. It also provides the user with content and activities based on the emotional state estimated by the emotion engine. Terminals include smartphones and smart glasses. The terminal has the following functions: Function to give movement to avatars Text and voice chat functionality Content playback function

[0381] User Users log into the application and interact with friends and virtual characters in a virtual environment. They communicate through text and voice chat, and enjoy music and video content based on their emotional state estimated by the emotion engine. They also participate in educational games and activities, experiencing adjustments in difficulty and content according to their emotional state. An example of a specific user scenario is as follows:

[0382] 1. The user launches the app and speaks, "I'm feeling a bit sad today." 2. The app analyzes the voice and the emotion engine determines that the user is feeling "sad." 3. The content recommendation system recommends soothing music and inspiring videos that suit the emotion of "sad." 4. The recommended content will be played automatically.

[0383] Examples of prompt statements User input: "I'm feeling a bit sad today." Analyzed emotion: "Sadness" Suitable content: "Healing music" "Inspirational videos" Suggested prompt: "If the user is feeling sad, please recommend soothing music that will soothe the soul. Also, please recommend an inspiring video."

[0384] In this way, content that is in tune with the user's emotions can be provided, providing a rich user experience.

[0385] The flow of the specific process in the application example 2 will be described with reference to FIG.

[0386] Step 1: The device receives the user's text chat or voice chat input. When the user speaks "I'm feeling a bit sad today," the device receives this voice data. The input data is voice data. The device then uses voice recognition software to convert this voice data into text data. The converted text data is output.

[0387] Step 2: The terminal transmits the converted text data to the server. The input here is text data, and the output to the server is also text data. Specifically, the terminal transmits the text data to the server via the network.

[0388] Step 3: The server inputs the received text data into an emotion engine. The emotion engine analyzes the user's emotional state from the text data. The input here is the text data, and the emotion analysis engine estimates the emotional state. The estimated emotional state, e.g., "sadness," is output.

[0389] Step 4: The server inputs the estimated emotional state into a content recommendation system, which recommends optimal music and video content based on the emotional state. The input is the emotional state, and the content recommendation system uses this data to select recommended content. The selected recommended content is output.

[0390] Step 5: The server transmits the recommended content to the terminal. The input is the recommended content data, and the output is also the recommended content data. The server transmits this to the terminal via the network.

[0391] Step 6: The terminal presents the received content to the user. Specifically, it uses the terminal's media player function to play music and videos. The input is the recommended content data, and the output is the content to be played that is presented to the user. The terminal plays the content using a screen display and speakers.

[0392] In this way, a system is realized that provides content in real time according to the user's emotions.

[0393] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input for the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0394] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0395] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[0396] [Fourth embodiment]

[0397] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[0398] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0399] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).

[0400] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. In addition, the microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[0401] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs the voice according to instructions from the processor 46.

[0402] Camera 42 is a small digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (e.g., an imaging range defined by an angle of view equivalent to the width of the field of vision of an average healthy person).

[0403] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54. The transmission and reception of various types of information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is performed in a secure state.

[0404] The control target 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, legs, etc. The posture and behavior of the robot 414 are controlled by controlling the motors of the arms, hands, legs, etc. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[0405] Fig. 8 shows an example of main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0406] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32, and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0407] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0408] In the robot 414, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0409] Next, a description will be given of the specific processing by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal".

[0410] An embodiment for implementing the present invention includes the following elements.

[0411] 1. Server: The server receives the user's login information and performs authentication. It also manages the user's avatar information and data required for game activities, and provides appropriate content in response to the user's requests. It also processes input to the generative AI and information shared with other users.

[0412] 2. Terminal: The terminal is where the user logs into the application and operates and communicates within the virtual environment. It transmits the user's text and voice chat input to the server, receives and displays the generative AI's responses, and transmits the user's avatar movements and game activity selections to the server.

[0413] 3. User: Users log into the application and interact with friends and virtual characters in a virtual environment, communicating through text and voice chat, animating their avatars to greet others, express emotions, and select educational games and activities to share with other users.

[0414] 4. Generative AI: A generative AI is a generative AI that generates responses based on user input. It receives information from the user's text chat and avatar movements, as well as information about the choices and suggestions of other users and AI characters, and generates appropriate responses. Generative AI runs on the server and processes input and output from the server.

[0415] All of these elements work together to allow users to interact with friends and virtual characters in a virtual environment and enjoy educational games and activities. The server manages user information and provides appropriate content. The device accepts user operations, communicates with the server, and displays the generative AI's responses. The generative AI generates responses based on the user's input and sends them to the device via the server.

[0416] The process flow will be explained below. Step 1: User logs into the application Step 2: The device sends the user's login information to the server Step 3: The server receives the user's login information and performs authentication. Step 4: The server sends the authentication result to the terminal. Step 5: The user selects an avatar to enter the virtual environment. Step 6: The device sends the user's avatar information to the server. Step 7: The server receives the user's avatar information and displays the avatar in the virtual environment. Step 8: The user interacts with friends and virtual characters through text and voice chat. Step 9: The device sends the user's chat input to the server Step 10: The server receives the user's chat input and passes it to the generative AI. Step 11: Generative AI generates a response based on the user's chat input Step 12: The server receives the response from the generative AI and sends it to the device. Step 13: The device displays the generative AI's response Step 14: The user gives the avatar movement to greet or express emotion. Step 15: The device sends the user's avatar movements to the server. Step 16: The server receives the user's avatar movements and moves the avatar in the virtual environment. Step 17: User selects educational games and activities Step 18: The device sends the user's selection to the server Step 19: The server serves games and activities based on the user's selection. Step 20: User experiences the game or activity Step 21: The user plays a game or activity with other users or AI characters. Step 22: The device sends the choices and suggestions of other users and AI characters to the server. Step 23: The server uses the choices and suggestions of other users and AI characters as inputs to the generative AI. Step 24: The generative AI generates a response based on the choices and suggestions of other users and the AI ​​character. Step 25: The server receives the response from the generative AI and sends it to the terminal. Step 26: The device displays the generative AI's response. Step 27: The user shares the results of the game or activity with other users. Step 28: The device sends the user's sharing contacts and chat contents to the server. Step 29: The server uses the user's shared contacts and chat contents as input to the generative AI. Step 30: Generative AI generates responses based on who the user is sharing with and what the chat is about. Step 31: The server receives the response from the generative AI and sends it to the terminal. Step 32: The device displays the generative AI's response. Step 33: Use information from other users who are registered as friends by the user Step 34: User sends invitation notification and points to friends Step 35: The device sends the user's friend registration information to the server. Step 36: The server receives the user's friend registration information and uses it as input for the generative AI. Step 37: The generative AI generates a response based on the user's friend registration information. Step 38: The server receives the response from the generative AI and sends it to the terminal. Step 39: The device displays the generative AI's response. Step 40: The user logs out and exits the application

[0417] Example 1 Next, a description will be given of Example 1. In the following description, the data processing device 12 is referred to as a "server" and the robot 414 is referred to as a "terminal."

[0418] Conventional information sharing applications have limited means for users to effectively interact in a virtual environment, making it difficult to obtain satisfactory results, especially in real-time communication and interactive experiences using generative AI. For example, there is a lack of systems that reflect user operations and chat content in real time and provide consistent responses. In addition, the quality of responses and accuracy of suggestions by generative AI have also been issues. In order to solve these problems, we provide the present invention.

[0419] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0420] In this invention, the server includes a means for authenticating the user's login information, a means for acquiring and managing the user's avatar information and game activity information, a means for receiving the user's text chat and voice chat input and sharing it with other users, and a means for generating responses based on the user's input using a generative AI model and sending the responses to the terminal. This allows the user to effectively interact with other users and virtual characters in real time within the virtual environment. Furthermore, the generative AI provides high-quality responses, improving the user's experience.

[0421] "User" refers to an individual who utilizes the system to interact and perform activities within the virtual environment. A "server" refers to a computing device that authenticates user login information, accesses a database to manage information, and responds to requests from users. "Terminal" refers to a device that allows a user to access the system and perform operations and send and receive information. Examples include smartphones and personal computers. "Login Information" refers to authentication information, such as a username and password, that a User enters to be authenticated to a System. "Authentication" refers to the process of verifying that a user can legitimately use a system based on the login information entered by the user. "Avatar" refers to a virtual character that a user uses to represent themselves within a virtual environment. "Game Activity Information" refers to data regarding activity within games and other virtual environments in which a User participates. "Text chat" refers to a means by which a user communicates with other users by inputting text. "Voice chat" refers to a means by which users communicate with other users using voice. A "generative AI model" refers to an artificial intelligence model that generates responses or suggestions based on user input. A "prompt sentence" refers to an input sentence that is provided to a generative AI model to generate some kind of response. "Feedback" refers to the evaluation or opinion given by the user regarding the response of the system or the generative AI model.

[0422] This invention is an information sharing application system that allows users to smoothly interact with other users and virtual characters in a virtual environment. This system is realized by the cooperation and operation of a server, a terminal, a user, and a generative AI model.

[0423] System Overview

[0424] server The server receives the user's login information and performs authentication. For example, an encryption library such as OpenSSL (registered trademark) is used for authentication. The server also stores the user's avatar information and past game activity information in a database (e.g., MySQL (registered trademark)) and manages this information. When the user sends a request, the server retrieves the corresponding data from the database and provides it to the user's terminal.

[0425] In addition, the server receives user input information and forwards it to the generative AI. Generative AI is usually implemented using machine learning libraries such as TensorFlow (registered trademark) and PyTorch (registered trademark). The server also manages shared information with other users. For example, if a user is chatting with a friend, the server receives the message and forwards it to the appropriate person.

[0426] Terminal The terminal provides an interface for users to access and log in to the application. It is responsible for sending text chat and voice chat information entered by the user to the server. It also transmits the movements of the avatar selected by the user and game activity to the server, enabling real-time operation.

[0427] The device receives the response from the server and displays it to the user. In particular, the response from the generative AI includes replies to the user's chat and suggestions for avatar movements. For example, in the case of a web application, this process realizes real-time communication on the browser using JavaScript (registered trademark) or WebSocket.

[0428] User Users log into the application using a terminal and interact with other users and virtual characters in the virtual environment. Users can communicate through text and voice chat, and can give their avatars animations to express greetings and emotions.

[0429] In addition, users can select educational games and activities and share them with other users. For example, if a user selects "Math Quiz," quiz questions will be provided by the generative AI, allowing users to compete or collaborate with other users to advance their learning.

[0430] Generative AI Model The generative AI model runs on a server and is responsible for generating responses based on user input. It receives information about the user's text chat, avatar movements, and other user and AI character choices and suggestions, and generates answers accordingly.

[0431] For example, Transformer-based models are often used as generative AI models, and when you input a prompt like "Please suggest that the user start a quiz," the AI ​​generates a response like "Let's ask, How about starting a quiz?"

[0432] Examples Examples of prompt statements 1. If the user wants their avatar to say "Hello!": Prompt: "Have the user say 'Hello!' to their avatar." Generative AI response: "Hello!"

[0433] 2. If the user wants to participate in an educational quiz: Prompt: "The user wants to participate in an educational math quiz." Generative AI response: "Great choice! Now, first question: What is 5 + 7?"

[0434] These elements work together to create a system that allows users to enjoy a wide variety of experiences within a virtual environment.

[0435] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0436] Program processing flow Step 1: User login process explanation User: Launches the application and enters username and password on the login screen. Terminal: Receives the user's login information and sends it to the server. Server: The received login information is compared against a database (e.g. MySQL (registered trademark)) and authentication is performed. Input is the user name and password, and output is the authentication result (success / failure).

[0437] operation The user enters "john_doe" and "password123" into the form and clicks the login button. The terminal encrypts this information and sends it to the server. The server checks the password against the records in its database, and if authentication is successful, it sends a "login successful" message to the terminal.

[0438] Step 2: Get your avatar and game activity information explanation Server: After the user successfully logs in, retrieve the avatar information and past game activity information from the database. Input is the user ID, and output is the retrieved avatar information and activity history. Terminal: displays information received from the server. Input is avatar information and activity information from the server, and output is the display of that information.

[0439] operation The server runs a database query to retrieve avatar information and past game history for user "john_doe". The device displays the acquired avatar information (e.g., "character wearing a blue hat") and past game history (e.g., "played math quiz on October 1st") on the screen.

[0440] Step 3: User interaction and communication explanation User: Controls an avatar and communicates with other users via text chat and voice chat. Terminal: Sends text and voice data entered by the user to the server. Input is text chat and voice data, and output is the result sent to the server. Server: Shares the received data with other users. Input is the received text or voice data, and output is the shared data.

[0441] operation A user types "Hello everyone!" in text chat and presses send. The terminal sends the text to the server. The server sends a notification to other logged-in users saying "john_doe: Hello everyone!"

[0442] Step 4: Generate prompts for generative AI and obtain responses explanation Server: Sends text chat content and avatar movement information to the generative AI. Input is text chat and avatar movement information, output is prompt text and generated responses. Generative AI: Processes the received prompt and generates an appropriate response. The input is the prompt and the output is the generated response. Server: Receives the response from the generative AI, performs further processing if necessary, and sends the response to the terminal. The input is the response from the generative AI, and the output is a notification to the terminal.

[0443] operation The server sends a prompt to the generative AI saying, "The user said 'Hello, everyone!' Please generate an appropriate response." The generative AI generates the response, "Hello john_doe! How are you?" The server receives the response from the generative AI and sends it to the terminal. The device will display the generative AI's response on the chat screen.

[0444] Step 5: Displaying the response and feeding it back to the server explanation Terminal: Receives the generative AI's response and displays it to the user. At the same time, sends the user's reaction and evaluation to the server. The input is the response from the generative AI, and the output is the display and feedback data for the user. User: Checks the generative AI's response and provides feedback if necessary. The input is the response from the generative AI, and the output is evaluation and feedback. Server: Receives user feedback and uses it to improve the generative AI model. The input is feedback data, and the output is feedback to the generative AI model.

[0445] operation The device will display the message "Hello john_doe! How are you?" on the chat screen. The user rates the response as "very good" and sends the feedback to the server. The server stores the feedback and later uses it as training data for the generative AI.

[0446] (Application example 1) Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[0447] Conventional educational entertainment and activities in virtual environments have lacked user interaction and personalization, resulting in low learning effectiveness and satisfaction. In addition, there were also insufficient methods for effectively utilizing responses generated in real time and mechanisms for deepening interactions between users.

[0448] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0449] In this invention, the server includes a means for enabling a user to communicate through text messaging or voice messaging, a means for giving a movement to a user's avatar to greet or express emotions, a means for providing educational entertainment or activities using a generation AI, a means for presenting educational entertainment or activities to a user and displaying the results through a visual output device or an audio output device, a means for acquiring a response by the generation AI in real time and presenting the response on a user interface, a means for providing educational entertainment or activities based on an answer obtained by inputting user attribute information and learning information into the generation AI, a data management means for analyzing a user's behavioral data and dialogue history and generating a more personalized response, a means for using the selection or suggestion of other users or AI agents based on the answer from the generation AI, a means for sharing the results of educational entertainment or activities between users and making suggestions to the user based on the answer from the generation AI, a sentence generation means for generating a prompt sentence and inputting it into the generation AI, and an evaluation means for evaluating and improving the response generated based on the prompt sentence. This allows a user to learn while enjoying personalized educational entertainment or activities, and to deepen interactions with other users and AI agents.

[0450] A "virtual environment" is a digital space created using computer technology in which users can have new experiences that are different from the real world. An "information sharing application" is software that enables multiple users to exchange and share information, including data, messages, media, etc., over the Internet. "Text messaging" means a means by which users communicate in real time by sending messages in text form. "Voice messaging" is a means by which users can send messages in the form of voice and communicate in real time. An "avatar" is a digital character that represents a user within a virtual space and is a means for expressing the user's movements and emotions. "Generative AI" is an artificial intelligence technique that generates text, images, etc. based on user input. "Educational entertainment" is games and activities that are designed to entertain users while also providing educational elements. A "visual output device" is a device, such as a digital display or a head-mounted display, that a user uses to obtain visual information. An "audio output device" is a device such as a speaker or headphones that a user uses to obtain audio information. A "user interface" is the arrangement and design of screens and input devices that allow a user to interact with and operate an application. "Attribute information" is data that indicates individual characteristics of a user, such as age, sex, and interests. "Learning information" is data related to the user's past learning history and knowledge level. "Behavioral data" is a record of what operations a user performed within the system. An "interaction history" is a record of messages and responses that have been exchanged between a user and a system in the past. An "AI agent" is a character that uses artificial intelligence and is designed to interact with users in a virtual space. A "prompt" is text that contains instructions or questions that are entered into the generation AI. A "sentence generation means" is a means for creating a prompt sentence to be supplied to the generation AI. An "evaluation means" is a means for evaluating the responses created by the generative AI to find areas for improvement.

[0451] The present invention provides a system for users to experience educational entertainment and activities in a virtual environment and to interact with other users and AI agents in real time. Specific embodiments for implementing the present invention are described below.

[0452] System Configuration The system mainly consists of three elements: the server, the terminal, and the user. The role of each element is explained below.

[0453] 1. Server Authentication means: The server receives the user's login information and performs authentication. For this purpose, a database (DB) system such as MySQL (registered trademark) or PostgreSQL (registered trademark) is used. Data management means: The server manages user avatar information, game activity data, and other user information. This is done using a cloud storage service (e.g., AWS (registered trademark) S3). Generative AI: The generative AI engine generates responses in real time based on user input. ChatGPT (registered trademark) and the like are used as generative AI engines. Communication method: The server uses WebSocket or HTTP protocols to manage communication with the user terminal.

[0454] 2. Terminal User input acceptance means: The terminal accepts the user's text and voice input and sends it to the server, using voice recognition software (e.g., Google® Speech-to-Text API) and a text input interface. Display means: The device displays the generated AI's responses and the results of game activities. This can be done using smart glasses or a head-mounted display. Movement control means: The terminal controls the movement of the avatar based on the user's operation. This is achieved by using motion capture technology.

[0455] 3. Users Input means: Users interact with other users and AI agents in the virtual environment through voice and text input. Avatar Control: Users control their own avatars to express actions and emotions. Activity Participation: Users participate in educational, entertainment, and activities.

[0456] Example of processing flow A specific example of the processing flow of the entire system is shown below.

[0457] 1. User Login The user enters login information from the terminal and transmits it to the server. The server uses the authentication means to authenticate the user.

[0458] 2. Activity selection The user selects educational entertainment and activities through the terminal's interface. The server receives the user's selection and provides the appropriate data.

[0459] 3. Generative AI response generation The user enters a question via text or voice. The server sends prompts to the generation AI, which generates responses in real time. Example response: "The user said, 'Do you want to proceed to the next question?' Generate an appropriate response."

[0460] 4. Display in the interface The server sends the generated response to the terminal, which presents it to the user through its visual and / or audio output devices.

[0461] In this manner, the present invention allows users to enjoy interactive educational entertainment and activities within a virtual environment.

[0462] The flow of the specific process in the application example 1 will be described with reference to FIG.

[0463] Step 1: The server receives the user's login information and performs authentication. The user ID and password entered by the user from the terminal are sent to the server. The server retrieves the corresponding user information from the database and checks whether it matches. The input is the user's login information, and the output is the authentication result (success or failure).

[0464] Step 2: A user selects an educational entertainment or activity using the interface of the terminal. The terminal transmits information about the activity selected by the user to the server. The server retrieves data corresponding to the selected activity from a database and transmits it to the terminal. The input is the information about the activity selected by the user, and the output is the data corresponding to the activity.

[0465] Step 3: The user inputs a question via text or voice. The text or voice data entered by the user into the terminal is sent to the server. The server converts the voice data into text (voice recognition) and sends it to the generation AI as a prompt. The input is the user's question (text, voice), and the output is the prompt.

[0466] Step 4: The server sends a prompt to the generation AI, which generates a response. The generation AI generates a text response based on the prompt and returns it to the server. The server receives the response and adjusts the response content, taking into account the user's attribute information and behavioral data. The input is the prompt and the user's attribute information, and the output is the generated response.

[0467] Step 5: The server sends the generated response to the terminal, which presents it to the user. The terminal displays the response message on a visual output device and optionally plays it through an audio output device. The input is the generated response and the output is the response presented in visual and audio output.

[0468] Step 6: The user confirms the generated response and decides on the next action: the user enters a further question or selects the next activity. The inputs are the user's confirmation and the next action, and the output is the information to proceed to the next step.

[0469] Step 7: The server provides an interaction mechanism for users to share the results of their educational activities. Users perform operations on their terminals to share their results, and the server manages the data. The input is the user's shared data, and the output is the shared information provided to other users.

[0470] The above is a specific processing flow of the program of the system that realizes the application example. This processing flow allows users to enjoy educational entertainment and activities in real time.

[0471] In addition, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0472] An embodiment for implementing the present invention includes the following elements.

[0473] 1. Server: The server realizes a system that combines an emotion engine that recognizes the user's emotions. The server analyzes the contents of the user's text chat and voice chat and the characteristics of the voice, and estimates the user's emotional state. The server also adjusts the difficulty and content of games and activities based on the user's emotional state estimated by the emotion engine.

[0474] 2. Terminal: The terminal is where the user logs into the application and operates and communicates within the virtual environment. The terminal transmits the user's text and voice chat input to the server, and the server provides games and activities based on the user's emotional state estimated by the emotion engine. The terminal also controls the user's avatar movement and interaction with other users.

[0475] 3. User: Users log into the application and interact with friends and virtual characters in a virtual environment. They communicate through text and voice chat, experience games and activities based on their emotional state estimated by an emotion engine, and interact with other users by giving their avatars animations to greet others and express emotions.

[0476] The above elements work together to realize a system that recognizes the user's emotions and provides games and activities accordingly. The server uses the emotion engine to estimate the user's emotional state and provide appropriate content. The terminal accepts the user's operations, communicates with the server, and displays responses via the emotion engine. The user can interact with friends and virtual characters in the virtual environment and enjoy content that matches their emotions.

[0477] The process flow will be explained below. Step 1: User logs into the application Step 2: The device sends the user's login information to the server Step 3: The server receives the user's login information and performs authentication. Step 4: The server sends the authentication result to the terminal. Step 5: The user selects an avatar to enter the virtual environment. Step 6: The device sends the user's avatar information to the server. Step 7: The server receives the user's avatar information and displays the avatar in the virtual environment. Step 8: The user interacts with friends and virtual characters through text and voice chat. Step 9: The device sends the user's chat input to the server Step 10: The server receives the user's chat input and passes it as input to the emotion engine. Step 11: The emotion engine analyzes the user's chat input and infers their emotional state Step 12: The server receives the results from the emotion engine and selects the appropriate game or activity. Step 13: The server sends the selected game or activity to the device. Step 14: The device displays the selected game or activity for the user to experience. Step 15: The user gives the avatar movement to greet or express emotion. Step 16: The terminal transmits the user's avatar movements to the server. Step 17: The server receives the user's avatar movements and moves the avatar in the virtual environment. Step 18: The user plays a game or activity with other users or AI characters. Step 19: The device sends the choices and suggestions of other users and AI characters to the server. Step 20: The server uses the choices and suggestions of other users and AI characters as inputs to the emotion engine. Step 21: The emotion engine generates responses based on the choices and suggestions of other users and AI characters. Step 22: The server receives the emotion engine's response and sends it to the terminal. Step 23: The device displays the emotion engine's response. Step 24: The user shares the results of the game or activity with other users. Step 25: The device sends the user's sharing contacts and chat contents to the server. Step 26: The server uses the user's shared contacts and chat contents as input to the emotion engine. Step 27: The emotion engine generates a response based on who the user is sharing with and what the chat is about. Step 28: The server receives the emotion engine's response and sends it to the terminal. Step 29: The device displays the emotion engine's response. Step 30: Use information about other users who are registered as friends by the user Step 31: User sends invitation notification and points to friends Step 32: The device sends the user's friend registration information to the server. Step 33: The server receives the user's friend registration information and uses it as input to the emotion engine. Step 34: The emotion engine generates a response based on the user's friend information. Step 35: The server receives the emotion engine's response and sends it to the terminal. Step 36: The device displays the emotion engine's response. Step 37: The user logs out and exits the application

[0478] Example 2 Next, a description will be given of Example 2. In the following description, the data processing device 12 is referred to as a "server" and the robot 414 is referred to as a "terminal."

[0479] Conventional communication systems and information sharing applications in virtual environments have the problem that it is difficult to adjust content in real time according to the user's emotional state. In addition, there is a lack of systems that can properly analyze the user's emotional state and provide games and activities based on it. This results in a uniform user experience and a lack of personalized support tailored to each individual's emotional state.

[0480] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for estimating the emotional state of the user by utilizing an emotion analysis engine, a means for adjusting the content of the game or activity based on the estimated emotional state, and a means for converting voice data into text data. This makes it possible to adjust the content in real time according to the emotional state of the user, and to provide a personalized experience.

[0481] A "virtual environment" is a computer-generated digital space that a user can interact with. A "communication application" is software that enables users to interact with other users and virtual characters through text and / or voice chat. "Text chat" is a means by which users exchange messages in real time through text input. "Voice chat" is a means by which users communicate in real time through voice input. "Character" refers to an avatar controlled by a user within a virtual environment, or a virtual entity controlled by a computer. An "emotion analysis engine" is an algorithm or software used to infer a user's emotional state from text or voice input. An "emotional state" refers to the psychological state a user feels at a particular point in time and can be classified as positive, negative, neutral, etc. "Games and Activities" are interactive entertainment or learning activities that users enjoy within a virtual environment. "Means for converting voice data into text data" refers to technology or software used to convert a user's voice input into text form. "Personalization" means optimizing systems and content to suit the characteristics and circumstances of each individual user. "Real-time" refers to responding immediately to user input and changes in the environment and processing without delay.

[0482] The following system configuration will be described as an embodiment of the present invention. The main components are a server, a terminal, and a user's operation method.

[0483] The server runs on a cloud service and is equipped with multiple function engines, including a sentiment analysis engine and a conversion engine. Specifically, the server uses a Natural Language Processing (NLP) API for sentiment analysis and a speech recognition API for converting voice data into text data. This makes it possible to analyze the user's emotional state and generate and provide interactive content in real time according to that state.

[0484] A terminal is a device that provides an interface for users to access applications and perform activities and communications within the virtual environment. Examples of terminals include PCs, smartphones, and tablets. The terminal has the function of sending the user's text chat and voice chat input to the server. The terminal also displays content to the user based on the emotional state received from the server. Furthermore, the terminal accepts user operations in real time and performs the actions and emotional expressions of the virtual character.

[0485] Users log into the application through their terminals and interact with other users and virtual characters in the virtual environment. The user's text and voice chat inputs are analyzed by an emotion analysis engine, and appropriate games and activities are provided based on the user's emotional state. For example, when a user says "hello," a character responds with a smile.

[0486] Examples

[0487] Example 1: A user interacts with a friend through text chat. 1. A user logs into the application from a terminal and starts a text chat. He types, "What's the weather like today?" 2. The device sends this text to the server. 3. The server uses a sentiment analysis engine to analyze "What's the weather like today?" and determines the emotional state as "neutral." 4. The server generates weather-related interactive activities as additional content and sends them to the terminal. 5. The terminal displays this information to the user, who can then enjoy further weather information or other activities.

[0488] Example 2: When a user interacts with a virtual character through voice chat 1. The user uses voice chat to talk to a virtual character, asking, "How are you?" 2. The terminal transmits the voice data to the server. 3. The server uses a voice recognition engine to convert the voice data into text data such as "How are you?" 4. The server passes the converted text data to a sentiment analysis engine and determines the emotional state as "positive." 5. The server generates an activity to respond cheerfully based on this positive emotional state. 6. The device displays this information to the user and the virtual character responds with something like "I'm fine! How about you?"

[0489] Examples of prompt statements "If the emotion engine analyzes the user's emotional state and estimates it to be 'happy,' please provide specific ideas for what kind of content (games or activities) should be provided."

[0490] From the above system configuration and specific examples, it can be understood that the emotion recognition and real-time content provision method according to the present invention can be clearly implemented.

[0491] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0492] Step 1: A user logs into the application from a terminal. A username and password are required as input. If the login is successful, the user profile and usage history are output from the server and displayed on the terminal. Specifically, the user enters his / her authentication information on the login screen and clicks the "Login" button.

[0493] Step 2: A user initiates a text or voice chat. As input, the text message or voice data is sent to the server. The server receives this data and prepares it for processing. For example, if a user sends a text message to a friend saying "How are you doing?", the device sends this message to the server.

[0494] Step 3: When the server receives the voice data, it converts the voice data into text data using Amazon (registered trademark) Transcribe. The server receives the voice data as input and generates converted text data as output. Specifically, the voice data of "How are you doing lately?" is converted into the text data of "How are you doing lately?"

[0495] Step 4: The server sends the received or converted text data to the Google Cloud Natural Language API for sentiment analysis. Text data is provided as input, and a sentiment score and a judgment result of the sentiment state are returned as output. For example, the text "How are you doing lately?" is analyzed to detect the sentiment state "neutral."

[0496] Step 5: The server generates appropriate content based on the results of the emotion analysis. It uses the emotional state determination as input and generates games and activities as output. Specifically, relaxing content is selected and generated for a "neutral" emotional state.

[0497] Step 6: The server sends the generated content to the terminal. As input, the server provides the generated content and sends information to the terminal. As output, the terminal displays this content to the user. For example, a puzzle game for relaxation is displayed on the terminal.

[0498] Step 7: The user uses the provided content. User operation data and additional feedback data are generated and sent to the server. Operation information and feedback from the user are received as input, and this is sent to the server as output. Specifically, when the user inputs his / her impressions while solving the puzzle, the data is sent to the server.

[0499] Step 8: The server again inputs the received feedback into the sentiment analysis engine and re-adjusts the content if necessary. Using the feedback data as input, the server provides the user with updated content or a re-analysis of the emotional state as output. For example, based on the feedback "This puzzle is a bit difficult", new content with a slightly lower difficulty level is generated.

[0500] (Application example 2) Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal".

[0501] Conventional applications for exchanging information in virtual environments lacked the ability to grasp the user's emotional state in real time and recommend and play optimal content. This made it difficult to provide content that was in tune with the user's emotions, limiting the user experience. Furthermore, there was no system that integrated educational elements with the provision of content based on emotions. A system that can solve these problems is needed.

[0502] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0503] In this invention, the server includes a means for receiving text chat or voice chat input by a user and presenting the received input to other users to realize exchanges by text chat or voice chat, a means for giving a movement to the user's avatar in response to the user's operation and having it express at least one of a greeting and an emotion, a means for inputting user information to a generation AI and providing the user with at least one of an educational game and an activity based on a response from the generation AI, and a means for analyzing the user's emotions in real time and recommending and playing appropriate entertainment (music or video content) based on the emotions. This makes it possible to provide content that is in line with the user's emotions, thereby providing a richer user experience.

[0504] An "information exchange application" is software that enables a user to communicate with friends and virtual characters within a virtual environment through text chat and voice chat. An "avatar" is a digital character that represents a user within a virtual environment and expresses the user's actions and emotions. "Generative AI" is artificial intelligence that automatically recommends educational games, activities, and content based on the user's input information and emotional state. "Emotion analysis" is a technology that analyzes the content of a user's text chat or voice chat and infers the user's emotional state. "Content recommendation" is a function that suggests optimal music and video content based on the user's emotional state and attribute information. "Educational games and activities" are interactive games and activities designed to learn or enhance knowledge. "Real-time" refers to immediate processing or response in the current time. A "virtual environment" is a computer-generated digital space with which a user interacts.

[0505] The embodiment for implementing the invention includes the following elements.

[0506] server The server has a built-in emotion engine that recognizes the user's emotions. Specifically, the emotion engine analyzes and estimates the user's emotional state from the contents of text and voice chat. The server utilizes a generative AI model to recommend and play optimal music and video content based on the user's emotional state. The server also adjusts the difficulty and content of educational games and activities according to the user's emotional state. The system includes the following software: emotion_engine(emotion engine) content_recommender (content recommendation system) text_analysis audio_recognition

[0507] Terminal The terminal is the device through which the user logs into the application and performs operations and communication within the virtual environment. The terminal receives the user's voice and text input and transmits it to the server. It also provides the user with content and activities based on the emotional state estimated by the emotion engine. Terminals include smartphones and smart glasses. The terminal has the following functions: Function to give movement to avatars Text and voice chat functionality Content playback function

[0508] User Users log into the application and interact with friends and virtual characters in a virtual environment. They communicate through text and voice chat, and enjoy music and video content based on their emotional state estimated by the emotion engine. They also participate in educational games and activities, experiencing adjustments in difficulty and content according to their emotional state. An example of a specific user scenario is as follows:

[0509] 1. The user launches the app and speaks, "I'm feeling a bit sad today." 2. The app analyzes the voice and the emotion engine determines that the user is feeling "sad." 3. The content recommendation system recommends soothing music and inspiring videos that suit the emotion of "sad." 4. The recommended content will be played automatically.

[0510] Examples of prompt statements User input: "I'm feeling a bit sad today." Analyzed emotion: "Sadness" Suitable content: "Healing music" "Inspirational videos" Suggested prompt: "If the user is feeling sad, please recommend soothing music that will soothe the soul. Also, please recommend an inspiring video."

[0511] In this way, content that is in tune with the user's emotions can be provided, providing a rich user experience.

[0512] The flow of the specific process in the application example 2 will be described with reference to FIG.

[0513] Step 1: The device receives the user's text chat or voice chat input. When the user speaks "I'm feeling a bit sad today," the device receives this voice data. The input data is voice data. The device then uses voice recognition software to convert this voice data into text data. The converted text data is output.

[0514] Step 2: The terminal transmits the converted text data to the server. The input here is text data, and the output to the server is also text data. Specifically, the terminal transmits the text data to the server via the network.

[0515] Step 3: The server inputs the received text data into an emotion engine. The emotion engine analyzes the user's emotional state from the text data. The input here is the text data, and the emotion analysis engine estimates the emotional state. The estimated emotional state, e.g., "sadness," is output.

[0516] Step 4: The server inputs the estimated emotional state into a content recommendation system, which recommends optimal music and video content based on the emotional state. The input is the emotional state, and the content recommendation system uses this data to select recommended content. The selected recommended content is output.

[0517] Step 5: The server transmits the recommended content to the terminal. The input is the recommended content data, and the output is also the recommended content data. The server transmits this to the terminal via the network.

[0518] Step 6: The terminal presents the received content to the user. Specifically, it uses the terminal's media player function to play music and videos. The input is the recommended content data, and the output is the content to be played that is presented to the user. The terminal plays the content using a screen display and speakers.

[0519] In this way, a system is realized that provides content in real time according to the user's emotions.

[0520] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires a voice indicating a user input for the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0521] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0522] In the above embodiment, an example was given in which the specific process was performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the robot 414.

[0523] The emotion identification model 59 as an emotion engine may determine the emotion of the user according to a specific mapping. Specifically, the emotion identification model 59 may determine the emotion of the user according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the emotion of the robot, and the identification processing unit 290 may perform identification processing using the emotion of the robot.

[0524] FIG. 9 is a diagram showing an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive emotions are arranged. The more outside the concentric circles, the more emotions that represent states and actions that arise from a state of mind are arranged. Emotions are a concept that includes emotions and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions that occur in the brain are arranged. On the right side of the concentric circles, emotions that are generally induced by situational judgment are arranged. On the upper and lower sides of the concentric circles, emotions that are generally generated from reactions that occur in the brain and are induced by situational judgment are arranged. In addition, on the upper side of the concentric circles, emotions of "pleasure" are arranged, and on the lower side, emotions of "discomfort" are arranged. In this way, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[0525] These emotions are distributed in the three o'clock direction of emotion map 400 and usually fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[0526] The inside of emotion map 400 represents what is going on inside one's mind, and the outside of emotion map 400 represents behavior, so the further out you go on emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[0527] Here, human emotions are based on various balances such as posture and blood sugar level, and when these balances are far from the ideal, it indicates an unpleasant state, and when they are close to the ideal, it indicates a pleasant state. Emotions can also be created for robots, cars, motorcycles, etc., based on various balances such as posture and battery level, so that when these balances are far from the ideal, it indicates an unpleasant state, and when they are close to the ideal, it indicates a pleasant state. The emotion map may be generated, for example, based on the emotion map of Dr. Mitsuyoshi (Research on speech emotion recognition and emotion brain physiological signal analysis system, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). On the left half of the emotion map, emotions belonging to an area called "reaction" where sensation is dominant are lined up. On the right half of the emotion map, emotions belonging to an area called "situation" where situation recognition is dominant are lined up.

[0528] The emotion map defines two emotions that promote learning. The first is the negative emotion around the middle of "repentance" or "remorse" on the situation side. In other words, this is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the positive emotion around "desire" on the response side. In other words, this is when the robot has positive feelings such as "I want more" or "I want to know more."

[0529] The emotion identification model 59 inputs the user input to a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the emotion of the user. This neural network is pre-trained based on multiple learning data that are combinations of the user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in Fig. 10. Fig. 10 shows an example in which multiple emotions, "relief," "calm," and "encouraging," have similar emotion values.

[0530] Although the system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, the system according to the present disclosure is not necessarily implemented in a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program that runs on a personal computer, or an application that runs on a smartphone or the like. The method according to the present disclosure may be provided to a user in the form of SaaS (Software as a Service).

[0531] In the above embodiment, an example is given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to input data.

[0532] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable computer-readable non-transitory storage medium such as a Universal Serial Bus (USB) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[0533] In addition, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 upon request from the data processing device 12.

[0534] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[0535] As the hardware resource for executing the specific process, various processors as shown below can be used. An example of the processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing the specific process by executing software, i.e., a program. Another example of the processor is a dedicated electric circuit, which is a processor having a circuit configuration designed exclusively for executing the specific process, such as a Field-Programmable Gate Array (FPGA), a Programmable Logic Device (PLD), or an Application Specific Integrated Circuit (ASIC). Each processor has a built-in or connected memory, and each processor executes the specific process by using the memory.

[0536] The hardware resource that executes the specific process may be one of these various processors, or may be a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific process may be a single processor.

[0537] As an example of a configuration using one processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a configuration using a processor that realizes the functions of the entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[0538] Furthermore, more specifically, the hardware structure of these various processors can be an electric circuit that combines circuit elements such as semiconductor elements. The specific processes described above are merely examples. It goes without saying that unnecessary steps may be deleted, new steps may be added, or the order of processes may be changed without departing from the spirit of the invention.

[0539] The above description and illustrations are detailed descriptions of the parts related to the technology of the present disclosure, and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, function, action, and effect is an example of the configuration, function, action, and effect of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above description and illustrations, within the scope of the gist of the technology of the present disclosure. In addition, in order to avoid confusion and to facilitate understanding of the parts related to the technology of the present disclosure, the above description and illustrations omit explanations of technical common sense that do not require explanation in order to enable the implementation of the technology of the present disclosure.

[0540] All publications, patent applications, and standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or standard was specifically and individually indicated to be incorporated by reference.

[0541] The following is further disclosed regarding the above embodiment.

[0542] <Appendix 1> A system for providing a social networking application that enables a user to interact with friends and virtual characters in a virtual environment, comprising: A means for enabling the users to communicate with each other through text chat or voice chat; A means for moving the user's avatar to say hello or express emotions; A means to leverage generative AI to provide educational games and activities; A system including:

[0543] <Appendix 2> In the system according to claim 1, A means for providing the educational game or activity based on an answer obtained by inputting attribute information and learning information of the user into the generation AI; A system including:

[0544] <Appendix 3> In the system according to claim 1 or 2, A means for utilizing the selections and suggestions of other users or AI characters based on the answers from the generating AI; and A means for sharing the results of the educational games and activities among the users and making suggestions to the users based on the answers of the generative AI; A system including:

[0545] <Appendix 4> In the system according to claim 1, The system includes a means for recognizing the emotions of the user by combining an emotion engine, and providing content according to the emotions.

[0546] "Example 1" (Claim 1) means for authenticating the user's login information; A means for acquiring and managing avatar information and game activity information of the user; means for receiving said user's text chat and / or voice chat input and sharing it with other users; means for generating a response based on the user's input using the generative AI model and transmitting the response to the terminal; A system including:

[0547] (Claim 2) A means for transmitting avatar movements and game activities to a server based on the user's operation and providing prompt sentences to a generative AI model; means for displaying responses from the generative AI model to a user and collecting user feedback; 2. The system of claim 1, comprising:

[0548] (Claim 3) means for suggesting other virtual characters or game activities to the user based on responses from the generative AI model; A means for the user to evaluate the response of the generative AI model and improve the generative AI model based on the evaluation result; 2. The system of claim 1, comprising:

[0549] "Application example 1" (Claim 1) A system for providing an information sharing application that enables a user to interact with other users and virtual characters in a virtual environment, comprising: means for enabling said users to interact with each other via text messaging and / or voice messaging; A means for giving a movement to the user's avatar to say hello or express emotions; A means to use generative AI to provide educational entertainment and activities; means for presenting said educational entertainment or activity to a user and displaying the results via a visual or audio output device; A means for acquiring a response from the generation AI in real time and presenting the response on a user interface; A system including:

[0550] (Claim 2) A means for providing the educational entertainment or activity based on an answer obtained by inputting attribute information and learning information of the user into the generating AI; A data management means for analyzing the user's behavioral data and dialogue history to generate a more personalized response; 2. The system of claim 1, comprising:

[0551] (Claim 3) A means for utilizing the selections and suggestions of other users or AI agents based on the answers from the generating AI; A means for sharing the results of the educational entertainment and activities among the users and making suggestions to the users based on the answers of the generating AI; A sentence generation means for generating the prompt sentence and inputting the prompt sentence to the generation AI; evaluation means for evaluating and improving responses generated based on said prompts; 3. The system of claim 1 or 2, comprising:

[0552] "Example 2 of combining emotion engines"

[0553] (Claim 1) A system for providing a communication application that enables a user to interact with other users or virtual characters in a virtual environment, comprising: A means for enabling the users to communicate with each other through text chat or voice chat; A means for giving a movement to the user's character to say hello or express emotions; means for estimating an emotional state of the user utilizing an emotion analysis engine; A means for adjusting game or activity content based on the estimated emotional state; and A means for converting voice data into text data; A system including:

[0554] (Claim 2) a means for providing the game or activity based on a result obtained by inputting attribute information and analysis information of the user into the emotion analysis engine; 2. The system of claim 1, comprising:

[0555] (Claim 3) means for selecting or suggesting other users or virtual characters based on results from said emotion analysis engine; A means for sharing the results of the game or activity among the users and making suggestions to the users based on the results of the emotion analysis engine; 2. The system of claim 1, comprising:

[0556] "Application example 2 when combining emotion engines" (Claim 1) A system for providing an information exchange application that enables a user to interact with friends and virtual characters in a virtual environment, comprising: A means for enabling said users to communicate with each other through text chat or voice chat; A means for giving a movement to the user's avatar to say hello or express emotions; A means to leverage generative AI to provide educational games and activities; A means for analyzing user emotions in real time and recommending and playing appropriate music and video content based on those emotions; A system including:

[0557] (Claim 2) A means for providing the educational game or activity based on an answer obtained by inputting attribute information and learning information of the user into the generation AI; A means for changing the content recommended by the generation AI in real time based on the user's sentiment analysis; 2. The system of claim 1, comprising:

[0558] (Claim 3) A means for utilizing the selections and suggestions of other users or AI characters based on the answers from the generating AI; A means for sharing the results of the educational games and activities among the users and making suggestions to the users based on the answers of the generating AI; A means of recommending and playing optimal content based on emotions; 3. The system of claim 1 or 2, comprising: [Explanation of symbols]

[0559] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A system for providing an information sharing application service that enables interaction between a first user and a second user or a virtual character in a virtual environment, comprising: a presentation means for receiving a text chat or voice chat input from the first user and presenting the received input to the second user, thereby realizing the interaction through the text chat or the voice chat; a control means for causing an avatar of the first user to move in response to an operation of the first user, and for causing the avatar to express at least one of a greeting and an emotion; providing means for inputting information of the first user into a generating AI and providing at least one of an educational game, an activity, and entertainment to the first user based on a response from the generating AI; A system including:

2. 2. The system of claim 1, wherein the information of the first user input to the generation AI includes input by the first user in the text chat or voice chat, movement of the first user's avatar in response to the operation of the first user, or an emotional state of the first user estimated by an emotion engine from the input.

3. The system of claim 1 , wherein the providing means adjusts at least one of the difficulty and content of the game or activity provided to the first user based on the response from the generating AI.

4. The system of claim 1 , further comprising a collection means for inputting information of the first user into a generative AI, presenting a response from the generative AI to the first user, and collecting feedback from the first user on the response for improving the generative AI's model.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A