System
The system addresses the limitations of existing virtual agent systems by analyzing user input to generate natural and real-time interactions with avatars, improving user experience through dynamic character behavior and animations.
Patent Information
- Application Number
- JP2024133612
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-08
- Publication Date
- 2026-02-20
AI Technical Summary
Existing virtual agent systems lack the means to realize natural interactions and natural conversations with users, and have particular problems with real-time performance and natural responses, and lack the means to realize natural interactions with users, and have limited interaction with avatars.
A system that analyzes user text input and touch events in real-time, using a large-scale language model to generate character behavior and animations, allowing for seamless and natural interactions.
Enables natural, real-time interactions with avatars by dynamically generating character behavior and animations based on user input, enhancing user experience and interaction quality.
Smart Images

Figure 2026030628000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Existing virtual agent systems have the problem that user interaction is limited to canned responses or pre-defined patterns, making it difficult to achieve natural interactions. Furthermore, because the character's behavior is patterned, the range of expression is narrow, making it difficult to achieve dynamic responses in real time. Furthermore, the cost of implementing these features is high. Therefore, a new interface is needed that allows users to interact with avatars more naturally in real time. [Means for solving the problem]
[0005] This invention provides a system that receives and analyzes text input and touch events from a user in real time. Specifically, it includes a means for analyzing the user's text input and inferring the user's intent, and a means for interpreting the user's touch events. It also uses a large-scale language model (LLM) to dynamically generate character behavior based on the analyzed user intent. It also includes a means for drawing animation frames in real time based on the generated behavior. This allows the character to seamlessly respond to the user's input text and touch events, achieving more natural interactions. This system can also generate behaviors that take into account the character's personality information, allowing for more personalized interactions with the user. Furthermore, it can provide smooth animations using technology that generates more than 100 animation frames per second.
[0006] "Text input" refers to the act of a user inputting a string of characters using a keyboard or a touch screen.
[0007] A "touch event" is interaction data that occurs when a user touches, presses, swipes, or performs other operations on a touch panel or touch screen.
[0008] "Intent analysis" is the process of inferring and understanding a user's intentions and will from the text they enter and the touch events they make.
[0009] A "large-scale language model (LLM)" is an artificial intelligence model trained on large amounts of text data, capable of understanding and generating natural language.
[0010] "Character behavior" refers to the movements or actions displayed by the virtual agent and is dynamically generated based on user input.
[0011] "Image generation technology" is a technology that uses algorithms to generate images and animation frames in real time.
[0012] An "animation frame" is an individual still image that makes up an animation, and movement is expressed by displaying these images in succession.
[0013] "Real-time rendering" is the process of instantly generating and displaying animation frames in response to user input and interaction.
[0014] "Seamless animation" refers to displaying successive animation frames at a high frame rate so that the movement appears smooth on the screen.
[0015] "Character personality information" is data relating to the unique personality and characteristics of a character, and based on this, the character's behavior is personalized to be more human-like. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] This invention is a system for realizing natural interactions between users and virtual agents. This system analyzes user text input and touch events in real time, dynamically generates character behaviors based on the analysis, and displays them as animations.
[0038] System Configuration
[0039] User Interface
[0040] The device has an input field for users to enter text and a touchscreen for detecting touch events, through which users interact with the system. For example, a user can type "hello" or tap on an avatar's face.
[0041] Sending and Receiving Data
[0042] The device immediately sends the text and touch events entered by the user to the server via data communication over the network. For example, when a user types "hello," the data is sent from the device to the server.
[0043] Data analysis
[0044] The server analyzes the received user text and touch events. It uses a large-scale language model (LLM) to infer the user's intent from their input. It also interprets the user's gestures based on the touch events. For example, the LLM can analyze the text "Hello" and generate an appropriate response such as "Hello! What's up?"
[0045] Behavior Generation
[0046] Based on the analysis results, the server generates appropriate behavior for the character, taking into account the character's personality information. For example, a character with a cheerful personality will be set to respond with a smile.
[0047] Generate animation
[0048] The server uses image generation technology to draw animation frames based on the generated behavior in real time, for example, generating 100 animation frames per second of an avatar smiling.
[0049] Viewing Data
[0050] The generated animation frames are sent from the server to the device and played back in real time on the device. A seamless animation is displayed to the user, realizing natural interaction between the user and the avatar. For example, when the user says "Hello," the avatar will respond with a smile, "Hello! What's wrong?"
[0051] Specific examples
[0052] Consider the case where a user inputs the text "Hello" to an avatar. At this time, the device sends the user's input to the server. The server uses LLM to analyze the input "Hello" and infer the user's intention. Based on this, the server generates a behavior for the avatar to respond with a smile, "Hello! What's wrong?" and uses image generation technology to render this as an animation frame. The rendered animation frame is sent to the device and played back in real time on the user's device. In this way, the user can experience natural interaction with the avatar.
[0053] The above is an embodiment of the "Touch and Link" system of the present invention, which allows users to interact with avatars naturally and in real time.
[0054] The processing flow will be explained below.
[0055] Step 1:
[0056] The user enters "Hello" in the text input field of the terminal and presses the send button.
[0057] Step 2:
[0058] The terminal receives the user's text input "Hello" and sends this data to the server.
[0059] Step 3:
[0060] The server receives the text data sent from the terminal.
[0061] Step 4:
[0062] The server provides text data to a large-scale language model (LLM) and begins analysis.
[0063] Step 5:
[0064] The server obtains the analysis results of the LLM and infers the user's intention. For example, if the LLM receives the input "Hello," it generates the response text "Hello! What's wrong?"
[0065] Step 6:
[0066] When the server receives a touch event, it analyzes it and interprets which part was touched and how. For example, if a user taps on the avatar's head, it obtains that information.
[0067] Step 7:
[0068] Based on the intent analysis and interpretation of the touch events, the server determines the appropriate behavior of the avatar, for example, the avatar responding with a smile, "Hello! What's up?"
[0069] Step 8:
[0070] The server uses image generation technology to generate animation frames based on the determined behavior, for example, generating 100 animation frames per second of an avatar smiling.
[0071] Step 9:
[0072] The animation frames generated by the server are divided into packets and sent to the terminal.
[0073] Step 10:
[0074] The device sequentially reads the animation frames received from the server and displays them on the user's screen in real time. For example, an animation of an avatar smoothly smiling and saying, "Hello! What's up?" is displayed.
[0075] The above is the specific program processing flow of the "Touch and Link" system of the present invention.
[0076] Example 1
[0077] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0078] Conventional interaction systems lack the means to realize natural interactions between users and virtual agents, and have particular problems with real-time performance and natural responses. The present invention aims to enhance natural interaction with users by more effectively analyzing user text input and touch events and displaying the behavior of dynamically generated characters in real time.
[0079] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0080] In this invention, the server includes means for receiving text input from a user and analyzing the user's intention, means for receiving touch events from the user and interpreting gestures, means for dynamically generating character behavior from the user's intention using a large-scale language model, means for drawing animation frames in real time based on the generated behavior using an image generation means, and means for transmitting the drawn animation frames to the user's display device and displaying them, thereby enabling natural interaction between the user and the virtual agent in real time.
[0081] "Text input" refers to the act of a user inputting characters and symbols using an input device such as a terminal.
[0082] "Intention" refers to the purpose or desire of what the user wants to communicate to the system or what they are asking for.
[0083] A "touch event" is an input signal that occurs when a user operates a touch screen, and includes gesture operations such as tapping, swiping, and pinching.
[0084] A "gesture" is a hand movement or operation pattern that a user uses when operating a touchscreen or device.
[0085] A "large-scale language model" is a machine learning model that is trained on a huge amount of text data and is used to perform natural language processing.
[0086] "Character behavior" refers to the actions and reactions that a character takes in response to user input and the environment.
[0087] "Image generation means" refers to a method for generating images or visual content using computer technology.
[0088] An "animation frame" is a still image that makes up an animation, and movement is expressed by displaying these images in succession.
[0089] A "display device" is hardware for visually displaying generated visual content, including a screen or display.
[0090] "Real-time" means responding or processing immediately to user input or events without delay.
[0091] The present invention provides a system for realizing natural interactions between a user and a virtual agent. This system analyzes the user's text input and touch events in real time, dynamically generates character behavior based on the analysis, and displays it as animation. Specific embodiments are described below.
[0092] 1. User Interface
[0093] The device has an input field where the user can enter text and a touchscreen that detects touch events, which the user uses to interact with the system. For example, the user can type "hello" or tap on an avatar's face.
[0094] 2. Sending and Receiving Data
[0095] The device immediately sends the text and touch events entered by the user to the server via data communication over the network. For example, when a user enters "hello," the text data is sent from the device to the server.
[0096] 3. Data Analysis
[0097] The server analyzes the received user text and touch events. It uses a large-scale language model (LLM) to infer the user's intent from the input. It also interprets the user's gestures based on the touch events. For example, the text "Hello" is analyzed and an appropriate response is generated: "Hello! What's wrong?"
[0098] 4. Behavior Generation
[0099] The server generates appropriate behavior for the character based on the analysis results. The character's personality information is also taken into consideration. For example, a character with a cheerful personality will be set to respond with a smile, and will respond to user input with a smile saying, "Hello! What's wrong?"
[0100] 5. Generating Animations
[0101] The server uses image generation technology to draw animation frames in real time based on the determined behavior. For example, 100 animation frames per second of an avatar smiling and saying "Hello! What's wrong?"
[0102] 6. Displaying Data
[0103] The generated animation frames are sent from the server to the device and played back in real time. This allows users to experience natural interactions with the avatar through seamless animation. For example, when the user says "Hello," the avatar will respond with a smile, "Hello! What's wrong?"
[0104] Specific examples
[0105] When a user inputs the text "Good morning" to an avatar, the device sends this input data "Good morning" to the server. The server uses LLM to analyze this text, understands the intent of the greeting, and generates a response such as "Good morning! Let's do our best today!". The server then generates an animation of the avatar smiling back based on this response. This animation frame is sent to the device and displayed in real time. The user can feel the avatar's response naturally, and simple interactions can lead to more complex and enjoyable interactions.
[0106] Prompt Sentence Examples
[0107] Below are some example prompts to input to the generative AI model:
[0108] Specific prompt:
[0109] When a user types "Hello," a cheerful avatar will respond with a smile, "Hello! What's up?" Specifically, you'll analyze the text, generate an appropriate response, and then animate it.
[0110] The above is an embodiment of the "Touch and Link" system of the present invention, which allows users to enjoy natural, real-time interaction with avatars.
[0111] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0112] Step 1:
[0113] User interface input
[0114] The device has an input field and a touch screen, which the user uses to input text into the system or to send touch events. For example, the user may type "hello" or tap on the face of an avatar. The input in this case is text data or touch event data.
[0115] Step 2:
[0116] Sending data
[0117] Any text input or touch events made by the user are immediately sent from the device to the server. Here, the input data (text or touch events) is sent via network communication and transmitted to the server in its original format. Specifically, the user's text "Hello" is sent from the device to the server.
[0118] Step 3:
[0119] Data analysis
[0120] The server analyzes the received user input data. It passes the input data (text and touch events) to a large-scale language model (LLM) that performs calculations to infer the user's intent. It also interprets gestures based on touch events. For example, the LLM analyzes the text "Hello" and generates an appropriate response, "Hello! What's wrong?" The analysis results are obtained as output.
[0121] Step 4:
[0122] Behavior Generation
[0123] The server dynamically generates character behavior based on the analysis results. The input here is the analysis results, and the output is specific character behavior information. Taking into account the character's personality information, for example, a character with a cheerful personality is set to respond with a smile.
[0124] Step 5:
[0125] Generate animation
[0126] The server uses image generation technology to draw animation frames based on the generated behavior information. The input is the character's behavior information, and animation frames are generated in real time based on that information. For example, 100 animation frames per second of a character smiling and saying "Hello! What's wrong?" are generated. The output is the generated animation frames.
[0127] Step 6:
[0128] Viewing Data
[0129] The generated animation frames are sent from the server to the device and played back in real time on the device. The input is the generated animation frames, which are sent to the device over the network and displayed seamlessly to the user. For example, in response to the user saying "Hello," an animation of an avatar smiling and saying "Hello! What's wrong?" is played back on the device. The output is the animation displayed on the device.
[0130] (Application example 1)
[0131] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0132] Conventional virtual agent systems have limited interaction with users, making it difficult to have natural conversations. Furthermore, they lack functionality as a shopping assistant, meaning they cannot provide sufficient support when users search for and purchase products. This results in a poor user experience and makes it difficult to promote sales effectively.
[0133] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0134] In this invention, the server includes means for receiving text input from a user and analyzing the user's intention, means for receiving touch events from the user and interpreting gestures, means for dynamically generating character behavior based on the user's intention using a large-scale language model, means for drawing animation frames in real time based on the generated behavior using image generation technology, means for transmitting the drawn animation frames to the user's terminal and displaying them, and means for receiving input from the user for searching for and purchasing products and providing product information based on the input. This enables natural interaction with the user, provides high performance as a shopping assistant, improves the user experience, and enables effective sales promotion.
[0135] "Text input from a user" refers to the act of a user inputting text information using a keyboard or touch screen.
[0136] A "means for analyzing user intent" is a means for analyzing what the text means and what it is trying to convey based on the text entered by the user.
[0137] A "touch event" is an event that occurs when a user taps, swipes, pinches, or performs other actions on a touchscreen.
[0138] The "means for interpreting a gesture" is a means for interpreting what gesture the user has made based on a touch event.
[0139] A "large-scale language model" is a machine learning model trained on massive amounts of text data, capable of understanding and generating natural language.
[0140] "Means for dynamically generating character behavior" refers to means for generating character movements and facial expressions in real time in response to user input and intentions.
[0141] "Image generation technology" is a technology that uses computer graphics to generate visual images and animations.
[0142] "Means for drawing animation frames in real time" refers to means for instantly creating and displaying animation frames based on user input and gestures.
[0143] An "animation frame" is an individual still image that makes up an animation.
[0144] "User terminal" refers to an electronic device used by a user, such as a smartphone, tablet, or PC.
[0145] The "displaying means" is a means for displaying the generated animation frames on the screen of the user's terminal.
[0146] "Input for searching and purchasing products" refers to the input operations performed by a user through a search bar or voice input to search for or purchase products.
[0147] "Means for providing product information" refers to the means for presenting information about products that users wish to search for and purchase.
[0148] "Character personality information" is information about the character's personality that is taken into consideration when determining the character's behavior and responses.
[0149] "Means for generating 100 or more animation frames per second" means means for generating 100 or more animation frames per second, thereby providing seamless animation.
[0150] The present invention is a system that realizes natural interactions with users and acts as a virtual shopping assistant. This system generates character behavior in real time based on user input, and supports product search and purchase.
[0151] System Configuration
[0152] User Interface
[0153] The device includes an input field for users to enter text or voice input, and a touchscreen for detecting touch events such as taps and swipes, through which users interact with the system. For example, a user can type "I want a smartwatch" or tap on a specific product.
[0154] Sending and Receiving Data
[0155] The device immediately sends text, touch events, and voice input from the user to the server via data communication over the network. For example, if a user types "I want a smartwatch," that data is sent from the device to the server.
[0156] Data analysis
[0157] The server analyzes the received user text, touch events, and voice data. It uses a large-scale language model (e.g., OpenAI's GPT-4) to infer the user's intent from their input. It also interprets the user's gestures based on the touch events. For example, a large-scale language model can analyze the text "I want a smartwatch" and generate the optimal response and product information.
[0158] Behavior Generation
[0159] Based on the analysis results, the server generates appropriate behavior for the character, taking into account the character's personality. For example, a character with a cheerful personality might respond with a smile, "What do you think of this smartwatch?"
[0160] Generate animation
[0161] The server uses image generation technology (e.g., Unity, Unreal Engine) to draw animation frames based on the generated behavior in real time. For example, it generates 60 animation frames per second of a virtual assistant pointing at a smartwatch.
[0162] Viewing Data
[0163] The generated animation frames are sent from the server to the device and played back in real time on the device. A seamless animation is displayed to the user, realizing natural interaction between the user and the virtual assistant. For example, if a user says, "I want a smartwatch," an animation is displayed in which the virtual assistant smiles and asks, "How about this smartwatch?"
[0164] Specific examples
[0165] Consider the case where a user texts the assistant, saying, "I want a smartwatch." At this time, the device sends the user's input to the server. The server uses GPT-4 to analyze the input, "I want a smartwatch," and infers the user's intent. Based on this, the server generates a behavior in which the virtual assistant responds with a smile, "How about this smartwatch?" and uses image generation technology to render this as an animation frame. The rendered animation frame is sent to the device and played back in real time on the user's device. In this way, the user experiences natural interactions with the virtual assistant, enabling smooth product search and purchase.
[0166] A specific example of a prompt sentence is as follows:
[0167] A user says, "I want a smartwatch." As a virtual shopping assistant, generate an appropriate response and consider the appropriate avatar behavior for that response.
[0168] This system allows users to interact with the virtual assistant naturally and in real time, making product searches and purchases more convenient.
[0169] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0170] Step 1:
[0171] The user provides text or voice input to the device. The device receives the user's input and sends it to the server as JSON format data. At this time, the input includes text or voice data such as "I want a smartwatch." The processing involves sending the input data to the server via the network. The output is the server receiving the input data.
[0172] Step 2:
[0173] The server analyzes the received text and voice data from the user. During the data analysis process, a large-scale language model (e.g., OpenAI's GPT-4) is used to infer the user's intent from their input. The data processing performed in this step involves text analysis and intent inference. Text data is given as input, and the analysis results are obtained as output. Specifically, the input "I want a smartwatch" is interpreted as "The user is looking for a smartwatch."
[0174] Step 3:
[0175] The server generates appropriate behavior for the virtual character based on the analysis results. In this process, the behavior is determined taking into account the character's personality information. For example, a character with a cheerful personality will respond with a smile. The input in this step is the analysis results, and the output is the generated character behavior. Specifically, the server generates the response "We recommend a smartwatch" and the corresponding character behavior.
[0176] Step 4:
[0177] The server uses image generation technology (e.g. Unity, Unreal Engine) to draw animation frames based on the generated behavior in real time. Character behavior information is given as input, and animation frames are the output. The specific operation in this step is for the animation software to generate frames.
[0178] Step 5:
[0179] The server sends the generated animation frames to the terminal. The input is the animation frames, and the output is the transmission of frames to the terminal. Specifically, the server sends animation data in real time over the network and prepares it for display on the user's terminal.
[0180] Step 6:
[0181] The device displays the received animation frames, providing the user with a seamless interaction experience. The input is the received animation frames, and the output is the display on the user's screen. Specifically, the device plays the animation frames, making it appear as if the virtual character is talking to the user.
[0182] This allows users to receive product information from the virtual shopping assistant through natural interaction.
[0183] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0184] This invention is a system incorporating an emotion analysis engine to realize natural interactions between users and virtual agents. This system analyzes user text input and touch events in real time, dynamically generates character behavior based on the analysis, and displays it as animation. Furthermore, by using the emotion engine to estimate the user's emotions and reflect them in the character's behavior, it realizes more natural, human-like responses.
[0185] System Configuration
[0186] User Interface
[0187] The device has a text entry field, a touchscreen, and voice input. Users can interact with the system through these, for example, by typing "hello" or by tapping on an avatar's face. They can also communicate using voice input.
[0188] Sending and Receiving Data
[0189] The device immediately transmits the user's text input, touch events, and voice input to the server using data communication over the network. For example, when a user types "hello," the data is sent from the device to the server.
[0190] Data analysis
[0191] The server analyzes the received user text, touch events, and voice data. It uses a large-scale language model (LLM) to infer the user's intent from the input. It also uses an emotion engine to infer the user's emotion. For example, the LLM analyzes the text "Hello" and generates an appropriate response such as "Hello! What's wrong?", and the emotion engine infers "joy."
[0192] Behavior Generation
[0193] Based on the analysis results, the server determines the appropriate behavior of the character, taking into account the character's personality and the user's emotional information. For example, if the user is happy, the character may respond with a bright smile.
[0194] Generate animation
[0195] The server uses image generation technology to draw animation frames in real time based on the generated behavior, for example, generating 100 animation frames per second of an avatar smiling and responding, "Hello! What's up?"
[0196] Viewing Data
[0197] The generated animation frames are sent from the server to the device and played back in real time on the device. A seamless animation is displayed to the user, realizing natural interaction between the user and the avatar. For example, when the user says "Hello," the avatar responds with a bright smile and says, "Hello! What's wrong?"
[0198] Specific examples
[0199] Consider the case where a user inputs text, "Hello," to an avatar, and simultaneously inputs voice data while smiling. At this time, the device sends the user's input to the server. The server uses LLM to analyze the input "Hello" and infer the user's intention. The emotion engine also infers the user's "happiness" through voice analysis. Based on this, the server generates a behavior for the avatar to respond with a smile, "Hello! What's wrong?" and uses image generation technology to render this as an animation frame. The rendered animation frame is sent to the device and played back in real time on the user's device. In this way, the user can experience more sophisticated and emotional interactions with the avatar.
[0200] The above is an embodiment of the "Touch and Link" system of the present invention, which combines an emotion engine. This system enables users to interact with avatars in a more human-like, natural way that reflects their emotions.
[0201] The processing flow will be explained below.
[0202] Step 1:
[0203] The user types "Hello" into the text input field of the device and presses the send button. The user also uses voice input to say "Hello."
[0204] Step 2:
[0205] The terminal receives the user's text input "Hello" and voice data and sends it to the server.
[0206] Step 3:
[0207] The server receives the text data and voice data sent from the terminal.
[0208] Step 4:
[0209] The server provides text data to a large-scale language model (LLM) and begins analysis.
[0210] Step 5:
[0211] The server receives the analysis results of the LLM and infers the intention from the user's text input. For example, in response to the input "Hello," it generates the response text "Hello! What's wrong?"
[0212] Step 6:
[0213] The server provides the voice data to the emotion engine, which analyzes the user's emotion. The emotion engine infers from the voice data that the user is feeling "joy."
[0214] Step 7:
[0215] When the server receives a touch event, it analyzes it and interprets which part was touched and how. For example, if a user taps on the avatar's head, it obtains that information.
[0216] Step 8:
[0217] The server determines the appropriate behavior of the avatar based on the intent analysis and emotion analysis results. Because the emotion engine estimates the user's "happiness," it determines the behavior of the avatar to respond with a smile, "Hello! What's wrong?"
[0218] Step 9:
[0219] The server uses image generation technology to generate animation frames based on the determined behavior, for example, generating 100 animation frames per second of an avatar smiling.
[0220] Step 10:
[0221] The animation frames generated by the server are divided into packets and sent to the terminal.
[0222] Step 11:
[0223] The device sequentially reads the animation frames received from the server and displays them on the user's screen in real time. For example, an animation of an avatar smoothly smiling and saying, "Hello! What's up?" is displayed.
[0224] The above is a specific flow of program processing that combines emotion engines in the "Touch and Link" system of the present invention.
[0225] Example 2
[0226] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0227] Conventional interactive character systems have had difficulty generating natural responses that accurately reflect the user's emotions and intentions. In particular, they must be able to handle not only text input but also touch events and voice input, and there is a demand for technology that can integrate and analyze these inputs and generate character behavior in real time. The objective of this invention is to provide a system that integrates various user input methods and realizes natural interactions that incorporate emotion analysis.
[0228] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes: means for receiving text input from a user and analyzing the user's intention; means for receiving touch events from the user and interpreting gestures; means for receiving voice input and analyzing the voice data; means for dynamically generating character behavior based on the user's intention using a large-scale language model; means for estimating the user's emotions using emotion analysis technology; means for determining the character's behavior based on the character's personality information and the user's emotion information; means for drawing animation frames in real time based on the behavior generated using image generation technology; and means for transmitting the drawn animation frames to the user's terminal and displaying them. This enables natural interaction that supports various user input means and reflects emotions in real time.
[0229] "Text input" refers to a method in which a user inputs characters using a keyboard, touch panel, or the like.
[0230] A "touch event" refers to a gesture, tap, swipe, or other action performed by a user on a touchscreen.
[0231] "Audio input" is the means by which a user sends voice to the system through a microphone.
[0232] A "large-scale language model" is a natural language processing system trained on vast amounts of text data to analyze user text input.
[0233] "Emotion analysis technology" is a technology that estimates a user's emotions from data such as voice and text.
[0234] "Character personality information" is information about the personality and characteristics set for a virtual character.
[0235] "Character behavior" refers to the actions and expressions made by a virtual character.
[0236] "Image generation technology" is a technology that uses computer graphics technology to generate still images and animations.
[0237] An "animation frame" is a sequence of still images that are displayed in succession to form a moving image.
[0238] A "server" is a computer system that receives data from users and analyzes and processes it.
[0239] A "terminal" is a computer system or device that is directly operated by a user and that performs input and display.
[0240] This invention is a system for realizing natural interactions between users and virtual characters. This system analyzes various user input methods (text input, touch events, and voice input), infers the user's intentions and emotions, and generates character behaviors in real time.
[0241] Hardware and Software
[0242] User Interface
[0243] The device includes a text input field, a touchscreen, and a microphone. This allows the user to interact with the system through text input, touch events, and voice input. For example, the user might type "hello" into the device, tap on the avatar's face, or speak "hello" into the microphone.
[0244] Sending and Receiving Data
[0245] The device transmits text input, touch events, and voice input data from the user to the server in real time using data communication over the Internet or a local network. For example, when the user types "hello," the data is immediately transmitted from the device to the server.
[0246] Data analysis
[0247] The server receives and analyzes various data (text, touch, and voice) sent by the user. It uses a large-scale language model (LLM) to analyze the user's text input and infer their intent. It also uses sentiment analysis technology to infer the user's emotions from voice data and touch events. For example, in response to the text "Hello," the LLM generates a response such as "Hello! What's wrong?", and the sentiment analysis technology infers the user's "joy."
[0248] Behavior Generation
[0249] The server determines the character's behavior based on the analysis results. This determination takes into account the character's personality information and the estimated user's emotional information. For example, if the user is happy, the character is set to respond with a bright smile.
[0250] Generate animation
[0251] The server uses image generation technology to generate animation frames based on the character's behavior in real time. Specifically, technologies such as DALL-E and GAN (Generative Adversarial Networks) are used. For example, it can generate 100 animation frames per second of an avatar smiling and saying, "Hello! What's up?"
[0252] Viewing Data
[0253] The generated animation frames are sent from the server to the device and displayed in real time on the device, allowing users to experience seamless and natural animation. For example, after typing "hello," the user can see an animation of an avatar smiling and replying, "Hello! What's up?"
[0254] Specific examples
[0255] Consider the case where a user uses a device to input text, such as "Hello," and simultaneously input voice data. In this case, the device sends the user's text input and voice data to the server. The server uses a large-scale language model to analyze the input "Hello" and generate an appropriate response, such as "Hello! What's wrong?" At the same time, it uses emotion analysis technology to estimate the user's "happiness" from the voice data. The server then uses this information to generate a behavior for the avatar to respond with a smile, saying "Hello! What's wrong?" and uses image generation technology to render this as an animation frame. The rendered animation frame is sent to the device and displayed in real time on the user's device.
[0256] For example, a prompt might look like this:
[0257] "The user types "hello" into the avatar and speaks to it with a smile using voice input. The device then sends the user's input to the server, which analyzes it and generates an appropriate response and character behavior."
[0258] The above is an embodiment of the invention. This system allows users to have more natural and emotionally appropriate interactions with virtual characters.
[0259] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0260] Step 1:
[0261] The user types "hello" into the device's text input field, taps the avatar's face, or speaks "hello" into the microphone. The device then captures the input data (text, touch, and voice) from the user. Input data such as "hello," as well as the coordinates of the touch event and voice data, are generated.
[0262] Step 2:
[0263] The device transmits the captured user input data to the server via the network. Specifically, the device converts the text, touch events, and voice data entered by the user into packets and transfers them to the server via the network. The output is data in packet format.
[0264] Step 3:
[0265] The server receives the user's input data sent from the device and uses a large-scale language model (LLM) to analyze the user's intent from the text data. For example, the server analyzes the text data "Hello" and generates a response such as "Hello! What's wrong?" The output is the response text as the analysis result.
[0266] Step 4:
[0267] The server analyzes the voice data and touch events using emotion analysis technology to estimate the user's emotion. For example, it analyzes the tone of the voice and the strength of the touch and estimates that the user has the emotion of "joy." The output is the estimated emotion information.
[0268] Step 5:
[0269] The server determines the avatar's behavior based on the intention and emotion analysis results of the LLM, taking into account the character's personality information. For example, if the user is happy, the avatar is set to respond with a bright smile. The output is the generated character's behavior data.
[0270] Step 6:
[0271] The server generates animation frames in real time based on the behavior determined using image generation technology. For example, using technologies such as DALL-E or GAN, it generates 100 animation frames per second in which an avatar smiles and responds, "Hello! What's wrong?" The output is the generated animation frames.
[0272] Step 7:
[0273] The server transmits the generated animation frames to the terminal via the network. At this time, the animation frames are converted into packets and transferred to the terminal via the network. The output is animation frame data in packet format.
[0274] Step 8:
[0275] The device receives the animation frames sent from the server and displays them in real time, allowing the user to see a seamless animation of the avatar. Specifically, an animation of the avatar smiling and saying "Hello! What's up?" is played on the device's display. The output is the displayed animation.
[0276] The above are the specific processing steps of the program for this system. Users can use various input means to naturally interact with the avatar.
[0277] (Application example 2)
[0278] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0279] In modern digital interfaces, interactions between users and virtual agents often lack naturalness. Furthermore, food delivery services lack effective information provision that takes user emotions into account. Therefore, there is a demand for interfaces that are more human-like and can respond to users' emotions.
[0280] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving text input from a user and analyzing the user's intention, means for receiving touch events from the user and interpreting gestures, means for estimating the user's emotion using an emotion engine, means for dynamically generating character behavior based on the user's intention using a large-scale language model (LLM), means for drawing animation frames in real time based on the generated behavior using image generation technology, means for transmitting the drawn animation frames to the user's terminal and displaying them, and means for providing information depending on the user's emotional state. This enables the user to have more natural and emotionally appropriate interactions with the virtual agent.
[0281] "Text input" is the act of a user providing a string of characters to a system using an input device.
[0282] "Analyzing intent" is the process of inferring a user's purpose or request from their text input or touch events.
[0283] A "touch event" is input information generated when a user touches a touchscreen device.
[0284] "Interpreting gestures" means recognizing the user's movements and instructions based on their touch events.
[0285] A "large-scale language model (LLM)" is a natural language processing model trained on massive amounts of text data, and can be applied to a variety of language tasks.
[0286] "Character behavior" refers to the actions and reactions of the virtual agent, which are dynamically generated by the system.
[0287] "Emotion engine" is a general term for algorithms and software that infer emotions from user input information.
[0288] An "animation frame" is an individual image that displays a sequence of character behavior.
[0289] "Drawing in real time" means generating and displaying animation frames instantly in response to user input.
[0290] A "user terminal" is an electronic device used by a user, such as a computer, smartphone, or tablet.
[0291] "Providing information depending on the emotional state" means providing information or suggestions appropriate to the user at any given time, taking into account the user's emotions.
[0292] A system embodying this invention provides natural, emotion-sensitive interactions between a user and a virtual agent by analyzing the user's text input, touch events, and emotions to generate natural responses in real time and display them on the user's device.
[0293] System Configuration
[0294] User Interface
[0295] The user's device has a text entry field, a touch screen, and voice input capabilities. The user interacts with the system by entering text, touching the screen, or communicating by voice. For example, the user enters the text "I want to feel better."
[0296] Sending and Receiving Data
[0297] The device immediately transmits the user's text input, touch events, and voice input to the server using data communication over the network. For example, when a user types "I need energy," the data is sent from the device to the server.
[0298] Data analysis
[0299] The server analyzes the received user text, touch events, and voice data. It uses a large-scale language model (LLM) to infer the user's intent and an emotion engine to infer the user's emotion. For example, the LLM analyzes the text "I need energy" and generates an appropriate response such as "I'm looking for food to give me energy," while the emotion engine infers "I'm tired."
[0300] Behavior Generation
[0301] The server then determines the appropriate behavior for the character based on the analysis results, taking into account the character's personality and the user's emotional information. For example, if the user is feeling tired, the server may suggest food delivery in an encouraging tone.
[0302] Generate animation
[0303] The server uses image generation technology to draw animation frames in real time based on the generated behavior, generating 100 animation frames per second of a virtual agent smiling and responding, for example, "I'm looking for food to cheer me up."
[0304] Viewing Data
[0305] The generated animation frames are sent from the server to the device and played back in real time on the user's device. Seamless, emotional animations are displayed to the user, enabling more natural interactions. For example, if a user types "I want to feel energized," the virtual agent will respond in real time with a smile, saying, "I'm looking for food to cheer you up."
[0306] Examples of specific examples and prompts
[0307] Examples:
[0308] The user inputs the text "I'm tired and want to feel energized." The device then sends the user's input to the server. The server uses a large-scale language model to analyze the input, "I'm tired and want to feel energized," and infers the user's intention. The emotion engine also infers "fatigue" from the input. Based on this, the server generates a behavior for the virtual agent to respond with "I'm looking for food to cheer me up," and uses image generation technology to render this as an animation frame. The rendered frame is then sent to the device and played back in real time on the user's device. The user can experience a sophisticated and emotional interaction with the virtual agent.
[0309] Example prompt sentence:
[0310] Analyze user input, infer emotions using the emotion engine, and generate context-sensitive responses.
[0311] User input: "I'm tired and need some encouragement."
[0312] In this way, the system according to the present invention is able to provide natural and emotional interactions taking into account the user's emotions.
[0313] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0314] Step 1:
[0315] The user inputs text, touch events, or voice input. The user's device receives this input through an input device (keyboard, touch screen, microphone). Specifically, the user inputs text such as "I'm tired and need some energy."
[0316] Step 2:
[0317] The device sends the user's input data to the server. The input data (text, touch information, voice data) is transferred immediately via network communication. The data sent at this time is in JSON format, etc.
[0318] Step 3:
[0319] The server analyzes the received text input using a large-scale language model (LLM). Specifically, the received text data ("I'm tired and want to be cheered up") is input into the LLM, and the server infers the user's intent (wanting to receive cheering suggestions).
[0320] Step 4:
[0321] The server uses an emotion engine to estimate the user's emotion from the input data. The emotion engine detects the emotion "fatigue" from the voice and text. The emotion engine analyzes the emotion using natural language processing models and speech analysis algorithms.
[0322] Step 5:
[0323] The server generates appropriate behavior for the virtual agent based on the user's intentions and emotions. It also takes into account the character's personality information to determine the optimal response. For example, it selects a response such as "Suggest foods that will cheer you up."
[0324] Step 6:
[0325] The server uses image generation technology to draw animation frames based on the virtual agent's behavior, generating 100 animation frames per second of a smiling virtual agent saying, "I'm looking for something to cheer me up."
[0326] Step 7:
[0327] The server then sends the generated animation frames to the user's device. This requires real-time data transfer, and the data is compressed and sent over a high-speed network.
[0328] Step 8:
[0329] The device then plays the received animation frames in real time, providing a seamless animation on the user's device, enabling natural, emotionally-responsive interaction with the virtual agent. The user is shown an animated response with a smile, saying, "I'm looking for some food to cheer me up."
[0330] The above is the flow of the system's program processing and the specific actions performed at each step. This system allows users to experience more natural and emotionally appropriate interactions.
[0331] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0332] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0333] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0334] [Second embodiment]
[0335] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0336] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0337] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0338] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0339] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0340] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0341] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0342] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0343] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0344] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0345] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0346] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0347] This invention is a system for realizing natural interactions between users and virtual agents. This system analyzes user text input and touch events in real time, dynamically generates character behaviors based on the analysis, and displays them as animations.
[0348] System Configuration
[0349] User Interface
[0350] The device has an input field for users to enter text and a touchscreen for detecting touch events, through which users interact with the system. For example, a user can type "hello" or tap on an avatar's face.
[0351] Sending and Receiving Data
[0352] The device immediately sends the text and touch events entered by the user to the server via data communication over the network. For example, when a user types "hello," the data is sent from the device to the server.
[0353] Data analysis
[0354] The server analyzes the received user text and touch events. It uses a large-scale language model (LLM) to infer the user's intent from their input. It also interprets the user's gestures based on the touch events. For example, the LLM can analyze the text "Hello" and generate an appropriate response such as "Hello! What's up?"
[0355] Behavior Generation
[0356] Based on the analysis results, the server generates appropriate behavior for the character, taking into account the character's personality information. For example, a character with a cheerful personality will be set to respond with a smile.
[0357] Generate animation
[0358] The server uses image generation technology to draw animation frames based on the generated behavior in real time, for example, generating 100 animation frames per second of an avatar smiling.
[0359] Viewing Data
[0360] The generated animation frames are sent from the server to the device and played back in real time on the device. A seamless animation is displayed to the user, realizing natural interaction between the user and the avatar. For example, when the user says "Hello," the avatar will respond with a smile, "Hello! What's wrong?"
[0361] Specific examples
[0362] Consider the case where a user inputs the text "Hello" to an avatar. At this time, the device sends the user's input to the server. The server uses LLM to analyze the input "Hello" and infer the user's intention. Based on this, the server generates a behavior for the avatar to respond with a smile, "Hello! What's wrong?" and uses image generation technology to render this as an animation frame. The rendered animation frame is sent to the device and played back in real time on the user's device. In this way, the user can experience natural interaction with the avatar.
[0363] The above is an embodiment of the "Touch and Link" system of the present invention, which allows users to interact with avatars naturally and in real time.
[0364] The processing flow will be explained below.
[0365] Step 1:
[0366] The user enters "Hello" in the text input field of the terminal and presses the send button.
[0367] Step 2:
[0368] The terminal receives the user's text input "Hello" and sends this data to the server.
[0369] Step 3:
[0370] The server receives the text data sent from the terminal.
[0371] Step 4:
[0372] The server provides text data to a large-scale language model (LLM) and begins analysis.
[0373] Step 5:
[0374] The server obtains the analysis results of the LLM and infers the user's intention. For example, if the LLM receives the input "Hello," it generates the response text "Hello! What's wrong?"
[0375] Step 6:
[0376] When the server receives a touch event, it analyzes it and interprets which part was touched and how. For example, if a user taps on the avatar's head, it obtains that information.
[0377] Step 7:
[0378] Based on the intent analysis and interpretation of the touch events, the server determines the appropriate behavior of the avatar, for example, the avatar responding with a smile, "Hello! What's up?"
[0379] Step 8:
[0380] The server uses image generation technology to generate animation frames based on the determined behavior, for example, generating 100 animation frames per second of an avatar smiling.
[0381] Step 9:
[0382] The animation frames generated by the server are divided into packets and sent to the terminal.
[0383] Step 10:
[0384] The device sequentially reads the animation frames received from the server and displays them on the user's screen in real time. For example, an animation of an avatar smoothly smiling and saying, "Hello! What's up?" is displayed.
[0385] The above is the specific program processing flow of the "Touch and Link" system of the present invention.
[0386] Example 1
[0387] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0388] Conventional interaction systems lack the means to realize natural interactions between users and virtual agents, and have particular problems with real-time performance and natural responses. The present invention aims to enhance natural interaction with users by more effectively analyzing user text input and touch events and displaying the behavior of dynamically generated characters in real time.
[0389] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0390] In this invention, the server includes means for receiving text input from a user and analyzing the user's intention, means for receiving touch events from the user and interpreting gestures, means for dynamically generating character behavior from the user's intention using a large-scale language model, means for drawing animation frames in real time based on the generated behavior using an image generation means, and means for transmitting the drawn animation frames to the user's display device and displaying them, thereby enabling natural interaction between the user and the virtual agent in real time.
[0391] "Text input" refers to the act of a user inputting characters and symbols using an input device such as a terminal.
[0392] "Intention" refers to the purpose or desire of what the user wants to communicate to the system or what they are asking for.
[0393] A "touch event" is an input signal that occurs when a user operates a touch screen, and includes gesture operations such as tapping, swiping, and pinching.
[0394] A "gesture" is a hand movement or operation pattern that a user uses when operating a touchscreen or device.
[0395] A "large-scale language model" is a machine learning model that is trained on a huge amount of text data and is used to perform natural language processing.
[0396] "Character behavior" refers to the actions and reactions that a character takes in response to user input and the environment.
[0397] "Image generation means" refers to a method for generating images or visual content using computer technology.
[0398] An "animation frame" is a still image that makes up an animation, and movement is expressed by displaying these images in succession.
[0399] A "display device" is hardware for visually displaying generated visual content, including a screen or display.
[0400] "Real-time" means responding or processing immediately to user input or events without delay.
[0401] The present invention provides a system for realizing natural interactions between a user and a virtual agent. This system analyzes the user's text input and touch events in real time, dynamically generates character behavior based on the analysis, and displays it as animation. Specific embodiments are described below.
[0402] 1. User Interface
[0403] The device has an input field where the user can enter text and a touchscreen that detects touch events, which the user uses to interact with the system. For example, the user can type "hello" or tap on an avatar's face.
[0404] 2. Sending and Receiving Data
[0405] The device immediately sends the text and touch events entered by the user to the server via data communication over the network. For example, when a user enters "hello," the text data is sent from the device to the server.
[0406] 3. Data Analysis
[0407] The server analyzes the received user text and touch events. It uses a large-scale language model (LLM) to infer the user's intent from the input. It also interprets the user's gestures based on the touch events. For example, the text "Hello" is analyzed and an appropriate response is generated: "Hello! What's wrong?"
[0408] 4. Behavior Generation
[0409] The server generates appropriate behavior for the character based on the analysis results. The character's personality information is also taken into consideration. For example, a character with a cheerful personality will be set to respond with a smile, and will respond to user input with a smile saying, "Hello! What's wrong?"
[0410] 5. Generating Animations
[0411] The server uses image generation technology to draw animation frames in real time based on the determined behavior. For example, 100 animation frames per second of an avatar smiling and saying "Hello! What's wrong?"
[0412] 6. Displaying Data
[0413] The generated animation frames are sent from the server to the device and played back in real time. This allows users to experience natural interactions with the avatar through seamless animation. For example, when the user says "Hello," the avatar will respond with a smile, "Hello! What's wrong?"
[0414] Specific examples
[0415] When a user inputs the text "Good morning" to an avatar, the device sends this input data "Good morning" to the server. The server uses LLM to analyze this text, understands the intent of the greeting, and generates a response such as "Good morning! Let's do our best today!". The server then generates an animation of the avatar smiling back based on this response. This animation frame is sent to the device and displayed in real time. The user can feel the avatar's response naturally, and simple interactions can lead to more complex and enjoyable interactions.
[0416] Prompt Sentence Examples
[0417] Below are some example prompts to input to the generative AI model:
[0418] Specific prompt:
[0419] When a user types "Hello," a cheerful avatar will respond with a smile, "Hello! What's up?" Specifically, you'll analyze the text, generate an appropriate response, and then animate it.
[0420] The above is an embodiment of the "Touch and Link" system of the present invention, which allows users to enjoy natural, real-time interaction with avatars.
[0421] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0422] Step 1:
[0423] User interface input
[0424] The device has an input field and a touch screen, which the user uses to input text into the system or to send touch events. For example, the user may type "hello" or tap on the face of an avatar. The input in this case is text data or touch event data.
[0425] Step 2:
[0426] Sending data
[0427] Any text input or touch events made by the user are immediately sent from the device to the server. Here, the input data (text or touch events) is sent via network communication and transmitted to the server in its original format. Specifically, the user's text "Hello" is sent from the device to the server.
[0428] Step 3:
[0429] Data analysis
[0430] The server analyzes the received user input data. It passes the input data (text and touch events) to a large-scale language model (LLM) that performs calculations to infer the user's intent. It also interprets gestures based on touch events. For example, the LLM analyzes the text "Hello" and generates an appropriate response, "Hello! What's wrong?" The analysis results are obtained as output.
[0431] Step 4:
[0432] Behavior Generation
[0433] The server dynamically generates character behavior based on the analysis results. The input here is the analysis results, and the output is specific character behavior information. Taking into account the character's personality information, for example, a character with a cheerful personality is set to respond with a smile.
[0434] Step 5:
[0435] Generate animation
[0436] The server uses image generation technology to draw animation frames based on the generated behavior information. The input is the character's behavior information, and animation frames are generated in real time based on that information. For example, 100 animation frames per second of a character smiling and saying "Hello! What's wrong?" are generated. The output is the generated animation frames.
[0437] Step 6:
[0438] Viewing Data
[0439] The generated animation frames are sent from the server to the device and played back in real time on the device. The input is the generated animation frames, which are sent to the device over the network and displayed seamlessly to the user. For example, in response to the user saying "Hello," an animation of an avatar smiling and saying "Hello! What's wrong?" is played back on the device. The output is the animation displayed on the device.
[0440] (Application example 1)
[0441] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0442] Conventional virtual agent systems have limited interaction with users, making it difficult to have natural conversations. Furthermore, they lack functionality as a shopping assistant, meaning they cannot provide sufficient support when users search for and purchase products. This results in a poor user experience and makes it difficult to promote sales effectively.
[0443] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0444] In this invention, the server includes means for receiving text input from a user and analyzing the user's intention, means for receiving touch events from the user and interpreting gestures, means for dynamically generating character behavior based on the user's intention using a large-scale language model, means for drawing animation frames in real time based on the generated behavior using image generation technology, means for transmitting the drawn animation frames to the user's terminal and displaying them, and means for receiving input from the user for searching for and purchasing products and providing product information based on the input. This enables natural interaction with the user, provides high performance as a shopping assistant, improves the user experience, and enables effective sales promotion.
[0445] "Text input from a user" refers to the act of a user inputting text information using a keyboard or touch screen.
[0446] A "means for analyzing user intent" is a means for analyzing what the text means and what it is trying to convey based on the text entered by the user.
[0447] A "touch event" is an event that occurs when a user taps, swipes, pinches, or performs other actions on a touchscreen.
[0448] The "means for interpreting a gesture" is a means for interpreting what gesture the user has made based on a touch event.
[0449] A "large-scale language model" is a machine learning model trained on massive amounts of text data, capable of understanding and generating natural language.
[0450] "Means for dynamically generating character behavior" refers to means for generating character movements and facial expressions in real time in response to user input and intentions.
[0451] "Image generation technology" is a technology that uses computer graphics to generate visual images and animations.
[0452] "Means for drawing animation frames in real time" refers to means for instantly creating and displaying animation frames based on user input and gestures.
[0453] An "animation frame" is an individual still image that makes up an animation.
[0454] "User terminal" refers to an electronic device used by a user, such as a smartphone, tablet, or PC.
[0455] The "displaying means" is a means for displaying the generated animation frames on the screen of the user's terminal.
[0456] "Input for searching and purchasing products" refers to the input operations performed by a user through a search bar or voice input to search for or purchase products.
[0457] "Means for providing product information" refers to the means for presenting information about products that users wish to search for and purchase.
[0458] "Character personality information" is information about the character's personality that is taken into consideration when determining the character's behavior and responses.
[0459] "Means for generating 100 or more animation frames per second" means means for generating 100 or more animation frames per second, thereby providing seamless animation.
[0460] The present invention is a system that realizes natural interactions with users and acts as a virtual shopping assistant. This system generates character behavior in real time based on user input, and supports product search and purchase.
[0461] System Configuration
[0462] User Interface
[0463] The device includes an input field for users to enter text or voice input, and a touchscreen for detecting touch events such as taps and swipes, through which users interact with the system. For example, a user can type "I want a smartwatch" or tap on a specific product.
[0464] Sending and Receiving Data
[0465] The device immediately sends text, touch events, and voice input from the user to the server via data communication over the network. For example, if a user types "I want a smartwatch," that data is sent from the device to the server.
[0466] Data analysis
[0467] The server analyzes the received user text, touch events, and voice data. It uses a large-scale language model (e.g., OpenAI's GPT-4) to infer the user's intent from their input. It also interprets the user's gestures based on the touch events. For example, a large-scale language model can analyze the text "I want a smartwatch" and generate the optimal response and product information.
[0468] Behavior Generation
[0469] Based on the analysis results, the server generates appropriate behavior for the character, taking into account the character's personality. For example, a character with a cheerful personality might respond with a smile, "What do you think of this smartwatch?"
[0470] Generate animation
[0471] The server uses image generation technology (e.g., Unity, Unreal Engine) to draw animation frames based on the generated behavior in real time. For example, it generates 60 animation frames per second of a virtual assistant pointing at a smartwatch.
[0472] Viewing Data
[0473] The generated animation frames are sent from the server to the device and played back in real time on the device. A seamless animation is displayed to the user, realizing natural interaction between the user and the virtual assistant. For example, if a user says, "I want a smartwatch," an animation is displayed in which the virtual assistant smiles and asks, "How about this smartwatch?"
[0474] Specific examples
[0475] Consider the case where a user texts the assistant, saying, "I want a smartwatch." At this time, the device sends the user's input to the server. The server uses GPT-4 to analyze the input, "I want a smartwatch," and infers the user's intent. Based on this, the server generates a behavior in which the virtual assistant responds with a smile, "How about this smartwatch?" and uses image generation technology to render this as an animation frame. The rendered animation frame is sent to the device and played back in real time on the user's device. In this way, the user experiences natural interactions with the virtual assistant, enabling smooth product search and purchase.
[0476] A specific example of a prompt sentence is as follows:
[0477] A user says, "I want a smartwatch." As a virtual shopping assistant, generate an appropriate response and consider the appropriate avatar behavior for that response.
[0478] This system allows users to interact with the virtual assistant naturally and in real time, making product searches and purchases more convenient.
[0479] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0480] Step 1:
[0481] The user provides text or voice input to the device. The device receives the user's input and sends it to the server as JSON format data. At this time, the input includes text or voice data such as "I want a smartwatch." The processing involves sending the input data to the server via the network. The output is the server receiving the input data.
[0482] Step 2:
[0483] The server analyzes the received text and voice data from the user. During the data analysis process, a large-scale language model (e.g., OpenAI's GPT-4) is used to infer the user's intent from their input. The data processing performed in this step involves text analysis and intent inference. Text data is given as input, and the analysis results are obtained as output. Specifically, the input "I want a smartwatch" is interpreted as "The user is looking for a smartwatch."
[0484] Step 3:
[0485] The server generates appropriate behavior for the virtual character based on the analysis results. In this process, the behavior is determined taking into account the character's personality information. For example, a character with a cheerful personality will respond with a smile. The input in this step is the analysis results, and the output is the generated character behavior. Specifically, the server generates the response "We recommend a smartwatch" and the corresponding character behavior.
[0486] Step 4:
[0487] The server uses image generation technology (e.g. Unity, Unreal Engine) to draw animation frames based on the generated behavior in real time. Character behavior information is given as input, and animation frames are the output. The specific operation in this step is for the animation software to generate frames.
[0488] Step 5:
[0489] The server sends the generated animation frames to the terminal. The input is the animation frames, and the output is the transmission of frames to the terminal. Specifically, the server sends animation data in real time over the network and prepares it for display on the user's terminal.
[0490] Step 6:
[0491] The device displays the received animation frames, providing the user with a seamless interaction experience. The input is the received animation frames, and the output is the display on the user's screen. Specifically, the device plays the animation frames, making it appear as if the virtual character is talking to the user.
[0492] This allows users to receive product information from the virtual shopping assistant through natural interaction.
[0493] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0494] This invention is a system incorporating an emotion analysis engine to realize natural interactions between users and virtual agents. This system analyzes user text input and touch events in real time, dynamically generates character behavior based on the analysis, and displays it as animation. Furthermore, by using the emotion engine to estimate the user's emotions and reflect them in the character's behavior, it realizes more natural, human-like responses.
[0495] System Configuration
[0496] User Interface
[0497] The device has a text entry field, a touchscreen, and voice input. Users can interact with the system through these, for example, by typing "hello" or by tapping on an avatar's face. They can also communicate using voice input.
[0498] Sending and Receiving Data
[0499] The device immediately transmits the user's text input, touch events, and voice input to the server using data communication over the network. For example, when a user types "hello," the data is sent from the device to the server.
[0500] Data analysis
[0501] The server analyzes the received user text, touch events, and voice data. It uses a large-scale language model (LLM) to infer the user's intent from the input. It also uses an emotion engine to infer the user's emotion. For example, the LLM analyzes the text "Hello" and generates an appropriate response such as "Hello! What's wrong?", and the emotion engine infers "joy."
[0502] Behavior Generation
[0503] Based on the analysis results, the server determines the appropriate behavior of the character, taking into account the character's personality and the user's emotional information. For example, if the user is happy, the character may respond with a bright smile.
[0504] Generate animation
[0505] The server uses image generation technology to draw animation frames in real time based on the generated behavior, for example, generating 100 animation frames per second of an avatar smiling and responding, "Hello! What's up?"
[0506] Viewing Data
[0507] The generated animation frames are sent from the server to the device and played back in real time on the device. A seamless animation is displayed to the user, realizing natural interaction between the user and the avatar. For example, when the user says "Hello," the avatar responds with a bright smile and says, "Hello! What's wrong?"
[0508] Specific examples
[0509] Consider the case where a user inputs text, "Hello," to an avatar, and simultaneously inputs voice data while smiling. At this time, the device sends the user's input to the server. The server uses LLM to analyze the input "Hello" and infer the user's intention. The emotion engine also infers the user's "happiness" through voice analysis. Based on this, the server generates a behavior for the avatar to respond with a smile, "Hello! What's wrong?" and uses image generation technology to render this as an animation frame. The rendered animation frame is sent to the device and played back in real time on the user's device. In this way, the user can experience more sophisticated and emotional interactions with the avatar.
[0510] The above is an embodiment of the "Touch and Link" system of the present invention, which combines an emotion engine. This system enables users to interact with avatars in a more human-like, natural way that reflects their emotions.
[0511] The processing flow will be explained below.
[0512] Step 1:
[0513] The user types "Hello" into the text input field of the device and presses the send button. The user also uses voice input to say "Hello."
[0514] Step 2:
[0515] The terminal receives the user's text input "Hello" and voice data and sends it to the server.
[0516] Step 3:
[0517] The server receives the text data and voice data sent from the terminal.
[0518] Step 4:
[0519] The server provides text data to a large-scale language model (LLM) and begins analysis.
[0520] Step 5:
[0521] The server receives the analysis results of the LLM and infers the intention from the user's text input. For example, in response to the input "Hello," it generates the response text "Hello! What's wrong?"
[0522] Step 6:
[0523] The server provides the voice data to the emotion engine, which analyzes the user's emotion. The emotion engine infers from the voice data that the user is feeling "joy."
[0524] Step 7:
[0525] When the server receives a touch event, it analyzes it and interprets which part was touched and how. For example, if a user taps on the avatar's head, it obtains that information.
[0526] Step 8:
[0527] The server determines the appropriate behavior of the avatar based on the intent analysis and emotion analysis results. Because the emotion engine estimates the user's "happiness," it determines the behavior of the avatar to respond with a smile, "Hello! What's wrong?"
[0528] Step 9:
[0529] The server uses image generation technology to generate animation frames based on the determined behavior, for example, generating 100 animation frames per second of an avatar smiling.
[0530] Step 10:
[0531] The animation frames generated by the server are divided into packets and sent to the terminal.
[0532] Step 11:
[0533] The device sequentially reads the animation frames received from the server and displays them on the user's screen in real time. For example, an animation of an avatar smoothly smiling and saying, "Hello! What's up?" is displayed.
[0534] The above is a specific flow of program processing that combines emotion engines in the "Touch and Link" system of the present invention.
[0535] Example 2
[0536] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0537] Conventional interactive character systems have had difficulty generating natural responses that accurately reflect the user's emotions and intentions. In particular, they must be able to handle not only text input but also touch events and voice input, and there is a demand for technology that can integrate and analyze these inputs and generate character behavior in real time. The objective of this invention is to provide a system that integrates various user input methods and realizes natural interactions that incorporate emotion analysis.
[0538] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes: means for receiving text input from a user and analyzing the user's intention; means for receiving touch events from the user and interpreting gestures; means for receiving voice input and analyzing the voice data; means for dynamically generating character behavior based on the user's intention using a large-scale language model; means for estimating the user's emotions using emotion analysis technology; means for determining the character's behavior based on the character's personality information and the user's emotion information; means for drawing animation frames in real time based on the behavior generated using image generation technology; and means for transmitting the drawn animation frames to the user's terminal and displaying them. This enables natural interaction that supports various user input means and reflects emotions in real time.
[0539] "Text input" refers to a method in which a user inputs characters using a keyboard, touch panel, or the like.
[0540] A "touch event" refers to a gesture, tap, swipe, or other action performed by a user on a touchscreen.
[0541] "Audio input" is the means by which a user sends voice to the system through a microphone.
[0542] A "large-scale language model" is a natural language processing system trained on vast amounts of text data to analyze user text input.
[0543] "Emotion analysis technology" is a technology that estimates a user's emotions from data such as voice and text.
[0544] "Character personality information" is information about the personality and characteristics set for a virtual character.
[0545] "Character behavior" refers to the actions and expressions made by a virtual character.
[0546] "Image generation technology" is a technology that uses computer graphics technology to generate still images and animations.
[0547] An "animation frame" is a sequence of still images that are displayed in succession to form a moving image.
[0548] A "server" is a computer system that receives data from users and analyzes and processes it.
[0549] A "terminal" is a computer system or device that is directly operated by a user and that performs input and display.
[0550] This invention is a system for realizing natural interactions between users and virtual characters. This system analyzes various user input methods (text input, touch events, and voice input), infers the user's intentions and emotions, and generates character behaviors in real time.
[0551] Hardware and Software
[0552] User Interface
[0553] The device includes a text input field, a touchscreen, and a microphone. This allows the user to interact with the system through text input, touch events, and voice input. For example, the user might type "hello" into the device, tap on the avatar's face, or speak "hello" into the microphone.
[0554] Sending and Receiving Data
[0555] The device transmits text input, touch events, and voice input data from the user to the server in real time using data communication over the Internet or a local network. For example, when the user types "hello," the data is immediately transmitted from the device to the server.
[0556] Data analysis
[0557] The server receives and analyzes various data (text, touch, and voice) sent by the user. It uses a large-scale language model (LLM) to analyze the user's text input and infer their intent. It also uses sentiment analysis technology to infer the user's emotions from voice data and touch events. For example, in response to the text "Hello," the LLM generates a response such as "Hello! What's wrong?", and the sentiment analysis technology infers the user's "joy."
[0558] Behavior Generation
[0559] The server determines the character's behavior based on the analysis results. This determination takes into account the character's personality information and the estimated user's emotional information. For example, if the user is happy, the character is set to respond with a bright smile.
[0560] Generate animation
[0561] The server uses image generation technology to generate animation frames based on the character's behavior in real time. Specifically, technologies such as DALL-E and GAN (Generative Adversarial Networks) are used. For example, it can generate 100 animation frames per second of an avatar smiling and saying, "Hello! What's up?"
[0562] Viewing Data
[0563] The generated animation frames are sent from the server to the device and displayed in real time on the device, allowing users to experience seamless and natural animation. For example, after typing "hello," the user can see an animation of an avatar smiling and replying, "Hello! What's up?"
[0564] Specific examples
[0565] Consider the case where a user uses a device to input text, such as "Hello," and simultaneously input voice data. In this case, the device sends the user's text input and voice data to the server. The server uses a large-scale language model to analyze the input "Hello" and generate an appropriate response, such as "Hello! What's wrong?" At the same time, it uses emotion analysis technology to estimate the user's "happiness" from the voice data. The server then uses this information to generate a behavior for the avatar to respond with a smile, saying "Hello! What's wrong?" and uses image generation technology to render this as an animation frame. The rendered animation frame is sent to the device and displayed in real time on the user's device.
[0566] For example, a prompt might look like this:
[0567] "The user types "hello" into the avatar and speaks to it with a smile using voice input. The device then sends the user's input to the server, which analyzes it and generates an appropriate response and character behavior."
[0568] The above is an embodiment of the invention. This system allows users to have more natural and emotionally appropriate interactions with virtual characters.
[0569] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0570] Step 1:
[0571] The user types "hello" into the device's text input field, taps the avatar's face, or speaks "hello" into the microphone. The device then captures the input data (text, touch, and voice) from the user. Input data such as "hello," as well as the coordinates of the touch event and voice data, are generated.
[0572] Step 2:
[0573] The device transmits the captured user input data to the server via the network. Specifically, the device converts the text, touch events, and voice data entered by the user into packets and transfers them to the server via the network. The output is data in packet format.
[0574] Step 3:
[0575] The server receives the user's input data sent from the device and uses a large-scale language model (LLM) to analyze the user's intent from the text data. For example, the server analyzes the text data "Hello" and generates a response such as "Hello! What's wrong?" The output is the response text as the analysis result.
[0576] Step 4:
[0577] The server analyzes the voice data and touch events using emotion analysis technology to estimate the user's emotion. For example, it analyzes the tone of the voice and the strength of the touch and estimates that the user has the emotion of "joy." The output is the estimated emotion information.
[0578] Step 5:
[0579] The server determines the avatar's behavior based on the intention and emotion analysis results of the LLM, taking into account the character's personality information. For example, if the user is happy, the avatar is set to respond with a bright smile. The output is the generated character's behavior data.
[0580] Step 6:
[0581] The server generates animation frames in real time based on the behavior determined using image generation technology. For example, using technologies such as DALL-E or GAN, it generates 100 animation frames per second in which an avatar smiles and responds, "Hello! What's wrong?" The output is the generated animation frames.
[0582] Step 7:
[0583] The server transmits the generated animation frames to the terminal via the network. At this time, the animation frames are converted into packets and transferred to the terminal via the network. The output is animation frame data in packet format.
[0584] Step 8:
[0585] The device receives the animation frames sent from the server and displays them in real time, allowing the user to see a seamless animation of the avatar. Specifically, an animation of the avatar smiling and saying "Hello! What's up?" is played on the device's display. The output is the displayed animation.
[0586] The above are the specific processing steps of the program for this system. Users can use various input means to naturally interact with the avatar.
[0587] (Application example 2)
[0588] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0589] In modern digital interfaces, interactions between users and virtual agents often lack naturalness. Furthermore, food delivery services lack effective information provision that takes user emotions into account. Therefore, there is a demand for interfaces that are more human-like and can respond to users' emotions.
[0590] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving text input from a user and analyzing the user's intention, means for receiving touch events from the user and interpreting gestures, means for estimating the user's emotion using an emotion engine, means for dynamically generating character behavior based on the user's intention using a large-scale language model (LLM), means for drawing animation frames in real time based on the generated behavior using image generation technology, means for transmitting the drawn animation frames to the user's terminal and displaying them, and means for providing information depending on the user's emotional state. This enables the user to have more natural and emotionally appropriate interactions with the virtual agent.
[0591] "Text input" is the act of a user providing a string of characters to a system using an input device.
[0592] "Analyzing intent" is the process of inferring a user's purpose or request from their text input or touch events.
[0593] A "touch event" is input information generated when a user touches a touchscreen device.
[0594] "Interpreting gestures" means recognizing the user's movements and instructions based on their touch events.
[0595] A "large-scale language model (LLM)" is a natural language processing model trained on massive amounts of text data, and can be applied to a variety of language tasks.
[0596] "Character behavior" refers to the actions and reactions of the virtual agent, which are dynamically generated by the system.
[0597] "Emotion engine" is a general term for algorithms and software that infer emotions from user input information.
[0598] An "animation frame" is an individual image that displays a sequence of character behavior.
[0599] "Drawing in real time" means generating and displaying animation frames instantly in response to user input.
[0600] A "user terminal" is an electronic device used by a user, such as a computer, smartphone, or tablet.
[0601] "Providing information depending on the emotional state" means providing information or suggestions appropriate to the user at any given time, taking into account the user's emotions.
[0602] A system embodying this invention provides natural, emotion-sensitive interactions between a user and a virtual agent by analyzing the user's text input, touch events, and emotions to generate natural responses in real time and display them on the user's device.
[0603] System Configuration
[0604] User Interface
[0605] The user's device has a text entry field, a touch screen, and voice input capabilities. The user interacts with the system by entering text, touching the screen, or communicating by voice. For example, the user enters the text "I want to feel better."
[0606] Sending and Receiving Data
[0607] The device immediately transmits the user's text input, touch events, and voice input to the server using data communication over the network. For example, when a user types "I need energy," the data is sent from the device to the server.
[0608] Data analysis
[0609] The server analyzes the received user text, touch events, and voice data. It uses a large-scale language model (LLM) to infer the user's intent and an emotion engine to infer the user's emotion. For example, the LLM analyzes the text "I need energy" and generates an appropriate response such as "I'm looking for food to give me energy," while the emotion engine infers "I'm tired."
[0610] Behavior Generation
[0611] The server then determines the appropriate behavior for the character based on the analysis results, taking into account the character's personality and the user's emotional information. For example, if the user is feeling tired, the server may suggest food delivery in an encouraging tone.
[0612] Generate animation
[0613] The server uses image generation technology to draw animation frames in real time based on the generated behavior, generating 100 animation frames per second of a virtual agent smiling and responding, for example, "I'm looking for food to cheer me up."
[0614] Viewing Data
[0615] The generated animation frames are sent from the server to the device and played back in real time on the user's device. Seamless, emotional animations are displayed to the user, enabling more natural interactions. For example, if a user types "I want to feel energized," the virtual agent will respond in real time with a smile, saying, "I'm looking for food to cheer you up."
[0616] Examples of specific examples and prompts
[0617] Examples:
[0618] The user inputs the text "I'm tired and want to feel energized." The device then sends the user's input to the server. The server uses a large-scale language model to analyze the input, "I'm tired and want to feel energized," and infers the user's intention. The emotion engine also infers "fatigue" from the input. Based on this, the server generates a behavior for the virtual agent to respond with "I'm looking for food to cheer me up," and uses image generation technology to render this as an animation frame. The rendered frame is then sent to the device and played back in real time on the user's device. The user can experience a sophisticated and emotional interaction with the virtual agent.
[0619] Example prompt sentence:
[0620] Analyze user input, infer emotions using the emotion engine, and generate context-sensitive responses.
[0621] User input: "I'm tired and need some encouragement."
[0622] In this way, the system according to the present invention is able to provide natural and emotional interactions taking into account the user's emotions.
[0623] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0624] Step 1:
[0625] The user inputs text, touch events, or voice input. The user's device receives this input through an input device (keyboard, touch screen, microphone). Specifically, the user inputs text such as "I'm tired and need some energy."
[0626] Step 2:
[0627] The device sends the user's input data to the server. The input data (text, touch information, voice data) is transferred immediately via network communication. The data sent at this time is in JSON format, etc.
[0628] Step 3:
[0629] The server analyzes the received text input using a large-scale language model (LLM). Specifically, the received text data ("I'm tired and want to be cheered up") is input into the LLM, and the server infers the user's intent (wanting to receive cheering suggestions).
[0630] Step 4:
[0631] The server uses an emotion engine to estimate the user's emotion from the input data. The emotion engine detects the emotion "fatigue" from the voice and text. The emotion engine analyzes the emotion using natural language processing models and speech analysis algorithms.
[0632] Step 5:
[0633] The server generates appropriate behavior for the virtual agent based on the user's intentions and emotions. It also takes into account the character's personality information to determine the optimal response. For example, it selects a response such as "Suggest foods that will cheer you up."
[0634] Step 6:
[0635] The server uses image generation technology to draw animation frames based on the virtual agent's behavior, generating 100 animation frames per second of a smiling virtual agent saying, "I'm looking for something to cheer me up."
[0636] Step 7:
[0637] The server then sends the generated animation frames to the user's device. This requires real-time data transfer, and the data is compressed and sent over a high-speed network.
[0638] Step 8:
[0639] The device then plays the received animation frames in real time, providing a seamless animation on the user's device, enabling natural, emotionally-responsive interaction with the virtual agent. The user is shown an animated response with a smile, saying, "I'm looking for some food to cheer me up."
[0640] The above is the flow of the system's program processing and the specific actions performed at each step. This system allows users to experience more natural and emotionally appropriate interactions.
[0641] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0642] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0643] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0644] [Third embodiment]
[0645] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0646] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0647] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0648] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0649] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0650] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0651] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0652] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0653] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0654] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0655] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0656] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0657] This invention is a system for realizing natural interactions between users and virtual agents. This system analyzes user text input and touch events in real time, dynamically generates character behaviors based on the analysis, and displays them as animations.
[0658] System Configuration
[0659] User Interface
[0660] The device has an input field for users to enter text and a touchscreen for detecting touch events, through which users interact with the system. For example, a user can type "hello" or tap on an avatar's face.
[0661] Sending and Receiving Data
[0662] The device immediately sends the text and touch events entered by the user to the server via data communication over the network. For example, when a user types "hello," the data is sent from the device to the server.
[0663] Data analysis
[0664] The server analyzes the received user text and touch events. It uses a large-scale language model (LLM) to infer the user's intent from their input. It also interprets the user's gestures based on the touch events. For example, the LLM can analyze the text "Hello" and generate an appropriate response such as "Hello! What's up?"
[0665] Behavior Generation
[0666] Based on the analysis results, the server generates appropriate behavior for the character, taking into account the character's personality information. For example, a character with a cheerful personality will be set to respond with a smile.
[0667] Generate animation
[0668] The server uses image generation technology to draw animation frames based on the generated behavior in real time, for example, generating 100 animation frames per second of an avatar smiling.
[0669] Viewing Data
[0670] The generated animation frames are sent from the server to the device and played back in real time on the device. A seamless animation is displayed to the user, realizing natural interaction between the user and the avatar. For example, when the user says "Hello," the avatar will respond with a smile, "Hello! What's wrong?"
[0671] Specific examples
[0672] Consider the case where a user inputs the text "Hello" to an avatar. At this time, the device sends the user's input to the server. The server uses LLM to analyze the input "Hello" and infer the user's intention. Based on this, the server generates a behavior for the avatar to respond with a smile, "Hello! What's wrong?" and uses image generation technology to render this as an animation frame. The rendered animation frame is sent to the device and played back in real time on the user's device. In this way, the user can experience natural interaction with the avatar.
[0673] The above is an embodiment of the "Touch and Link" system of the present invention, which allows users to interact with avatars naturally and in real time.
[0674] The processing flow will be explained below.
[0675] Step 1:
[0676] The user enters "Hello" in the text input field of the terminal and presses the send button.
[0677] Step 2:
[0678] The terminal receives the user's text input "Hello" and sends this data to the server.
[0679] Step 3:
[0680] The server receives the text data sent from the terminal.
[0681] Step 4:
[0682] The server provides text data to a large-scale language model (LLM) and begins analysis.
[0683] Step 5:
[0684] The server obtains the analysis results of the LLM and infers the user's intention. For example, if the LLM receives the input "Hello," it generates the response text "Hello! What's wrong?"
[0685] Step 6:
[0686] When the server receives a touch event, it analyzes it and interprets which part was touched and how. For example, if a user taps on the avatar's head, it obtains that information.
[0687] Step 7:
[0688] Based on the intent analysis and interpretation of the touch events, the server determines the appropriate behavior of the avatar, for example, the avatar responding with a smile, "Hello! What's up?"
[0689] Step 8:
[0690] The server uses image generation technology to generate animation frames based on the determined behavior, for example, generating 100 animation frames per second of an avatar smiling.
[0691] Step 9:
[0692] The animation frames generated by the server are divided into packets and sent to the terminal.
[0693] Step 10:
[0694] The device sequentially reads the animation frames received from the server and displays them on the user's screen in real time. For example, an animation of an avatar smoothly smiling and saying, "Hello! What's up?" is displayed.
[0695] The above is the specific program processing flow of the "Touch and Link" system of the present invention.
[0696] Example 1
[0697] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0698] Conventional interaction systems lack the means to realize natural interactions between users and virtual agents, and have particular problems with real-time performance and natural responses. The present invention aims to enhance natural interaction with users by more effectively analyzing user text input and touch events and displaying the behavior of dynamically generated characters in real time.
[0699] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0700] In this invention, the server includes means for receiving text input from a user and analyzing the user's intention, means for receiving touch events from the user and interpreting gestures, means for dynamically generating character behavior from the user's intention using a large-scale language model, means for drawing animation frames in real time based on the generated behavior using an image generation means, and means for transmitting the drawn animation frames to the user's display device and displaying them, thereby enabling natural interaction between the user and the virtual agent in real time.
[0701] "Text input" refers to the act of a user inputting characters and symbols using an input device such as a terminal.
[0702] "Intention" refers to the purpose or desire of what the user wants to communicate to the system or what they are asking for.
[0703] A "touch event" is an input signal that occurs when a user operates a touch screen, and includes gesture operations such as tapping, swiping, and pinching.
[0704] A "gesture" is a hand movement or operation pattern that a user uses when operating a touchscreen or device.
[0705] A "large-scale language model" is a machine learning model that is trained on a huge amount of text data and is used to perform natural language processing.
[0706] "Character behavior" refers to the actions and reactions that a character takes in response to user input and the environment.
[0707] "Image generation means" refers to a method for generating images or visual content using computer technology.
[0708] An "animation frame" is a still image that makes up an animation, and movement is expressed by displaying these images in succession.
[0709] A "display device" is hardware for visually displaying generated visual content, including a screen or display.
[0710] "Real-time" means responding or processing immediately to user input or events without delay.
[0711] The present invention provides a system for realizing natural interactions between a user and a virtual agent. This system analyzes the user's text input and touch events in real time, dynamically generates character behavior based on the analysis, and displays it as animation. Specific embodiments are described below.
[0712] 1. User Interface
[0713] The device has an input field where the user can enter text and a touchscreen that detects touch events, which the user uses to interact with the system. For example, the user can type "hello" or tap on an avatar's face.
[0714] 2. Sending and Receiving Data
[0715] The device immediately sends the text and touch events entered by the user to the server via data communication over the network. For example, when a user enters "hello," the text data is sent from the device to the server.
[0716] 3. Data Analysis
[0717] The server analyzes the received user text and touch events. It uses a large-scale language model (LLM) to infer the user's intent from the input. It also interprets the user's gestures based on the touch events. For example, the text "Hello" is analyzed and an appropriate response is generated: "Hello! What's wrong?"
[0718] 4. Behavior Generation
[0719] The server generates appropriate behavior for the character based on the analysis results. The character's personality information is also taken into consideration. For example, a character with a cheerful personality will be set to respond with a smile, and will respond to user input with a smile saying, "Hello! What's wrong?"
[0720] 5. Generating Animations
[0721] The server uses image generation technology to draw animation frames in real time based on the determined behavior. For example, 100 animation frames per second of an avatar smiling and saying "Hello! What's wrong?"
[0722] 6. Displaying Data
[0723] The generated animation frames are sent from the server to the device and played back in real time. This allows users to experience natural interactions with the avatar through seamless animation. For example, when the user says "Hello," the avatar will respond with a smile, "Hello! What's wrong?"
[0724] Specific examples
[0725] When a user inputs the text "Good morning" to an avatar, the device sends this input data "Good morning" to the server. The server uses LLM to analyze this text, understands the intent of the greeting, and generates a response such as "Good morning! Let's do our best today!". The server then generates an animation of the avatar smiling back based on this response. This animation frame is sent to the device and displayed in real time. The user can feel the avatar's response naturally, and simple interactions can lead to more complex and enjoyable interactions.
[0726] Prompt Sentence Examples
[0727] Below are some example prompts to input to the generative AI model:
[0728] Specific prompt:
[0729] When a user types "Hello," a cheerful avatar will respond with a smile, "Hello! What's up?" Specifically, you'll analyze the text, generate an appropriate response, and then animate it.
[0730] The above is an embodiment of the "Touch and Link" system of the present invention, which allows users to enjoy natural, real-time interaction with avatars.
[0731] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0732] Step 1:
[0733] User interface input
[0734] The device has an input field and a touch screen, which the user uses to input text into the system or to send touch events. For example, the user may type "hello" or tap on the face of an avatar. The input in this case is text data or touch event data.
[0735] Step 2:
[0736] Sending data
[0737] Any text input or touch events made by the user are immediately sent from the device to the server. Here, the input data (text or touch events) is sent via network communication and transmitted to the server in its original format. Specifically, the user's text "Hello" is sent from the device to the server.
[0738] Step 3:
[0739] Data analysis
[0740] The server analyzes the received user input data. It passes the input data (text and touch events) to a large-scale language model (LLM) that performs calculations to infer the user's intent. It also interprets gestures based on touch events. For example, the LLM analyzes the text "Hello" and generates an appropriate response, "Hello! What's wrong?" The analysis results are obtained as output.
[0741] Step 4:
[0742] Behavior Generation
[0743] The server dynamically generates character behavior based on the analysis results. The input here is the analysis results, and the output is specific character behavior information. Taking into account the character's personality information, for example, a character with a cheerful personality is set to respond with a smile.
[0744] Step 5:
[0745] Generate animation
[0746] The server uses image generation technology to draw animation frames based on the generated behavior information. The input is the character's behavior information, and animation frames are generated in real time based on that information. For example, 100 animation frames per second of a character smiling and saying "Hello! What's wrong?" are generated. The output is the generated animation frames.
[0747] Step 6:
[0748] Viewing Data
[0749] The generated animation frames are sent from the server to the device and played back in real time on the device. The input is the generated animation frames, which are sent to the device over the network and displayed seamlessly to the user. For example, in response to the user saying "Hello," an animation of an avatar smiling and saying "Hello! What's wrong?" is played back on the device. The output is the animation displayed on the device.
[0750] (Application example 1)
[0751] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0752] Conventional virtual agent systems have limited interaction with users, making it difficult to have natural conversations. Furthermore, they lack functionality as a shopping assistant, meaning they cannot provide sufficient support when users search for and purchase products. This results in a poor user experience and makes it difficult to promote sales effectively.
[0753] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0754] In this invention, the server includes means for receiving text input from a user and analyzing the user's intention, means for receiving touch events from the user and interpreting gestures, means for dynamically generating character behavior based on the user's intention using a large-scale language model, means for drawing animation frames in real time based on the generated behavior using image generation technology, means for transmitting the drawn animation frames to the user's terminal and displaying them, and means for receiving input from the user for searching for and purchasing products and providing product information based on the input. This enables natural interaction with the user, provides high performance as a shopping assistant, improves the user experience, and enables effective sales promotion.
[0755] "Text input from a user" refers to the act of a user inputting text information using a keyboard or touch screen.
[0756] A "means for analyzing user intent" is a means for analyzing what the text means and what it is trying to convey based on the text entered by the user.
[0757] A "touch event" is an event that occurs when a user taps, swipes, pinches, or performs other actions on a touchscreen.
[0758] The "means for interpreting a gesture" is a means for interpreting what gesture the user has made based on a touch event.
[0759] A "large-scale language model" is a machine learning model trained on massive amounts of text data, capable of understanding and generating natural language.
[0760] "Means for dynamically generating character behavior" refers to means for generating character movements and facial expressions in real time in response to user input and intentions.
[0761] "Image generation technology" is a technology that uses computer graphics to generate visual images and animations.
[0762] "Means for drawing animation frames in real time" refers to means for instantly creating and displaying animation frames based on user input and gestures.
[0763] An "animation frame" is an individual still image that makes up an animation.
[0764] "User terminal" refers to an electronic device used by a user, such as a smartphone, tablet, or PC.
[0765] The "displaying means" is a means for displaying the generated animation frames on the screen of the user's terminal.
[0766] "Input for searching and purchasing products" refers to the input operations performed by a user through a search bar or voice input to search for or purchase products.
[0767] "Means for providing product information" refers to the means for presenting information about products that users wish to search for and purchase.
[0768] "Character personality information" is information about the character's personality that is taken into consideration when determining the character's behavior and responses.
[0769] "Means for generating 100 or more animation frames per second" means means for generating 100 or more animation frames per second, thereby providing seamless animation.
[0770] The present invention is a system that realizes natural interactions with users and acts as a virtual shopping assistant. This system generates character behavior in real time based on user input, and supports product search and purchase.
[0771] System Configuration
[0772] User Interface
[0773] The device includes an input field for users to enter text or voice input, and a touchscreen for detecting touch events such as taps and swipes, through which users interact with the system. For example, a user can type "I want a smartwatch" or tap on a specific product.
[0774] Sending and Receiving Data
[0775] The device immediately sends text, touch events, and voice input from the user to the server via data communication over the network. For example, if a user types "I want a smartwatch," that data is sent from the device to the server.
[0776] Data analysis
[0777] The server analyzes the received user text, touch events, and voice data. It uses a large-scale language model (e.g., OpenAI's GPT-4) to infer the user's intent from their input. It also interprets the user's gestures based on the touch events. For example, a large-scale language model can analyze the text "I want a smartwatch" and generate the optimal response and product information.
[0778] Behavior Generation
[0779] Based on the analysis results, the server generates appropriate behavior for the character, taking into account the character's personality. For example, a character with a cheerful personality might respond with a smile, "What do you think of this smartwatch?"
[0780] Generate animation
[0781] The server uses image generation technology (e.g., Unity, Unreal Engine) to draw animation frames based on the generated behavior in real time. For example, it generates 60 animation frames per second of a virtual assistant pointing at a smartwatch.
[0782] Viewing Data
[0783] The generated animation frames are sent from the server to the device and played back in real time on the device. A seamless animation is displayed to the user, realizing natural interaction between the user and the virtual assistant. For example, if a user says, "I want a smartwatch," an animation is displayed in which the virtual assistant smiles and asks, "How about this smartwatch?"
[0784] Specific examples
[0785] Consider the case where a user texts the assistant, saying, "I want a smartwatch." At this time, the device sends the user's input to the server. The server uses GPT-4 to analyze the input, "I want a smartwatch," and infers the user's intent. Based on this, the server generates a behavior in which the virtual assistant responds with a smile, "How about this smartwatch?" and uses image generation technology to render this as an animation frame. The rendered animation frame is sent to the device and played back in real time on the user's device. In this way, the user experiences natural interactions with the virtual assistant, enabling smooth product search and purchase.
[0786] A specific example of a prompt sentence is as follows:
[0787] A user says, "I want a smartwatch." As a virtual shopping assistant, generate an appropriate response and consider the appropriate avatar behavior for that response.
[0788] This system allows users to interact with the virtual assistant naturally and in real time, making product searches and purchases more convenient.
[0789] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0790] Step 1:
[0791] The user provides text or voice input to the device. The device receives the user's input and sends it to the server as JSON format data. At this time, the input includes text or voice data such as "I want a smartwatch." The processing involves sending the input data to the server via the network. The output is the server receiving the input data.
[0792] Step 2:
[0793] The server analyzes the received text and voice data from the user. During the data analysis process, a large-scale language model (e.g., OpenAI's GPT-4) is used to infer the user's intent from their input. The data processing performed in this step involves text analysis and intent inference. Text data is given as input, and the analysis results are obtained as output. Specifically, the input "I want a smartwatch" is interpreted as "The user is looking for a smartwatch."
[0794] Step 3:
[0795] The server generates appropriate behavior for the virtual character based on the analysis results. In this process, the behavior is determined taking into account the character's personality information. For example, a character with a cheerful personality will respond with a smile. The input in this step is the analysis results, and the output is the generated character behavior. Specifically, the server generates the response "We recommend a smartwatch" and the corresponding character behavior.
[0796] Step 4:
[0797] The server uses image generation technology (e.g. Unity, Unreal Engine) to draw animation frames based on the generated behavior in real time. Character behavior information is given as input, and animation frames are the output. The specific operation in this step is for the animation software to generate frames.
[0798] Step 5:
[0799] The server sends the generated animation frames to the terminal. The input is the animation frames, and the output is the transmission of frames to the terminal. Specifically, the server sends animation data in real time over the network and prepares it for display on the user's terminal.
[0800] Step 6:
[0801] The device displays the received animation frames, providing the user with a seamless interaction experience. The input is the received animation frames, and the output is the display on the user's screen. Specifically, the device plays the animation frames, making it appear as if the virtual character is talking to the user.
[0802] This allows users to receive product information from the virtual shopping assistant through natural interaction.
[0803] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0804] This invention is a system incorporating an emotion analysis engine to realize natural interactions between users and virtual agents. This system analyzes user text input and touch events in real time, dynamically generates character behavior based on the analysis, and displays it as animation. Furthermore, by using the emotion engine to estimate the user's emotions and reflect them in the character's behavior, it realizes more natural, human-like responses.
[0805] System Configuration
[0806] User Interface
[0807] The device has a text entry field, a touchscreen, and voice input. Users can interact with the system through these, for example, by typing "hello" or by tapping on an avatar's face. They can also communicate using voice input.
[0808] Sending and Receiving Data
[0809] The device immediately transmits the user's text input, touch events, and voice input to the server using data communication over the network. For example, when a user types "hello," the data is sent from the device to the server.
[0810] Data analysis
[0811] The server analyzes the received user text, touch events, and voice data. It uses a large-scale language model (LLM) to infer the user's intent from the input. It also uses an emotion engine to infer the user's emotion. For example, the LLM analyzes the text "Hello" and generates an appropriate response such as "Hello! What's wrong?", and the emotion engine infers "joy."
[0812] Behavior Generation
[0813] Based on the analysis results, the server determines the appropriate behavior of the character, taking into account the character's personality and the user's emotional information. For example, if the user is happy, the character may respond with a bright smile.
[0814] Generate animation
[0815] The server uses image generation technology to draw animation frames in real time based on the generated behavior, for example, generating 100 animation frames per second of an avatar smiling and responding, "Hello! What's up?"
[0816] Viewing Data
[0817] The generated animation frames are sent from the server to the device and played back in real time on the device. A seamless animation is displayed to the user, realizing natural interaction between the user and the avatar. For example, when the user says "Hello," the avatar responds with a bright smile and says, "Hello! What's wrong?"
[0818] Specific examples
[0819] Consider the case where a user inputs text, "Hello," to an avatar, and simultaneously inputs voice data while smiling. At this time, the device sends the user's input to the server. The server uses LLM to analyze the input "Hello" and infer the user's intention. The emotion engine also infers the user's "happiness" through voice analysis. Based on this, the server generates a behavior for the avatar to respond with a smile, "Hello! What's wrong?" and uses image generation technology to render this as an animation frame. The rendered animation frame is sent to the device and played back in real time on the user's device. In this way, the user can experience more sophisticated and emotional interactions with the avatar.
[0820] The above is an embodiment of the "Touch and Link" system of the present invention, which combines an emotion engine. This system enables users to interact with avatars in a more human-like, natural way that reflects their emotions.
[0821] The processing flow will be explained below.
[0822] Step 1:
[0823] The user types "Hello" into the text input field of the device and presses the send button. The user also uses voice input to say "Hello."
[0824] Step 2:
[0825] The terminal receives the user's text input "Hello" and voice data and sends it to the server.
[0826] Step 3:
[0827] The server receives the text data and voice data sent from the terminal.
[0828] Step 4:
[0829] The server provides text data to a large-scale language model (LLM) and begins analysis.
[0830] Step 5:
[0831] The server receives the analysis results of the LLM and infers the intention from the user's text input. For example, in response to the input "Hello," it generates the response text "Hello! What's wrong?"
[0832] Step 6:
[0833] The server provides the voice data to the emotion engine, which analyzes the user's emotion. The emotion engine infers from the voice data that the user is feeling "joy."
[0834] Step 7:
[0835] When the server receives a touch event, it analyzes it and interprets which part was touched and how. For example, if a user taps on the avatar's head, it obtains that information.
[0836] Step 8:
[0837] The server determines the appropriate behavior of the avatar based on the intent analysis and emotion analysis results. Because the emotion engine estimates the user's "happiness," it determines the behavior of the avatar to respond with a smile, "Hello! What's wrong?"
[0838] Step 9:
[0839] The server uses image generation technology to generate animation frames based on the determined behavior, for example, generating 100 animation frames per second of an avatar smiling.
[0840] Step 10:
[0841] The animation frames generated by the server are divided into packets and sent to the terminal.
[0842] Step 11:
[0843] The device sequentially reads the animation frames received from the server and displays them on the user's screen in real time. For example, an animation of an avatar smoothly smiling and saying, "Hello! What's up?" is displayed.
[0844] The above is a specific flow of program processing that combines emotion engines in the "Touch and Link" system of the present invention.
[0845] Example 2
[0846] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0847] Conventional interactive character systems have had difficulty generating natural responses that accurately reflect the user's emotions and intentions. In particular, they must be able to handle not only text input but also touch events and voice input, and there is a demand for technology that can integrate and analyze these inputs and generate character behavior in real time. The objective of this invention is to provide a system that integrates various user input methods and realizes natural interactions that incorporate emotion analysis.
[0848] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes: means for receiving text input from a user and analyzing the user's intention; means for receiving touch events from the user and interpreting gestures; means for receiving voice input and analyzing the voice data; means for dynamically generating character behavior based on the user's intention using a large-scale language model; means for estimating the user's emotions using emotion analysis technology; means for determining the character's behavior based on the character's personality information and the user's emotion information; means for drawing animation frames in real time based on the behavior generated using image generation technology; and means for transmitting the drawn animation frames to the user's terminal and displaying them. This enables natural interaction that supports various user input means and reflects emotions in real time.
[0849] "Text input" refers to a method in which a user inputs characters using a keyboard, touch panel, or the like.
[0850] A "touch event" refers to a gesture, tap, swipe, or other action performed by a user on a touchscreen.
[0851] "Audio input" is the means by which a user sends voice to the system through a microphone.
[0852] A "large-scale language model" is a natural language processing system trained on vast amounts of text data to analyze user text input.
[0853] "Emotion analysis technology" is a technology that estimates a user's emotions from data such as voice and text.
[0854] "Character personality information" is information about the personality and characteristics set for a virtual character.
[0855] "Character behavior" refers to the actions and expressions made by a virtual character.
[0856] "Image generation technology" is a technology that uses computer graphics technology to generate still images and animations.
[0857] An "animation frame" is a sequence of still images that are displayed in succession to form a moving image.
[0858] A "server" is a computer system that receives data from users and analyzes and processes it.
[0859] A "terminal" is a computer system or device that is directly operated by a user and that performs input and display.
[0860] This invention is a system for realizing natural interactions between users and virtual characters. This system analyzes various user input methods (text input, touch events, and voice input), infers the user's intentions and emotions, and generates character behaviors in real time.
[0861] Hardware and Software
[0862] User Interface
[0863] The device includes a text input field, a touchscreen, and a microphone. This allows the user to interact with the system through text input, touch events, and voice input. For example, the user might type "hello" into the device, tap on the avatar's face, or speak "hello" into the microphone.
[0864] Sending and Receiving Data
[0865] The device transmits text input, touch events, and voice input data from the user to the server in real time using data communication over the Internet or a local network. For example, when the user types "hello," the data is immediately transmitted from the device to the server.
[0866] Data analysis
[0867] The server receives and analyzes various data (text, touch, and voice) sent by the user. It uses a large-scale language model (LLM) to analyze the user's text input and infer their intent. It also uses sentiment analysis technology to infer the user's emotions from voice data and touch events. For example, in response to the text "Hello," the LLM generates a response such as "Hello! What's wrong?", and the sentiment analysis technology infers the user's "joy."
[0868] Behavior Generation
[0869] The server determines the character's behavior based on the analysis results. This determination takes into account the character's personality information and the estimated user's emotional information. For example, if the user is happy, the character is set to respond with a bright smile.
[0870] Generate animation
[0871] The server uses image generation technology to generate animation frames based on the character's behavior in real time. Specifically, technologies such as DALL-E and GAN (Generative Adversarial Networks) are used. For example, it can generate 100 animation frames per second of an avatar smiling and saying, "Hello! What's up?"
[0872] Viewing Data
[0873] The generated animation frames are sent from the server to the device and displayed in real time on the device, allowing users to experience seamless and natural animation. For example, after typing "hello," the user can see an animation of an avatar smiling and replying, "Hello! What's up?"
[0874] Specific examples
[0875] Consider the case where a user uses a device to input text, such as "Hello," and simultaneously input voice data. In this case, the device sends the user's text input and voice data to the server. The server uses a large-scale language model to analyze the input "Hello" and generate an appropriate response, such as "Hello! What's wrong?" At the same time, it uses emotion analysis technology to estimate the user's "happiness" from the voice data. The server then uses this information to generate a behavior for the avatar to respond with a smile, saying "Hello! What's wrong?" and uses image generation technology to render this as an animation frame. The rendered animation frame is sent to the device and displayed in real time on the user's device.
[0876] For example, a prompt might look like this:
[0877] "The user types "hello" into the avatar and speaks to it with a smile using voice input. The device then sends the user's input to the server, which analyzes it and generates an appropriate response and character behavior."
[0878] The above is an embodiment of the invention. This system allows users to have more natural and emotionally appropriate interactions with virtual characters.
[0879] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0880] Step 1:
[0881] The user types "hello" into the device's text input field, taps the avatar's face, or speaks "hello" into the microphone. The device then captures the input data (text, touch, and voice) from the user. Input data such as "hello," as well as the coordinates of the touch event and voice data, are generated.
[0882] Step 2:
[0883] The device transmits the captured user input data to the server via the network. Specifically, the device converts the text, touch events, and voice data entered by the user into packets and transfers them to the server via the network. The output is data in packet format.
[0884] Step 3:
[0885] The server receives the user's input data sent from the device and uses a large-scale language model (LLM) to analyze the user's intent from the text data. For example, the server analyzes the text data "Hello" and generates a response such as "Hello! What's wrong?" The output is the response text as the analysis result.
[0886] Step 4:
[0887] The server analyzes the voice data and touch events using emotion analysis technology to estimate the user's emotion. For example, it analyzes the tone of the voice and the strength of the touch and estimates that the user has the emotion of "joy." The output is the estimated emotion information.
[0888] Step 5:
[0889] The server determines the avatar's behavior based on the intention and emotion analysis results of the LLM, taking into account the character's personality information. For example, if the user is happy, the avatar is set to respond with a bright smile. The output is the generated character's behavior data.
[0890] Step 6:
[0891] The server generates animation frames in real time based on the behavior determined using image generation technology. For example, using technologies such as DALL-E or GAN, it generates 100 animation frames per second in which an avatar smiles and responds, "Hello! What's wrong?" The output is the generated animation frames.
[0892] Step 7:
[0893] The server transmits the generated animation frames to the terminal via the network. At this time, the animation frames are converted into packets and transferred to the terminal via the network. The output is animation frame data in packet format.
[0894] Step 8:
[0895] The device receives the animation frames sent from the server and displays them in real time, allowing the user to see a seamless animation of the avatar. Specifically, an animation of the avatar smiling and saying "Hello! What's up?" is played on the device's display. The output is the displayed animation.
[0896] The above are the specific processing steps of the program for this system. Users can use various input means to naturally interact with the avatar.
[0897] (Application example 2)
[0898] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0899] In modern digital interfaces, interactions between users and virtual agents often lack naturalness. Furthermore, food delivery services lack effective information provision that takes user emotions into account. Therefore, there is a demand for interfaces that are more human-like and can respond to users' emotions.
[0900] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving text input from a user and analyzing the user's intention, means for receiving touch events from the user and interpreting gestures, means for estimating the user's emotion using an emotion engine, means for dynamically generating character behavior based on the user's intention using a large-scale language model (LLM), means for drawing animation frames in real time based on the generated behavior using image generation technology, means for transmitting the drawn animation frames to the user's terminal and displaying them, and means for providing information depending on the user's emotional state. This enables the user to have more natural and emotionally appropriate interactions with the virtual agent.
[0901] "Text input" is the act of a user providing a string of characters to a system using an input device.
[0902] "Analyzing intent" is the process of inferring a user's purpose or request from their text input or touch events.
[0903] A "touch event" is input information generated when a user touches a touchscreen device.
[0904] "Interpreting gestures" means recognizing the user's movements and instructions based on their touch events.
[0905] A "large-scale language model (LLM)" is a natural language processing model trained on massive amounts of text data, and can be applied to a variety of language tasks.
[0906] "Character behavior" refers to the actions and reactions of the virtual agent, which are dynamically generated by the system.
[0907] "Emotion engine" is a general term for algorithms and software that infer emotions from user input information.
[0908] An "animation frame" is an individual image that displays a sequence of character behavior.
[0909] "Drawing in real time" means generating and displaying animation frames instantly in response to user input.
[0910] A "user terminal" is an electronic device used by a user, such as a computer, smartphone, or tablet.
[0911] "Providing information depending on the emotional state" means providing information or suggestions appropriate to the user at any given time, taking into account the user's emotions.
[0912] A system embodying this invention provides natural, emotion-sensitive interactions between a user and a virtual agent by analyzing the user's text input, touch events, and emotions to generate natural responses in real time and display them on the user's device.
[0913] System Configuration
[0914] User Interface
[0915] The user's device has a text entry field, a touch screen, and voice input capabilities. The user interacts with the system by entering text, touching the screen, or communicating by voice. For example, the user enters the text "I want to feel better."
[0916] Sending and Receiving Data
[0917] The device immediately transmits the user's text input, touch events, and voice input to the server using data communication over the network. For example, when a user types "I need energy," the data is sent from the device to the server.
[0918] Data analysis
[0919] The server analyzes the received user text, touch events, and voice data. It uses a large-scale language model (LLM) to infer the user's intent and an emotion engine to infer the user's emotion. For example, the LLM analyzes the text "I need energy" and generates an appropriate response such as "I'm looking for food to give me energy," while the emotion engine infers "I'm tired."
[0920] Behavior Generation
[0921] The server then determines the appropriate behavior for the character based on the analysis results, taking into account the character's personality and the user's emotional information. For example, if the user is feeling tired, the server may suggest food delivery in an encouraging tone.
[0922] Generate animation
[0923] The server uses image generation technology to draw animation frames in real time based on the generated behavior, generating 100 animation frames per second of a virtual agent smiling and responding, for example, "I'm looking for food to cheer me up."
[0924] Viewing Data
[0925] The generated animation frames are sent from the server to the device and played back in real time on the user's device. Seamless, emotional animations are displayed to the user, enabling more natural interactions. For example, if a user types "I want to feel energized," the virtual agent will respond in real time with a smile, saying, "I'm looking for food to cheer you up."
[0926] Examples of specific examples and prompts
[0927] Examples:
[0928] The user inputs the text "I'm tired and want to feel energized." The device then sends the user's input to the server. The server uses a large-scale language model to analyze the input, "I'm tired and want to feel energized," and infers the user's intention. The emotion engine also infers "fatigue" from the input. Based on this, the server generates a behavior for the virtual agent to respond with "I'm looking for food to cheer me up," and uses image generation technology to render this as an animation frame. The rendered frame is then sent to the device and played back in real time on the user's device. The user can experience a sophisticated and emotional interaction with the virtual agent.
[0929] Example prompt sentence:
[0930] Analyze user input, infer emotions using the emotion engine, and generate context-sensitive responses.
[0931] User input: "I'm tired and need some encouragement."
[0932] In this way, the system according to the present invention is able to provide natural and emotional interactions taking into account the user's emotions.
[0933] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0934] Step 1:
[0935] The user inputs text, touch events, or voice input. The user's device receives this input through an input device (keyboard, touch screen, microphone). Specifically, the user inputs text such as "I'm tired and need some energy."
[0936] Step 2:
[0937] The device sends the user's input data to the server. The input data (text, touch information, voice data) is transferred immediately via network communication. The data sent at this time is in JSON format, etc.
[0938] Step 3:
[0939] The server analyzes the received text input using a large-scale language model (LLM). Specifically, the received text data ("I'm tired and want to be cheered up") is input into the LLM, and the server infers the user's intent (wanting to receive cheering suggestions).
[0940] Step 4:
[0941] The server uses an emotion engine to estimate the user's emotion from the input data. The emotion engine detects the emotion "fatigue" from the voice and text. The emotion engine analyzes the emotion using natural language processing models and speech analysis algorithms.
[0942] Step 5:
[0943] The server generates appropriate behavior for the virtual agent based on the user's intentions and emotions. It also takes into account the character's personality information to determine the optimal response. For example, it selects a response such as "Suggest foods that will cheer you up."
[0944] Step 6:
[0945] The server uses image generation technology to draw animation frames based on the virtual agent's behavior, generating 100 animation frames per second of a smiling virtual agent saying, "I'm looking for something to cheer me up."
[0946] Step 7:
[0947] The server then sends the generated animation frames to the user's device. This requires real-time data transfer, and the data is compressed and sent over a high-speed network.
[0948] Step 8:
[0949] The device then plays the received animation frames in real time, providing a seamless animation on the user's device, enabling natural, emotionally-responsive interaction with the virtual agent. The user is shown an animated response with a smile, saying, "I'm looking for some food to cheer me up."
[0950] The above is the flow of the system's program processing and the specific actions performed at each step. This system allows users to experience more natural and emotionally appropriate interactions.
[0951] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0952] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0953] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[0954] [Fourth embodiment]
[0955] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[0956] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0957] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0958] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[0959] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0960] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0961] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0962] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[0963] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0964] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0965] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0966] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0967] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0968] This invention is a system for realizing natural interactions between users and virtual agents. This system analyzes user text input and touch events in real time, dynamically generates character behaviors based on the analysis, and displays them as animations.
[0969] System Configuration
[0970] User Interface
[0971] The device has an input field for users to enter text and a touchscreen for detecting touch events, through which users interact with the system. For example, a user can type "hello" or tap on an avatar's face.
[0972] Sending and Receiving Data
[0973] The device immediately sends the text and touch events entered by the user to the server via data communication over the network. For example, when a user types "hello," the data is sent from the device to the server.
[0974] Data analysis
[0975] The server analyzes the received user text and touch events. It uses a large-scale language model (LLM) to infer the user's intent from their input. It also interprets the user's gestures based on the touch events. For example, the LLM can analyze the text "Hello" and generate an appropriate response such as "Hello! What's up?"
[0976] Behavior Generation
[0977] Based on the analysis results, the server generates appropriate behavior for the character, taking into account the character's personality information. For example, a character with a cheerful personality will be set to respond with a smile.
[0978] Generate animation
[0979] The server uses image generation technology to draw animation frames based on the generated behavior in real time, for example, generating 100 animation frames per second of an avatar smiling.
[0980] Viewing Data
[0981] The generated animation frames are sent from the server to the device and played back in real time on the device. A seamless animation is displayed to the user, realizing natural interaction between the user and the avatar. For example, when the user says "Hello," the avatar will respond with a smile, "Hello! What's wrong?"
[0982] Specific examples
[0983] Consider the case where a user inputs the text "Hello" to an avatar. At this time, the device sends the user's input to the server. The server uses LLM to analyze the input "Hello" and infer the user's intention. Based on this, the server generates a behavior for the avatar to respond with a smile, "Hello! What's wrong?" and uses image generation technology to render this as an animation frame. The rendered animation frame is sent to the device and played back in real time on the user's device. In this way, the user can experience natural interaction with the avatar.
[0984] The above is an embodiment of the "Touch and Link" system of the present invention, which allows users to interact with avatars naturally and in real time.
[0985] The processing flow will be explained below.
[0986] Step 1:
[0987] The user enters "Hello" in the text input field of the terminal and presses the send button.
[0988] Step 2:
[0989] The terminal receives the user's text input "Hello" and sends this data to the server.
[0990] Step 3:
[0991] The server receives the text data sent from the terminal.
[0992] Step 4:
[0993] The server provides text data to a large-scale language model (LLM) and begins analysis.
[0994] Step 5:
[0995] The server obtains the analysis results of the LLM and infers the user's intention. For example, if the LLM receives the input "Hello," it generates the response text "Hello! What's wrong?"
[0996] Step 6:
[0997] When the server receives a touch event, it analyzes it and interprets which part was touched and how. For example, if a user taps on the avatar's head, it obtains that information.
[0998] Step 7:
[0999] Based on the intent analysis and interpretation of the touch events, the server determines the appropriate behavior of the avatar, for example, the avatar responding with a smile, "Hello! What's up?"
[1000] Step 8:
[1001] The server uses image generation technology to generate animation frames based on the determined behavior, for example, generating 100 animation frames per second of an avatar smiling.
[1002] Step 9:
[1003] The animation frames generated by the server are divided into packets and sent to the terminal.
[1004] Step 10:
[1005] The device sequentially reads the animation frames received from the server and displays them on the user's screen in real time. For example, an animation of an avatar smoothly smiling and saying, "Hello! What's up?" is displayed.
[1006] The above is the specific program processing flow of the "Touch and Link" system of the present invention.
[1007] Example 1
[1008] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1009] Conventional interaction systems lack the means to realize natural interactions between users and virtual agents, and have particular problems with real-time performance and natural responses. The present invention aims to enhance natural interaction with users by more effectively analyzing user text input and touch events and displaying the behavior of dynamically generated characters in real time.
[1010] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1011] In this invention, the server includes means for receiving text input from a user and analyzing the user's intention, means for receiving touch events from the user and interpreting gestures, means for dynamically generating character behavior from the user's intention using a large-scale language model, means for drawing animation frames in real time based on the generated behavior using an image generation means, and means for transmitting the drawn animation frames to the user's display device and displaying them, thereby enabling natural interaction between the user and the virtual agent in real time.
[1012] "Text input" refers to the act of a user inputting characters and symbols using an input device such as a terminal.
[1013] "Intention" refers to the purpose or desire of what the user wants to communicate to the system or what they are asking for.
[1014] A "touch event" is an input signal that occurs when a user operates a touch screen, and includes gesture operations such as tapping, swiping, and pinching.
[1015] A "gesture" is a hand movement or operation pattern that a user uses when operating a touchscreen or device.
[1016] A "large-scale language model" is a machine learning model that is trained on a huge amount of text data and is used to perform natural language processing.
[1017] "Character behavior" refers to the actions and reactions that a character takes in response to user input and the environment.
[1018] "Image generation means" refers to a method for generating images or visual content using computer technology.
[1019] An "animation frame" is a still image that makes up an animation, and movement is expressed by displaying these images in succession.
[1020] A "display device" is hardware for visually displaying generated visual content, including a screen or display.
[1021] "Real-time" means responding or processing immediately to user input or events without delay.
[1022] The present invention provides a system for realizing natural interactions between a user and a virtual agent. This system analyzes the user's text input and touch events in real time, dynamically generates character behavior based on the analysis, and displays it as animation. Specific embodiments are described below.
[1023] 1. User Interface
[1024] The device has an input field where the user can enter text and a touchscreen that detects touch events, which the user uses to interact with the system. For example, the user can type "hello" or tap on an avatar's face.
[1025] 2. Sending and Receiving Data
[1026] The device immediately sends the text and touch events entered by the user to the server via data communication over the network. For example, when a user enters "hello," the text data is sent from the device to the server.
[1027] 3. Data Analysis
[1028] The server analyzes the received user text and touch events. It uses a large-scale language model (LLM) to infer the user's intent from the input. It also interprets the user's gestures based on the touch events. For example, the text "Hello" is analyzed and an appropriate response is generated: "Hello! What's wrong?"
[1029] 4. Behavior Generation
[1030] The server generates appropriate behavior for the character based on the analysis results. The character's personality information is also taken into consideration. For example, a character with a cheerful personality will be set to respond with a smile, and will respond to user input with a smile saying, "Hello! What's wrong?"
[1031] 5. Generating Animations
[1032] The server uses image generation technology to draw animation frames in real time based on the determined behavior. For example, 100 animation frames per second of an avatar smiling and saying "Hello! What's wrong?"
[1033] 6. Displaying Data
[1034] The generated animation frames are sent from the server to the device and played back in real time. This allows users to experience natural interactions with the avatar through seamless animation. For example, when the user says "Hello," the avatar will respond with a smile, "Hello! What's wrong?"
[1035] Specific examples
[1036] When a user inputs the text "Good morning" to an avatar, the device sends this input data "Good morning" to the server. The server uses LLM to analyze this text, understands the intent of the greeting, and generates a response such as "Good morning! Let's do our best today!". The server then generates an animation of the avatar smiling back based on this response. This animation frame is sent to the device and displayed in real time. The user can feel the avatar's response naturally, and simple interactions can lead to more complex and enjoyable interactions.
[1037] Prompt Sentence Examples
[1038] Below are some example prompts to input to the generative AI model:
[1039] Specific prompt:
[1040] When a user types "Hello," a cheerful avatar will respond with a smile, "Hello! What's up?" Specifically, you'll analyze the text, generate an appropriate response, and then animate it.
[1041] The above is an embodiment of the "Touch and Link" system of the present invention, which allows users to enjoy natural, real-time interaction with avatars.
[1042] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1043] Step 1:
[1044] User interface input
[1045] The device has an input field and a touch screen, which the user uses to input text into the system or to send touch events. For example, the user may type "hello" or tap on the face of an avatar. The input in this case is text data or touch event data.
[1046] Step 2:
[1047] Sending data
[1048] Any text input or touch events made by the user are immediately sent from the device to the server. Here, the input data (text or touch events) is sent via network communication and transmitted to the server in its original format. Specifically, the user's text "Hello" is sent from the device to the server.
[1049] Step 3:
[1050] Data analysis
[1051] The server analyzes the received user input data. It passes the input data (text and touch events) to a large-scale language model (LLM) that performs calculations to infer the user's intent. It also interprets gestures based on touch events. For example, the LLM analyzes the text "Hello" and generates an appropriate response, "Hello! What's wrong?" The analysis results are obtained as output.
[1052] Step 4:
[1053] Behavior Generation
[1054] The server dynamically generates character behavior based on the analysis results. The input here is the analysis results, and the output is specific character behavior information. Taking into account the character's personality information, for example, a character with a cheerful personality is set to respond with a smile.
[1055] Step 5:
[1056] Generate animation
[1057] The server uses image generation technology to draw animation frames based on the generated behavior information. The input is the character's behavior information, and animation frames are generated in real time based on that information. For example, 100 animation frames per second of a character smiling and saying "Hello! What's wrong?" are generated. The output is the generated animation frames.
[1058] Step 6:
[1059] Viewing Data
[1060] The generated animation frames are sent from the server to the device and played back in real time on the device. The input is the generated animation frames, which are sent to the device over the network and displayed seamlessly to the user. For example, in response to the user saying "Hello," an animation of an avatar smiling and saying "Hello! What's wrong?" is played back on the device. The output is the animation displayed on the device.
[1061] (Application example 1)
[1062] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1063] Conventional virtual agent systems have limited interaction with users, making it difficult to have natural conversations. Furthermore, they lack functionality as a shopping assistant, meaning they cannot provide sufficient support when users search for and purchase products. This results in a poor user experience and makes it difficult to promote sales effectively.
[1064] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1065] In this invention, the server includes means for receiving text input from a user and analyzing the user's intention, means for receiving touch events from the user and interpreting gestures, means for dynamically generating character behavior based on the user's intention using a large-scale language model, means for drawing animation frames in real time based on the generated behavior using image generation technology, means for transmitting the drawn animation frames to the user's terminal and displaying them, and means for receiving input from the user for searching for and purchasing products and providing product information based on the input. This enables natural interaction with the user, provides high performance as a shopping assistant, improves the user experience, and enables effective sales promotion.
[1066] "Text input from a user" refers to the act of a user inputting text information using a keyboard or touch screen.
[1067] A "means for analyzing user intent" is a means for analyzing what the text means and what it is trying to convey based on the text entered by the user.
[1068] A "touch event" is an event that occurs when a user taps, swipes, pinches, or performs other actions on a touchscreen.
[1069] The "means for interpreting a gesture" is a means for interpreting what gesture the user has made based on a touch event.
[1070] A "large-scale language model" is a machine learning model trained on massive amounts of text data, capable of understanding and generating natural language.
[1071] "Means for dynamically generating character behavior" refers to means for generating character movements and facial expressions in real time in response to user input and intentions.
[1072] "Image generation technology" is a technology that uses computer graphics to generate visual images and animations.
[1073] "Means for drawing animation frames in real time" refers to means for instantly creating and displaying animation frames based on user input and gestures.
[1074] An "animation frame" is an individual still image that makes up an animation.
[1075] "User terminal" refers to an electronic device used by a user, such as a smartphone, tablet, or PC.
[1076] The "displaying means" is a means for displaying the generated animation frames on the screen of the user's terminal.
[1077] "Input for searching and purchasing products" refers to the input operations performed by a user through a search bar or voice input to search for or purchase products.
[1078] "Means for providing product information" refers to the means for presenting information about products that users wish to search for and purchase.
[1079] "Character personality information" is information about the character's personality that is taken into consideration when determining the character's behavior and responses.
[1080] "Means for generating 100 or more animation frames per second" means means for generating 100 or more animation frames per second, thereby providing seamless animation.
[1081] The present invention is a system that realizes natural interactions with users and acts as a virtual shopping assistant. This system generates character behavior in real time based on user input, and supports product search and purchase.
[1082] System Configuration
[1083] User Interface
[1084] The device includes an input field for users to enter text or voice input, and a touchscreen for detecting touch events such as taps and swipes, through which users interact with the system. For example, a user can type "I want a smartwatch" or tap on a specific product.
[1085] Sending and Receiving Data
[1086] The device immediately sends text, touch events, and voice input from the user to the server via data communication over the network. For example, if a user types "I want a smartwatch," that data is sent from the device to the server.
[1087] Data analysis
[1088] The server analyzes the received user text, touch events, and voice data. It uses a large-scale language model (e.g., OpenAI's GPT-4) to infer the user's intent from their input. It also interprets the user's gestures based on the touch events. For example, a large-scale language model can analyze the text "I want a smartwatch" and generate the optimal response and product information.
[1089] Behavior Generation
[1090] Based on the analysis results, the server generates appropriate behavior for the character, taking into account the character's personality. For example, a character with a cheerful personality might respond with a smile, "What do you think of this smartwatch?"
[1091] Generate animation
[1092] The server uses image generation technology (e.g., Unity, Unreal Engine) to draw animation frames based on the generated behavior in real time. For example, it generates 60 animation frames per second of a virtual assistant pointing at a smartwatch.
[1093] Viewing Data
[1094] The generated animation frames are sent from the server to the device and played back in real time on the device. A seamless animation is displayed to the user, realizing natural interaction between the user and the virtual assistant. For example, if a user says, "I want a smartwatch," an animation is displayed in which the virtual assistant smiles and asks, "How about this smartwatch?"
[1095] Specific examples
[1096] Consider the case where a user texts the assistant, saying, "I want a smartwatch." At this time, the device sends the user's input to the server. The server uses GPT-4 to analyze the input, "I want a smartwatch," and infers the user's intent. Based on this, the server generates a behavior in which the virtual assistant responds with a smile, "How about this smartwatch?" and uses image generation technology to render this as an animation frame. The rendered animation frame is sent to the device and played back in real time on the user's device. In this way, the user experiences natural interactions with the virtual assistant, enabling smooth product search and purchase.
[1097] A specific example of a prompt sentence is as follows:
[1098] A user says, "I want a smartwatch." As a virtual shopping assistant, generate an appropriate response and consider the appropriate avatar behavior for that response.
[1099] This system allows users to interact with the virtual assistant naturally and in real time, making product searches and purchases more convenient.
[1100] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1101] Step 1:
[1102] The user provides text or voice input to the device. The device receives the user's input and sends it to the server as JSON format data. At this time, the input includes text or voice data such as "I want a smartwatch." The processing involves sending the input data to the server via the network. The output is the server receiving the input data.
[1103] Step 2:
[1104] The server analyzes the received text and voice data from the user. During the data analysis process, a large-scale language model (e.g., OpenAI's GPT-4) is used to infer the user's intent from their input. The data processing performed in this step involves text analysis and intent inference. Text data is given as input, and the analysis results are obtained as output. Specifically, the input "I want a smartwatch" is interpreted as "The user is looking for a smartwatch."
[1105] Step 3:
[1106] The server generates appropriate behavior for the virtual character based on the analysis results. In this process, the behavior is determined taking into account the character's personality information. For example, a character with a cheerful personality will respond with a smile. The input in this step is the analysis results, and the output is the generated character behavior. Specifically, the server generates the response "We recommend a smartwatch" and the corresponding character behavior.
[1107] Step 4:
[1108] The server uses image generation technology (e.g. Unity, Unreal Engine) to draw animation frames based on the generated behavior in real time. Character behavior information is given as input, and animation frames are the output. The specific operation in this step is for the animation software to generate frames.
[1109] Step 5:
[1110] The server sends the generated animation frames to the terminal. The input is the animation frames, and the output is the transmission of frames to the terminal. Specifically, the server sends animation data in real time over the network and prepares it for display on the user's terminal.
[1111] Step 6:
[1112] The device displays the received animation frames, providing the user with a seamless interaction experience. The input is the received animation frames, and the output is the display on the user's screen. Specifically, the device plays the animation frames, making it appear as if the virtual character is talking to the user.
[1113] This allows users to receive product information from the virtual shopping assistant through natural interaction.
[1114] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1115] This invention is a system incorporating an emotion analysis engine to realize natural interactions between users and virtual agents. This system analyzes user text input and touch events in real time, dynamically generates character behavior based on the analysis, and displays it as animation. Furthermore, by using the emotion engine to estimate the user's emotions and reflect them in the character's behavior, it realizes more natural, human-like responses.
[1116] System Configuration
[1117] User Interface
[1118] The device has a text entry field, a touchscreen, and voice input. Users can interact with the system through these, for example, by typing "hello" or by tapping on an avatar's face. They can also communicate using voice input.
[1119] Sending and Receiving Data
[1120] The device immediately transmits the user's text input, touch events, and voice input to the server using data communication over the network. For example, when a user types "hello," the data is sent from the device to the server.
[1121] Data analysis
[1122] The server analyzes the received user text, touch events, and voice data. It uses a large-scale language model (LLM) to infer the user's intent from the input. It also uses an emotion engine to infer the user's emotion. For example, the LLM analyzes the text "Hello" and generates an appropriate response such as "Hello! What's wrong?", and the emotion engine infers "joy."
[1123] Behavior Generation
[1124] Based on the analysis results, the server determines the appropriate behavior of the character, taking into account the character's personality and the user's emotional information. For example, if the user is happy, the character may respond with a bright smile.
[1125] Generate animation
[1126] The server uses image generation technology to draw animation frames in real time based on the generated behavior, for example, generating 100 animation frames per second of an avatar smiling and responding, "Hello! What's up?"
[1127] Viewing Data
[1128] The generated animation frames are sent from the server to the device and played back in real time on the device. A seamless animation is displayed to the user, realizing natural interaction between the user and the avatar. For example, when the user says "Hello," the avatar responds with a bright smile and says, "Hello! What's wrong?"
[1129] Specific examples
[1130] Consider the case where a user inputs text, "Hello," to an avatar, and simultaneously inputs voice data while smiling. At this time, the device sends the user's input to the server. The server uses LLM to analyze the input "Hello" and infer the user's intention. The emotion engine also infers the user's "happiness" through voice analysis. Based on this, the server generates a behavior for the avatar to respond with a smile, "Hello! What's wrong?" and uses image generation technology to render this as an animation frame. The rendered animation frame is sent to the device and played back in real time on the user's device. In this way, the user can experience more sophisticated and emotional interactions with the avatar.
[1131] The above is an embodiment of the "Touch and Link" system of the present invention, which combines an emotion engine. This system enables users to interact with avatars in a more human-like, natural way that reflects their emotions.
[1132] The processing flow will be explained below.
[1133] Step 1:
[1134] The user types "Hello" into the text input field of the device and presses the send button. The user also uses voice input to say "Hello."
[1135] Step 2:
[1136] The terminal receives the user's text input "Hello" and voice data and sends it to the server.
[1137] Step 3:
[1138] The server receives the text data and voice data sent from the terminal.
[1139] Step 4:
[1140] The server provides text data to a large-scale language model (LLM) and begins analysis.
[1141] Step 5:
[1142] The server receives the analysis results of the LLM and infers the intention from the user's text input. For example, in response to the input "Hello," it generates the response text "Hello! What's wrong?"
[1143] Step 6:
[1144] The server provides the voice data to the emotion engine, which analyzes the user's emotion. The emotion engine infers from the voice data that the user is feeling "joy."
[1145] Step 7:
[1146] When the server receives a touch event, it analyzes it and interprets which part was touched and how. For example, if a user taps on the avatar's head, it obtains that information.
[1147] Step 8:
[1148] The server determines the appropriate behavior of the avatar based on the intent analysis and emotion analysis results. Because the emotion engine estimates the user's "happiness," it determines the behavior of the avatar to respond with a smile, "Hello! What's wrong?"
[1149] Step 9:
[1150] The server uses image generation technology to generate animation frames based on the determined behavior, for example, generating 100 animation frames per second of an avatar smiling.
[1151] Step 10:
[1152] The animation frames generated by the server are divided into packets and sent to the terminal.
[1153] Step 11:
[1154] The device sequentially reads the animation frames received from the server and displays them on the user's screen in real time. For example, an animation of an avatar smoothly smiling and saying, "Hello! What's up?" is displayed.
[1155] The above is a specific flow of program processing that combines emotion engines in the "Touch and Link" system of the present invention.
[1156] Example 2
[1157] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1158] Conventional interactive character systems have had difficulty generating natural responses that accurately reflect the user's emotions and intentions. In particular, they must be able to handle not only text input but also touch events and voice input, and there is a demand for technology that can integrate and analyze these inputs and generate character behavior in real time. The objective of this invention is to provide a system that integrates various user input methods and realizes natural interactions that incorporate emotion analysis.
[1159] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes: means for receiving text input from a user and analyzing the user's intention; means for receiving touch events from the user and interpreting gestures; means for receiving voice input and analyzing the voice data; means for dynamically generating character behavior based on the user's intention using a large-scale language model; means for estimating the user's emotions using emotion analysis technology; means for determining the character's behavior based on the character's personality information and the user's emotion information; means for drawing animation frames in real time based on the behavior generated using image generation technology; and means for transmitting the drawn animation frames to the user's terminal and displaying them. This enables natural interaction that supports various user input means and reflects emotions in real time.
[1160] "Text input" refers to a method in which a user inputs characters using a keyboard, touch panel, or the like.
[1161] A "touch event" refers to a gesture, tap, swipe, or other action performed by a user on a touchscreen.
[1162] "Audio input" is the means by which a user sends voice to the system through a microphone.
[1163] A "large-scale language model" is a natural language processing system trained on vast amounts of text data to analyze user text input.
[1164] "Emotion analysis technology" is a technology that estimates a user's emotions from data such as voice and text.
[1165] "Character personality information" is information about the personality and characteristics set for a virtual character.
[1166] "Character behavior" refers to the actions and expressions made by a virtual character.
[1167] "Image generation technology" is a technology that uses computer graphics technology to generate still images and animations.
[1168] An "animation frame" is a sequence of still images that are displayed in succession to form a moving image.
[1169] A "server" is a computer system that receives data from users and analyzes and processes it.
[1170] A "terminal" is a computer system or device that is directly operated by a user and that performs input and display.
[1171] This invention is a system for realizing natural interactions between users and virtual characters. This system analyzes various user input methods (text input, touch events, and voice input), infers the user's intentions and emotions, and generates character behaviors in real time.
[1172] Hardware and Software
[1173] User Interface
[1174] The device includes a text input field, a touchscreen, and a microphone. This allows the user to interact with the system through text input, touch events, and voice input. For example, the user might type "hello" into the device, tap on the avatar's face, or speak "hello" into the microphone.
[1175] Sending and Receiving Data
[1176] The device transmits text input, touch events, and voice input data from the user to the server in real time using data communication over the Internet or a local network. For example, when the user types "hello," the data is immediately transmitted from the device to the server.
[1177] Data analysis
[1178] The server receives and analyzes various data (text, touch, and voice) sent by the user. It uses a large-scale language model (LLM) to analyze the user's text input and infer their intent. It also uses sentiment analysis technology to infer the user's emotions from voice data and touch events. For example, in response to the text "Hello," the LLM generates a response such as "Hello! What's wrong?", and the sentiment analysis technology infers the user's "joy."
[1179] Behavior Generation
[1180] The server determines the character's behavior based on the analysis results. This determination takes into account the character's personality information and the estimated user's emotional information. For example, if the user is happy, the character is set to respond with a bright smile.
[1181] Generate animation
[1182] The server uses image generation technology to generate animation frames based on the character's behavior in real time. Specifically, technologies such as DALL-E and GAN (Generative Adversarial Networks) are used. For example, it can generate 100 animation frames per second of an avatar smiling and saying, "Hello! What's up?"
[1183] Viewing Data
[1184] The generated animation frames are sent from the server to the device and displayed in real time on the device, allowing users to experience seamless and natural animation. For example, after typing "hello," the user can see an animation of an avatar smiling and replying, "Hello! What's up?"
[1185] Specific examples
[1186] Consider the case where a user uses a device to input text, such as "Hello," and simultaneously input voice data. In this case, the device sends the user's text input and voice data to the server. The server uses a large-scale language model to analyze the input "Hello" and generate an appropriate response, such as "Hello! What's wrong?" At the same time, it uses emotion analysis technology to estimate the user's "happiness" from the voice data. The server then uses this information to generate a behavior for the avatar to respond with a smile, saying "Hello! What's wrong?" and uses image generation technology to render this as an animation frame. The rendered animation frame is sent to the device and displayed in real time on the user's device.
[1187] For example, a prompt might look like this:
[1188] "The user types "hello" into the avatar and speaks to it with a smile using voice input. The device then sends the user's input to the server, which analyzes it and generates an appropriate response and character behavior."
[1189] The above is an embodiment of the invention. This system allows users to have more natural and emotionally appropriate interactions with virtual characters.
[1190] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1191] Step 1:
[1192] The user types "hello" into the device's text input field, taps the avatar's face, or speaks "hello" into the microphone. The device then captures the input data (text, touch, and voice) from the user. Input data such as "hello," as well as the coordinates of the touch event and voice data, are generated.
[1193] Step 2:
[1194] The device transmits the captured user input data to the server via the network. Specifically, the device converts the text, touch events, and voice data entered by the user into packets and transfers them to the server via the network. The output is data in packet format.
[1195] Step 3:
[1196] The server receives the user's input data sent from the device and uses a large-scale language model (LLM) to analyze the user's intent from the text data. For example, the server analyzes the text data "Hello" and generates a response such as "Hello! What's wrong?" The output is the response text as the analysis result.
[1197] Step 4:
[1198] The server analyzes the voice data and touch events using emotion analysis technology to estimate the user's emotion. For example, it analyzes the tone of the voice and the strength of the touch and estimates that the user has the emotion of "joy." The output is the estimated emotion information.
[1199] Step 5:
[1200] The server determines the avatar's behavior based on the intention and emotion analysis results of the LLM, taking into account the character's personality information. For example, if the user is happy, the avatar is set to respond with a bright smile. The output is the generated character's behavior data.
[1201] Step 6:
[1202] The server generates animation frames in real time based on the behavior determined using image generation technology. For example, using technologies such as DALL-E or GAN, it generates 100 animation frames per second in which an avatar smiles and responds, "Hello! What's wrong?" The output is the generated animation frames.
[1203] Step 7:
[1204] The server transmits the generated animation frames to the terminal via the network. At this time, the animation frames are converted into packets and transferred to the terminal via the network. The output is animation frame data in packet format.
[1205] Step 8:
[1206] The device receives the animation frames sent from the server and displays them in real time, allowing the user to see a seamless animation of the avatar. Specifically, an animation of the avatar smiling and saying "Hello! What's up?" is played on the device's display. The output is the displayed animation.
[1207] The above are the specific processing steps of the program for this system. Users can use various input means to naturally interact with the avatar.
[1208] (Application example 2)
[1209] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1210] In modern digital interfaces, interactions between users and virtual agents often lack naturalness. Furthermore, food delivery services lack effective information provision that takes user emotions into account. Therefore, there is a demand for interfaces that are more human-like and can respond to users' emotions.
[1211] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving text input from a user and analyzing the user's intention, means for receiving touch events from the user and interpreting gestures, means for estimating the user's emotion using an emotion engine, means for dynamically generating character behavior based on the user's intention using a large-scale language model (LLM), means for drawing animation frames in real time based on the generated behavior using image generation technology, means for transmitting the drawn animation frames to the user's terminal and displaying them, and means for providing information depending on the user's emotional state. This enables the user to have more natural and emotionally appropriate interactions with the virtual agent.
[1212] "Text input" is the act of a user providing a string of characters to a system using an input device.
[1213] "Analyzing intent" is the process of inferring a user's purpose or request from their text input or touch events.
[1214] A "touch event" is input information generated when a user touches a touchscreen device.
[1215] "Interpreting gestures" means recognizing the user's movements and instructions based on their touch events.
[1216] A "large-scale language model (LLM)" is a natural language processing model trained on massive amounts of text data, and can be applied to a variety of language tasks.
[1217] "Character behavior" refers to the actions and reactions of the virtual agent, which are dynamically generated by the system.
[1218] "Emotion engine" is a general term for algorithms and software that infer emotions from user input information.
[1219] An "animation frame" is an individual image that displays a sequence of character behavior.
[1220] "Drawing in real time" means generating and displaying animation frames instantly in response to user input.
[1221] A "user terminal" is an electronic device used by a user, such as a computer, smartphone, or tablet.
[1222] "Providing information depending on the emotional state" means providing information or suggestions appropriate to the user at any given time, taking into account the user's emotions.
[1223] A system embodying this invention provides natural, emotion-sensitive interactions between a user and a virtual agent by analyzing the user's text input, touch events, and emotions to generate natural responses in real time and display them on the user's device.
[1224] System Configuration
[1225] User Interface
[1226] The user's device has a text entry field, a touch screen, and voice input capabilities. The user interacts with the system by entering text, touching the screen, or communicating by voice. For example, the user enters the text "I want to feel better."
[1227] Sending and Receiving Data
[1228] The device immediately transmits the user's text input, touch events, and voice input to the server using data communication over the network. For example, when a user types "I need energy," the data is sent from the device to the server.
[1229] Data analysis
[1230] The server analyzes the received user text, touch events, and voice data. It uses a large-scale language model (LLM) to infer the user's intent and an emotion engine to infer the user's emotion. For example, the LLM analyzes the text "I need energy" and generates an appropriate response such as "I'm looking for food to give me energy," while the emotion engine infers "I'm tired."
[1231] Behavior Generation
[1232] The server then determines the appropriate behavior for the character based on the analysis results, taking into account the character's personality and the user's emotional information. For example, if the user is feeling tired, the server may suggest food delivery in an encouraging tone.
[1233] Generate animation
[1234] The server uses image generation technology to draw animation frames in real time based on the generated behavior, generating 100 animation frames per second of a virtual agent smiling and responding, for example, "I'm looking for food to cheer me up."
[1235] Viewing Data
[1236] The generated animation frames are sent from the server to the device and played back in real time on the user's device. Seamless, emotional animations are displayed to the user, enabling more natural interactions. For example, if a user types "I want to feel energized," the virtual agent will respond in real time with a smile, saying, "I'm looking for food to cheer you up."
[1237] Examples of specific examples and prompts
[1238] Examples:
[1239] The user inputs the text "I'm tired and want to feel energized." The device then sends the user's input to the server. The server uses a large-scale language model to analyze the input, "I'm tired and want to feel energized," and infers the user's intention. The emotion engine also infers "fatigue" from the input. Based on this, the server generates a behavior for the virtual agent to respond with "I'm looking for food to cheer me up," and uses image generation technology to render this as an animation frame. The rendered frame is then sent to the device and played back in real time on the user's device. The user can experience a sophisticated and emotional interaction with the virtual agent.
[1240] Example prompt sentence:
[1241] Analyze user input, infer emotions using the emotion engine, and generate context-sensitive responses.
[1242] User input: "I'm tired and need some encouragement."
[1243] In this way, the system according to the present invention is able to provide natural and emotional interactions taking into account the user's emotions.
[1244] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1245] Step 1:
[1246] The user inputs text, touch events, or voice input. The user's device receives this input through an input device (keyboard, touch screen, microphone). Specifically, the user inputs text such as "I'm tired and need some energy."
[1247] Step 2:
[1248] The device sends the user's input data to the server. The input data (text, touch information, voice data) is transferred immediately via network communication. The data sent at this time is in JSON format, etc.
[1249] Step 3:
[1250] The server analyzes the received text input using a large-scale language model (LLM). Specifically, the received text data ("I'm tired and want to be cheered up") is input into the LLM, and the server infers the user's intent (wanting to receive cheering suggestions).
[1251] Step 4:
[1252] The server uses an emotion engine to estimate the user's emotion from the input data. The emotion engine detects the emotion "fatigue" from the voice and text. The emotion engine analyzes the emotion using natural language processing models and speech analysis algorithms.
[1253] Step 5:
[1254] The server generates appropriate behavior for the virtual agent based on the user's intentions and emotions. It also takes into account the character's personality information to determine the optimal response. For example, it selects a response such as "Suggest foods that will cheer you up."
[1255] Step 6:
[1256] The server uses image generation technology to draw animation frames based on the virtual agent's behavior, generating 100 animation frames per second of a smiling virtual agent saying, "I'm looking for something to cheer me up."
[1257] Step 7:
[1258] The server then sends the generated animation frames to the user's device. This requires real-time data transfer, and the data is compressed and sent over a high-speed network.
[1259] Step 8:
[1260] The device then plays the received animation frames in real time, providing a seamless animation on the user's device, enabling natural, emotionally-responsive interaction with the virtual agent. The user is shown an animated response with a smile, saying, "I'm looking for some food to cheer me up."
[1261] The above is the flow of the system's program processing and the specific actions performed at each step. This system allows users to experience more natural and emotionally appropriate interactions.
[1262] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1263] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1264] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1265] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1266] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1267] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1268] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1269] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1270] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1271] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1272] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1273] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1274] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1275] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1276] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1277] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1278] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1279] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1280] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1281] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1282] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1283] The following is further disclosed regarding the above embodiment.
[1284] (Claim 1)
[1285] means for receiving text input from a user and analyzing the user's intent;
[1286] means for receiving touch events from a user and interpreting gestures;
[1287] A means for dynamically generating character behavior based on user intentions using a large-scale language model (LLM);
[1288] means for rendering animation frames in real time based on the generated behavior using image generation techniques;
[1289] means for transmitting the rendered animation frames to a user's terminal for display;
[1290] A system including:
[1291] (Claim 2)
[1292] 10. The system of claim 1, further comprising means for generating character behavior taking into account character personality information.
[1293] (Claim 3)
[1294] 10. The system of claim 1, further comprising means for generating 100 or more animation frames per second based on user text input and touch events to provide seamless animation.
[1295] "Example 1"
[1296] (Claim 1)
[1297] means for receiving text input from a user and analyzing the user's intent;
[1298] means for receiving touch events from a user and interpreting gestures;
[1299] A means for dynamically generating character behavior based on a user's intentions using a large-scale language model;
[1300] means for rendering animation frames in real time based on the behavior generated using the image generation means;
[1301] means for transmitting the rendered animation frames to a user's display device for display;
[1302] A system including:
[1303] (Claim 2)
[1304] 10. The system of claim 1, further comprising means for generating character behavior taking into account character personality information.
[1305] (Claim 3)
[1306] 10. The system of claim 1, further comprising means for generating 100 or more animation frames per second based on user text input and touch events to provide seamless animation.
[1307] "Application Example 1"
[1308] (Claim 1)
[1309] means for receiving text input from a user and analyzing the user's intent;
[1310] means for receiving touch events from a user and interpreting gestures;
[1311] A means for dynamically generating character behavior based on a user's intentions using a large-scale language model;
[1312] means for rendering animation frames in real time based on the generated behavior using image generation techniques;
[1313] means for transmitting the rendered animation frames to a user's terminal for display;
[1314] A means for receiving input from a user to search for and purchase products and providing product information based on the input;
[1315] A system including:
[1316] (Claim 2)
[1317] 10. The system of claim 1, further comprising means for generating character behavior taking into account character personality information.
[1318] (Claim 3)
[1319] 10. The system of claim 1, further comprising means for generating 100 or more animation frames per second based on user text input and touch events to provide seamless animation.
[1320] "Example 2: Combining Emotion Engines"
[1321] (Claim 1)
[1322] means for receiving text input from a user and analyzing the user's intent;
[1323] means for receiving touch events from a user and interpreting gestures;
[1324] means for receiving voice input and analyzing the voice data;
[1325] A means for dynamically generating character behavior based on a user's intentions using a large-scale language model;
[1326] means for estimating a user's emotion using emotion analysis techniques;
[1327] means for determining the behavior of the character based on the character's personality information and the user's emotional information;
[1328] means for rendering animation frames in real time based on the generated behavior using image generation techniques;
[1329] means for transmitting the rendered animation frames to a user's terminal for display;
[1330] A system including:
[1331] (Claim 2)
[1332] 10. The system of claim 1, further comprising means for generating character behavior taking into account character personality information.
[1333] (Claim 3)
[1334] 10. The system of claim 1, further comprising means for generating 100 or more animation frames per second based on user text input and touch events to provide seamless animation.
[1335] "Application example 2 when combining emotion engines"
[1336] (Claim 1)
[1337] means for receiving text input from a user and analyzing the user's intent;
[1338] means for receiving touch events from a user and interpreting gestures;
[1339] A means for dynamically generating character behavior based on user intentions using a large-scale language model (LLM);
[1340] means for estimating a user's emotion using an emotion engine;
[1341] means for generating an appropriate response based on the estimated emotion;
[1342] means for rendering animation frames in real time based on the generated behavior using image generation techniques;
[1343] means for transmitting the rendered animation frames to a user's terminal for display;
[1344] means for providing information dependent on the emotional state of the user;
[1345] A system including:
[1346] (Claim 2)
[1347] 2. The system according to claim 1, further comprising means for generating character behavior in consideration of character personality information and user emotion information.
[1348] (Claim 3)
[1349] 10. The system of claim 1, further comprising means for generating 100 or more animation frames per second based on user text input and touch events to provide seamless animation. [Explanation of symbols]
[1350] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for receiving text input from a user and analyzing the user's intent; means for receiving touch events from a user and interpreting gestures; A means for dynamically generating character behavior based on a user's intentions using a large-scale language model; means for rendering animation frames in real time based on the generated behavior using image generation techniques; means for transmitting the rendered animation frames to a user's terminal for display; A system including:
2. 2. The system according to claim 1, further comprising means for generating character behavior in consideration of character personality information.
3. 10. The system of claim 1, further comprising means for generating 100 or more animation frames per second based on user text input and touch events to provide seamless animation.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A