System
The system addresses the limitations of existing interactive storytelling systems by receiving and analyzing user prompts, generating character actions, and adjusting story complexity and emotion, enhancing user engagement and creativity.
Patent Information
- Application Number
- JP2024118253
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-23
- Publication Date
- 2026-02-04
AI Technical Summary
Existing systems lack the ability to flexibly respond to user input in real-time and reflect user creativity, limiting the entertainment value and user experience in interactive storytelling and virtual environments.
A system that receives user prompts, analyzes them using natural language processing, generates character actions, and displays the results, allowing for adjustable story complexity based on user skills and incorporating emotion recognition to provide a dynamic and personalized experience.
Enables users to hone their technical skills and creativity through interactive storytelling, providing a dynamic and personalized experience by generating character actions and adjusting story content in real-time based on user input and emotions.
Smart Images

Figure 2026017471000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] With the recent development of AI technology, the demand for new user experiences and creative storytelling is increasing. However, understanding of the techniques for providing prompts to AI is not yet widespread, and opportunities to hone both technical skills and creativity are limited. As a result, users are not able to fully enjoy the opportunities to learn new AI technologies and storytelling techniques. [Means for solving the problem]
[0005] The present invention solves the above problems by using the following means.
[0006] First, it provides a means for receiving prompts from the user. Next, it uses a means for analyzing the received prompts and generating character actions. This analysis employs natural language processing technology. It also includes a means for displaying the generated character actions and has a means for changing the complexity of the story according to the user's prompting skills. It also stores the prompt history and uses it for analysis. It also has a network connection means for sending and receiving data between the server and the user's device. This provides a new gaming experience that allows users to hone their technical skills and creativity at the same time through interaction between the AI and the user.
[0007] A "user" is an entity that operates the system and inputs prompts to trigger character actions.
[0008] A "prompt" is a textual instruction that a user inputs to the system, and is a source of information that determines the character's actions.
[0009] "Means for receiving" refers to the process by which the system obtains the prompt input from the user.
[0010] "Means of analysis" refers to the process of understanding the prompt received and determining the character's specific actions based on its content.
[0011] "Character behavior" refers to the actions and reactions of the virtual character generated by the system based on the analyzed prompts.
[0012] "Means for displaying" refers to the process of visualizing the generated character actions and story text to the user.
[0013] "Natural language processing" refers to the technology that enables computers to understand, analyze, and generate human language.
[0014] "Story text" refers to the narrative text related to the character's actions, which is displayed to the user.
[0015] "Means of varying complexity" refers to the process of adjusting the content and difficulty of the story according to the user's prompting skills.
[0016] "Network connection means" refers to a communication means for sending and receiving data between a server and a user terminal.
[0017] "Prompt history" refers to a record of prompts that the user has entered so far, and indicates the data that will be used for analysis. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] The present invention is a system for receiving prompts from a user, analyzing the prompts, generating character actions, and displaying the actions to progress a story. Specific embodiments for implementing this system will be described below.
[0040] System Overview
[0041] This system operates mainly with a server, a terminal, and a user. The server mainly analyzes prompts and generates character actions, while the terminal is responsible for inputting prompts from the user and displaying the received character actions and story text.
[0042] Program processing
[0043] Entering and submitting prompts
[0044] First, the user inputs a prompt into the input field of the terminal and presses the send button, at which point the terminal sends the input prompt to the server.
[0045] Parsing prompts
[0046] The server passes the received prompt to an AI analysis engine, which uses natural language processing technology to understand the meaning of the prompt and determine the character's specific actions.
[0047] Character behavior and story text generation
[0048] Next, the AI analysis engine generates character actions based on the analysis results, and also generates story text related to the character actions. The server compiles this data and sends it to the device.
[0049] Screen display
[0050] The device displays the received character actions and story text on the screen, allowing the user to see in real time how their prompts are reflected.
[0051] Specific examples
[0052] Example 1: Adventure game
[0053] The user types in "Find a way to enter the forest" and presses the send button. The device sends this prompt to the server. The server uses an AI analysis engine to analyze the prompt "Find a way to enter the forest" and determines that the character should take the action of "taking out a map and starting to search for a way." The server then generates the story text, "The character unfolds the map and starts to search for a way. After walking a little further, he finds a hidden entrance into the forest." This data is sent to the device, which displays it on the screen.
[0054] Example 2: Mystery game
[0055] The user types "Find clues" and presses the send button. The device sends this prompt to the server. The server uses an AI analysis engine to analyze the prompt "Find clues" and determines the character's action to "begin to investigate the room in detail." Next, it generates story text that reads, "The character investigates the entire room and finds hidden clues." This data is sent to the device, which displays it on the screen.
[0056] As described above, the system receives and analyzes prompts, generates and displays character actions, and allows users to progress through the story, allowing users to hone their technical skills and creativity through interaction with AI.
[0057] The processing flow will be explained below.
[0058] Step 1:
[0059] The user enters a prompt (e.g., "Find a way into the forest") into the input field of the terminal and presses the send button.
[0060] Step 2:
[0061] The terminal sends the prompts entered by the user to the server, where the input data is converted into packets and sent over the network.
[0062] Step 3:
[0063] The server passes the received prompt to the AI analysis engine. Since the prompt is in text format and cannot be understood by the AI as is, it is analyzed via a natural language processing API.
[0064] Step 4:
[0065] The AI analysis engine analyzes the meaning of the received prompt. Specifically, it goes through processes such as grammatical analysis, keyword extraction, and semantic analysis to understand the user's instruction to "find a way into the forest."
[0066] Step 5:
[0067] The AI analysis engine generates specific character actions based on the prompt (e.g., "take out a map and find the way"), and this generated action data is returned to the server in a data structure such as JSON format.
[0068] Step 6:
[0069] The server receives the character behavior data returned from the AI analysis engine and generates corresponding story text (e.g., "The character unfolds the map and begins to search for a path. After walking a little further, he finds a hidden entrance into the forest.").
[0070] Step 7:
[0071] The server combines the generated character behavior data and story text into a single packet and sends it to the terminal.
[0072] Step 8:
[0073] The device analyzes the data it receives and prepares it for display on the screen, specifically generating character animations and formatting story text.
[0074] Step 9:
[0075] The device displays the character's action scenes and story text on the screen, allowing the user to visually confirm how their prompts have been reflected.
[0076] The above explains the processing content of the program of this system, broken down into specific processing steps. The present invention provides a new type of game in which the actions of an AI character are determined by user prompts and are displayed as a story.
[0077] Example 1
[0078] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0079] Conventional virtual character systems have struggled to respond flexibly and in real time to user input and progress the story. Furthermore, they lack the ability to reflect the user's creative input, resulting in a lack of entertainment value. This limits the user experience and makes it difficult to maintain sustained interest.
[0080] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0081] In this invention, the server includes means for receiving input data from a user, means for analyzing the received input data and generating virtual character behaviors, and means for displaying the generated virtual character behaviors, thereby enabling the server to analyze user prompts in real time and generate appropriate virtual character behaviors and story texts.
[0082] "User" refers to the person who operates the system and provides input data.
[0083] "Input data" refers to text or commands that a user enters into a system.
[0084] "Means for receiving" refers to the functions and processes for obtaining input data provided by a user.
[0085] "Means of analysis" refers to the function or process that understands received input data and processes it to extract meaning.
[0086] "Virtual character behavior" refers to a specific action taken by a virtual character that is generated based on the analysis results.
[0087] "Generating means" refers to the function or process that produces virtual character behavior and narrative text based on the analysis of input data.
[0088] "Display means" refers to a function or process for visually presenting the generated virtual character actions and narrative text to the user.
[0089] "Natural language processing technology" refers to technology that enables computers to understand, interpret, and generate human language.
[0090] "Story text" refers to the text of a story that is generated based on the actions of the generated virtual character.
[0091] A "system" refers to a collection of functions and processes realized by combining the above means.
[0092] The present invention is a system for receiving input data from a user, analyzing the data, and generating and displaying the behavior of a virtual character. Specific embodiments of the present invention will be described in detail below.
[0093] Program processing flow
[0094] This system mainly consists of three components: a server, a terminal, and a user. The server analyzes prompts and generates virtual character actions, while the terminal receives prompts from the user and displays the received virtual character actions and story text.
[0095] Hardware and software used
[0096] server
[0097] The server uses a high-performance computer, and uses software such as a generative AI model (e.g., GPT-3) to analyze prompts and generate virtual character behavior and narrative text.
[0098] Terminal
[0099] The terminals used can be devices capable of input and display, such as personal computers, smartphones, tablets, etc. The user inputs prompts through the terminal and views the results sent from the server.
[0100] Specific explanation of data processing and data calculation
[0101] 1. The user enters the prompt into the input field on the terminal and presses the send button.
[0102] 2. The terminal sends the entered prompt to the server using an HTTP request.
[0103] 3. The server passes the received prompt to a generative AI model (e.g., GPT-3) for analysis.
[0104] 4. The generative AI model uses natural language processing techniques to analyze the input prompt and understand its meaning.
[0105] 5. Based on the analysis results, the virtual character's actions are generated, along with narrative text related to the generated actions.
[0106] 6. The server sends the generated virtual character actions and story text to the device, again using an HTTP request.
[0107] 7. The terminal displays the received data on the screen, allowing the user to see how their prompts were reflected.
[0108] Specific examples
[0109] Example 1: Adventure game
[0110] The user types "Find a way into the forest" and presses the send button. The device sends this prompt to the server. The server uses a generative AI model (e.g., "GPT-3") to analyze the prompt "Find a way into the forest." Based on the analysis results, the server generates an action for the character to "take out a map and start looking for a way." Next, it generates narrative text such as "The character unfolds the map and starts looking for a way. After walking a little further, he finds a hidden entrance into the forest." This data is sent to the device, which displays it on the screen.
[0111] Example 2: Mystery game
[0112] The user types "Find clues" and presses the send button. The device sends this prompt to the server. The server uses a generative AI model (e.g., "GPT-3") to analyze the prompt "Find clues." Based on the analysis results, the server generates an action for the character to "begin to investigate the room in detail." Next, it generates narrative text such as "The character investigates the entire room and finds hidden clues." This data is sent to the device, which displays it on the screen.
[0113] As described above, the present invention is a system that receives and analyzes user prompts, generates and displays the actions of virtual characters, and allows users to progress through their own story, honing their technical skills and creativity through interaction with AI.
[0114] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0115] Step 1:
[0116] The user enters a prompt into the input field of the terminal and presses the send button. The input data is in the form of a specific command or question in text format. For example, "Find a way into the forest." This is the input data. The terminal receives the input data through the user's operation.
[0117] Step 2:
[0118] The terminal generates and sends an HTTP request to send the received input data to the server. The input data is a prompt sentence in text format, and is sent to the server as the payload of the HTTP request. This transfers the input data from the terminal to the server.
[0119] Step 3:
[0120] The server receives the HTTP request sent from the terminal and extracts the prompt text. The received data is the text data included in the payload of the HTTP request. This text data is passed to the next process for analysis.
[0121] Step 4:
[0122] The server passes the received prompt sentence to a generative AI model (for example, GPT-3) for analysis. The input is the text data of the prompt, and the output is the analysis result. The generative AI model uses natural language processing technology to analyze the meaning of the prompt and determine the specific actions of the virtual character. This analysis results in actions such as "taking out a map and starting to look for directions."
[0123] Step 5:
[0124] The server generates narrative text related to the virtual character's actions based on the analysis results obtained by the generative AI model. The input is the analysis results, and the output is the generated character's actions and narrative text. For example, the generated text might read, "The character unfolds the map and begins to search for a path. After walking a little further, he finds a hidden entrance into the forest."
[0125] Step 6:
[0126] The server sends the generated virtual character actions and story text to the terminal as an HTTP response. The input is the generated data, and the output is the HTTP response. This transfers data from the server to the terminal.
[0127] Step 7:
[0128] The device extracts the virtual character's actions and story text from the received HTTP response and displays them on the screen. The input is the received data, and the output is a display format that the user can visually confirm. Specifically, the screen displays the message, "The character unfolds the map and begins to search for a path. After walking a little further, he finds a hidden entrance into the forest." This allows the user to see in real time how the prompts they entered have been reflected.
[0129] As described above, specific data processing and data calculations are carried out at each step, and ultimately the system displays appropriate virtual character behavior and story text to the user.
[0130] (Application example 1)
[0131] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0132] A problem with virtual stores is the lack of a system that responds appropriately to customer requests and questions about products. Another issue is the lack of an interface that allows virtual store clerk characters to make suggestions tailored to the customer's needs. This can lead to a lack of improvement in the customer experience and a decrease in customer satisfaction.
[0133] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0134] In this invention, the server includes means for receiving prompts from a user, means for analyzing the received prompts and generating character actions and suggestions, means for displaying the generated character actions and suggestions, and means for making suggestions in response to requests from customers of the virtual store, thereby enabling the virtual store clerk character to take appropriate actions and make appropriate suggestions in response to requests and questions input by customers.
[0135] The "means for receiving prompts from a user" is an interface for acquiring requests or questions entered by customers of the virtual store via a digital device.
[0136] The "means for analyzing prompts and generating character behavior and proposal content" is a system that uses natural language processing technology to understand the input content from customers and generates appropriate behavior for a virtual store clerk character and product proposals based on that.
[0137] The "means for displaying the generated character's actions and proposals" refers to a display or user interface that visually conveys to the customer the character's actions and proposals generated based on the analysis results.
[0138] The "means for making proposals in response to customer requests in a virtual store" is a system that has the function of allowing a virtual store clerk character to make appropriate product proposals and provide information regarding the products and information that customers are looking for.
[0139] The present invention is a system for generating and displaying character actions and suggestions in response to customer requests and questions in a virtual store. A specific embodiment for implementing this system will be described.
[0140] System configuration
[0141] This system operates mainly with a server, a terminal, and a user. The server mainly analyzes prompts and generates character actions and suggestions, while the terminal is responsible for inputting prompts from the user and displaying character actions, suggestions, and story text received from the server.
[0142] Program processing
[0143] server
[0144] The server has an interface for receiving prompts from the user. It also analyzes the received prompts using natural language processing technology and generates the actions and suggestions of a virtual store clerk character based on the analysis results. A generative AI model is used to generate the character's actions and suggestions. The generated data is sent to the terminal.
[0145] Terminal
[0146] The terminal provides an input field for the user to enter a prompt. When the user enters and submits the prompt, the terminal sends it to the server. Upon receiving the analysis results from the server, the terminal displays the character's actions and suggestions, along with the associated story text.
[0147] User
[0148] The user inputs a request or question into the input field on the device. For example, they can type "Show me shoes that go with this dress" and press the send button. The user can then check the character's actions and suggestions sent from the server on the device screen.
[0149] Technical details
[0150] Hardware and software used
[0151] Hardware: Digital devices such as smartphones, tablets, and PCs
[0152] software:
[0153] Server side: Flask (a Python micro web framework)
[0154] Client-side: Web browser or dedicated application
[0155] Natural Language Processing: Generative AI models such as OpenAI GPT
[0156] Communication: requests module using HTTP protocol
[0157] Data processing and calculation
[0158] 1. The user types a prompt (such as "Show me the shoes that go with this dress") into the device's input field.
[0159] 2. The terminal sends the prompt to the server via an HTTP request.
[0160] 3. The server receives the prompt and passes it to a generative AI model for analysis, using natural language processing techniques.
[0161] 4. As a result of the analysis, character actions and suggested content are generated. For example, a character action such as "I selected three suggested shoes from the catalog and displayed them on the screen" and story text such as "The virtual salesperson selected three shoes from the catalog and explained the features of each pair and how they would match your dress" are generated.
[0162] 5. The server sends the generated data to the terminal as an HTTP response.
[0163] 6. The terminal displays the received data on the user interface, and the user confirms it.
[0164] As described above, the system of the present invention receives user prompts and generates appropriate character actions and suggestions based on the prompts, thereby improving the customer experience in a virtual store.
[0165] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0166] Step 1:
[0167] The user inputs a request or question into the input field of the terminal. For example, the user inputs a prompt sentence such as "Show me the shoes that go with this dress" and presses the send button. The input prompt sentence is provided to the terminal as input.
[0168] Step 2:
[0169] The terminal sends the prompt text entered by the user to the server. Specifically, it creates an HTTP request and sends the prompt text as a payload to the specified endpoint of the server. At this point, the input is the prompt text, and the output is the request data sent to the server.
[0170] Step 3:
[0171] The server analyzes the received prompt. Here, the server passes the prompt to a generative AI model (e.g., OpenAI GPT) for analysis. During this analysis process, the input is the prompt, and the output is the character's actions and suggestions as a result of the analysis.
[0172] Step 4:
[0173] The server generates character actions and suggestions based on the analysis results. For example, the analysis results might generate a character action such as "I selected three suggested shoes from the catalog and displayed them on the screen," and a story text such as "The virtual salesperson selected three shoes from the catalog and explained the features of each pair and how they would match your dress." The input for this step is the analysis results from the generative AI model, and the output is the generated character actions and story text.
[0174] Step 5:
[0175] The server sends the generated character actions and suggestions to the device. Specifically, it returns data created based on the analysis results to the device as an HTTP response. At this point, the input is the generated character actions and story text, and the output is the data sent to the device.
[0176] Step 6:
[0177] The device displays the character actions and suggestions received from the server on the user interface, allowing the user to visually confirm them. In this step, the device analyzes the received data and renders it as a view to be displayed on the screen. The input is the character actions and story text sent from the server, and the output is the visual data displayed on the user interface.
[0178] Through the above processing steps, a system is realized in which appropriate character actions and suggestions are provided to users in real time in response to customer requests in a virtual store.
[0179] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0180] The present invention is a system that receives prompts from a user, analyzes them, generates and displays character actions, and also recognizes the user's emotions and reflects them in the character actions and story text. Specific embodiments for implementing this system will be described below.
[0181] System Overview
[0182] This system mainly consists of a server, a terminal, a user, and an emotion engine. The server mainly analyzes prompts and generates character actions, while the emotion engine recognizes the user's emotions. The terminal is responsible for inputting prompts from the user, acquiring user emotion information, and displaying received character actions and story text.
[0183] Program processing
[0184] Entering and submitting prompts
[0185] First, the user enters a prompt (e.g., "Find a way into the forest") into an input field on the device and presses the send button. Then, the device sends the user's emotional information (e.g., facial recognition and voice tone analysis) to the emotion engine.
[0186] Parsing prompts
[0187] The server passes the received prompt to an AI analysis engine, which uses natural language processing technology to understand the meaning of the prompt and determine the character's specific actions.
[0188] Reflecting emotional information
[0189] The emotion engine analyzes the user's emotions and sends the results to the server. The server uses this emotional information to make fine adjustments to the character's behavior and story text. For example, if the user is excited, the character's behavior may be changed more dynamically.
[0190] Character behavior and story text generation
[0191] Next, the AI analysis engine generates character actions based on the prompt analysis results and emotional information. It also generates story text related to the character actions. The server then compiles this data and sends it to the device.
[0192] Screen display
[0193] The device displays the received character actions and story text on the screen, allowing the user to see in real time how their prompts and emotions are reflected.
[0194] Specific examples
[0195] Example 1: Adventure game
[0196] When the user types "Find a way into the forest" and presses the send button, the device simultaneously sends the user's facial recognition information to the emotion engine. The emotion engine recognizes the user's emotion as "curious" and sends this information to the server. The server uses its AI analysis engine to analyze the prompt, determines the action to "take out a map and find a way," and generates a scene in which the character happily checks the map, reflecting the emotion of "curious." Next, the server generates story text that reads, "The character unfolds the map and begins to search for a way with gusto. After walking a little further, he excitedly finds a hidden entrance to the forest." This data is sent to the device, which displays it on the screen.
[0197] Example 2: Mystery game
[0198] When the user types "find clues" and presses the send button, the device simultaneously sends the user's voice tone information to the emotion engine. The emotion engine recognizes the user's emotion as "anxiety" and sends this information to the server. The server uses its AI analysis engine to analyze the prompt, determines the action to "start investigating the room in detail," and generates a scene of the character moving quickly, reflecting the emotion of "anxiety." Next, story text is generated: "The character investigates the room anxiously and finds hidden clues." This data is sent to the device, which displays it on the screen.
[0199] The above explains the processing content of the program of this system, broken down into specific processing steps. The present invention provides a more advanced and personalized user experience by dynamically changing the behavior and story of the AI character according to the user's prompts and emotions.
[0200] The processing flow will be explained below.
[0201] Step 1:
[0202] The user enters a prompt (e.g., "Find a way into the forest") into the input field of the terminal and presses the send button.
[0203] Step 2:
[0204] The device sends the input prompt to the server, and at the same time, sends the user's facial recognition information and emotion data such as voice tone to the emotion engine.
[0205] Step 3:
[0206] The server receives the prompt and passes it to the AI analysis engine, which uses natural language processing to analyze the prompt and determine the character's actions based on its content.
[0207] Step 4:
[0208] The emotion engine analyzes the user's emotion data and sends the recognized emotional state (e.g., "curious" or "anxious") to the server.
[0209] Step 5:
[0210] The server integrates the character behavior data returned from the AI analysis engine with the emotion data from the emotion engine, and adjusts the character behavior scene to reflect the user's emotions.
[0211] Step 6:
[0212] The server generates story text based on the integrated data. For example, if the recognized emotion is "curious," the generated story text would be, "The character unfolded the map and began to search for the way with gusto."
[0213] Step 7:
[0214] The server transmits the generated character behavior data and the adjusted story text to the terminal.
[0215] Step 8:
[0216] The device analyzes the data it receives and displays the character's action scenes and story text on the screen.
[0217] Step 9:
[0218] Users can visually see how their prompts and emotions are reflected on the device screen.
[0219] In this way, character behaviors and story text are dynamically generated and displayed to the user based on user-entered prompts and recognized emotions. This process occurs in real time, providing the user with a highly personalized, interactive storytelling experience.
[0220] Example 2
[0221] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0222] Conventional character behavior generation systems generate behavior based only on prompts entered by the user, which means they are unable to provide a dynamic user experience that reflects the user's emotions. Furthermore, they are unable to adjust story text or character behavior based on the user's emotions, making it difficult to provide a personalized, highly interactive experience.
[0223] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving text input from a user, means for analyzing the received text input and generating an iconic character behavior, means for displaying the generated iconic character behavior, means for acquiring user emotion information, means for analyzing the acquired emotion information, and means for adjusting the iconic character behavior and the story text using the analyzed emotion information. This allows the character behavior and story text to be dynamically changed based on the user's emotion, enabling a more advanced and personalized user experience.
[0224] "Text input" refers to textual instructions or commands entered by a user using a terminal.
[0225] "Iconic character behavior" refers to the visual representation or animated character behavior that is generated based on the analyzed prompt and emotional information.
[0226] "Display means" refers to a system component for visually presenting the generated iconographic character actions and narrative text on the terminal screen.
[0227] "Emotional information" refers to the user's emotional state analyzed based on the user's facial expression, vocal tone, or other biometric signals.
[0228] "Means for acquiring" refers to a camera, microphone, or other sensors for collecting user emotional information from the terminal.
[0229] The "analysis means" refers to an analysis system or algorithm for processing the acquired emotion information to identify the user's emotion.
[0230] "Narrative" refers to a textual story generated based on the analyzed prompt and emotional information.
[0231] The present invention relates to a system that receives prompts from a user, analyzes them, generates and displays character actions, and also recognizes the user's emotions and reflects them in the character actions and story text. A detailed description of an embodiment of this system will be given below.
[0232] System configuration
[0233] This system mainly consists of a server, a terminal, a user, and an emotion engine. The server mainly analyzes prompts and generates character actions, while the emotion engine recognizes the user's emotions. The terminal is responsible for inputting prompts from the user, acquiring user emotional information, and displaying the received character actions and story text.
[0234] Hardware and software used
[0235] Hardware
[0236] Server: Use high-performance cloud servers (e.g., Amazon Web Services or Google Cloud Platform).
[0237] Device: Smartphone or personal computer (e.g., iPhone, Android device, Windows personal computer)
[0238] Emotion engine: Various input devices of the device, including sensors such as cameras and microphones.
[0239] software
[0240] Natural language processing engine: Generative AI model (e.g., OpenAI GPT-4)
[0241] Facial recognition technology: Microsoft Azure Face API, Google Cloud Vision API, etc.
[0242] Speech recognition technology: Google Speech-to-Text, IBM Watson Speech to Text, etc.
[0243] Specific Examples
[0244] Here is a concrete example of the system's processing flow. First, the user uses the terminal to input a prompt sentence. For example, they might input "Find a way to enter the forest" and press the send button. The terminal then sends this prompt sentence to the server. The terminal also uses the camera and microphone to obtain the user's emotional information and sends it to the emotion engine.
[0245] The emotion engine analyzes the user's emotional information, such as facial expressions and tone of voice, and sends the results to the server. The server receives this and analyzes the prompt using the AI analysis engine (generative AI model). Character behavior is generated based on the prompt analysis results and emotional information.
[0246] The server then fine-tunes the character's behavior based on the emotional information. For example, if the user's emotion is recognized as "curious," the character's behavior will be displayed as being playful. Furthermore, the AI analysis engine generates a narrative sentence such as, "The character unfolded the map and began to search for a path with gusto. After walking a little further, he excitedly found a hidden entrance to the forest."
[0247] The server sends the generated character actions and narrative text to the device, which receives them and displays them on the screen, allowing the user to see in real time how their prompts and emotions are reflected.
[0248] Examples of prompt statements
[0249] "Tell a story that scares the character."
[0250] "Discover the castle's secrets"
[0251] "Wait for further instructions"
[0252] In this way, the system provides a highly personalized interactive experience based on the user's prompts and emotional information.
[0253] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0254] Step 1:
[0255] The user inputs a prompt sentence into the terminal and presses the send button. For example, the user inputs "Find a way to enter the forest" and clicks the send button. The prompt sentence "Find a way to enter the forest" is obtained as input. The terminal sends this prompt sentence to the server. Furthermore, the terminal obtains the user's emotional information (e.g., facial recognition and voice tone) and sends it to the emotion engine. The emotional information is obtained from data provided by the user through the camera and microphone.
[0256] Step 2:
[0257] The server passes the received prompt to the AI analysis engine. The input is the prompt received by the server from the device: "Find a way into the forest." The AI analysis engine analyzes this prompt, understands its meaning, and identifies the user's intention. As a specific operation, the AI analysis engine determines the character's action to "take out a map and find a way." The resulting character action, "take out a map and find a way," is output.
[0258] Step 3:
[0259] The emotion engine receives and analyzes emotion information sent from the device. The input is emotion information such as the user's facial expression analysis and voice tone sent from the device. The emotion engine analyzes this and identifies the user's emotion. In concrete terms, the emotion engine analyzes the user's emotion as "curious." As a result, the analyzed emotion information "curious" is output.
[0260] Step 4:
[0261] The server integrates the character's actions from the AI analysis engine with the emotional information from the emotion engine, and adjusts the character's actions based on that. The inputs are the character's action "take out a map and search for a way" and the emotional information "curious." As a specific action, the server generates an animation of the character happily checking the map. As a result, the animation of the adjusted character's actions is output.
[0262] Step 5:
[0263] The AI analysis engine generates a narrative based on the character's actions and emotional information. The inputs are the character's action of "taking out a map and searching for a way" and emotional information of "curious." As a specific action, the AI analysis engine generates the narrative sentence, "The character unfolded the map and began to search for a way with gusto. After walking a little further, he excitedly found a hidden entrance to the forest." The resulting narrative sentence is output.
[0264] Step 6:
[0265] The server sends the generated animation of the character's actions and the narrative text to the terminal. The input is the animation of the character's actions and the narrative text that the server obtains from the AI analysis engine. The server sends this data to the terminal via the network. As a result, the data sent to the terminal is output.
[0266] Step 7:
[0267] The device displays the received animation of the character's actions and narrative text on the screen. The input is the animation of the character's actions and narrative text sent from the server. The device receives this and presents it visually to the user. In concrete terms, the device displays on the screen an animation of the character happily checking a map and the narrative text: "The character unfolded the map and began to search for a path with gusto. After walking a little further, he excitedly found a hidden entrance to the forest." As a result, the user can see a display that reflects their prompts and emotions in real time.
[0268] (Application example 2)
[0269] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0270] Conventional interactive content systems typically generate character behavior and story text based on simple prompts from the user, but are unable to dynamically change them to take into account the user's emotional information. This results in a lack of personalized user experience and reduced satisfaction. Furthermore, the story development tends to be fixed, making it difficult to provide real-time interactivity that responds to the user's emotions.
[0271] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving prompts from the user, means for analyzing the received prompts and generating character actions, means for displaying the generated character actions, means for acquiring user emotional information, means for analyzing the acquired emotional information, and means for reflecting the emotional information in character actions and story text. This allows the character actions and story text to be dynamically changed based on the user prompts and the emotional information analyzed in real time, enabling a highly personalized user experience.
[0272] "Means for receiving prompts from the user" refers to a function that allows the system to receive text or voice commands entered by the user.
[0273] The "means for analyzing the received prompt and generating character behavior" is a function that analyzes the received prompt using natural language processing technology and determines the specific behavior of the character based on the results.
[0274] The "means for displaying the generated character's actions" is a function for displaying the actions of the generated character on a screen or the like so that the user can visually confirm them.
[0275] The "means for acquiring user emotion information" is a function for collecting data for determining emotions from the user's facial expressions, tone of voice, etc.
[0276] The "means for analyzing acquired emotional information" is a function for analyzing collected emotional data of a user and identifying the emotional state of the user at that time.
[0277] "Means for reflecting emotional information in character behavior and story text" is a function that dynamically modifies character behavior and story text based on analyzed emotional information of the user.
[0278] The present invention is a system that receives prompts from a user, analyzes them, generates and displays character actions, and also recognizes the user's emotions and reflects them in the character actions and story text. Specific embodiments for implementing this system will be described below.
[0279] System Overview
[0280] This system mainly consists of a server, a terminal, a user, and an emotion engine. The server mainly analyzes prompts and generates character actions, while the emotion engine recognizes the user's emotions. The terminal is responsible for inputting prompts from the user, acquiring user emotion information, and displaying received character actions and story text.
[0281] Hardware and software used
[0282] The system configuration uses the following hardware and software:
[0283] Hardware: Smartphone or head-mounted display (with camera)
[0284] software:
[0285] DeepFace: A Python library for analyzing user emotions using facial recognition
[0286] OpenAI GPT-3: An AI model for natural language processing and narrative text generation
[0287] TextBlob: A Python library for simple natural language processing
[0288] Specific processing and data calculations
[0289] 1. Enter and submit the prompt
[0290] The user enters the prompt into the input field on the device and presses the send button, at which point the device captures the user's face with a camera and sends the facial recognition data to the emotion engine.
[0291] 2. Emotional Information Analysis
[0292] The facial recognition data sent by the device is analyzed by an emotion engine to identify the user's emotional state. For example, the DeepFace library can be used to detect emotions such as "curious" or "anxious" from the user's facial expressions.
[0293] 3. Prompt analysis and character behavior generation
[0294] The server analyzes the received prompts using OpenAI GPT-3 and generates specific character actions, such as "taking out a map and finding the way."
[0295] 4. Reflecting Emotional Information and Generating Story Text
[0296] Based on the analysis results of the emotion engine, the server reflects the user's emotional information in the character's behavior. Using OpenAI GPT-3, story text is generated based on the user's emotional state. For example, if the emotion "curious" is recognized, the element "enjoyed" is added to the story text.
[0297] 5. Display of character actions and story text
[0298] The generated character actions and story text are sent to the device and displayed on the device screen, allowing the user to see in real time how their prompts and emotions are reflected.
[0299] Examples of specific examples and prompts
[0300] Specific examples
[0301] When a user types the prompt "Find a way into the forest" into a smartphone app, their face is captured by a camera and the emotion is recognized as "curious." As a result, a scene is generated in which the character happily unfolds a map, and the story text reads, "The character unfolds the map and begins to joyfully search for a way. After walking a little further, he excitedly finds a hidden entrance into the forest."
[0302] Examples of prompt statements
[0303] If the user inputs the prompt "Find a way to save the princess" and the emotion is recognized as "anxiety," the action "The character hurriedly searches around the castle and finds the hidden door" is generated and displayed along with the story text "The character searches around the castle in an anxious manner and discovers the hidden door."
[0304] As described above, this system dynamically changes character behavior and story text based on user prompts and real-time analyzed emotional information, enabling a highly personalized user experience.
[0305] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0306] Step 1:
[0307] The user types a prompt into the terminal.
[0308] The user enters a prompt into the input field of the device and presses the send button. For example, the user enters "Find a way into the forest." The device receives the prompt. The device also captures the user's face using the built-in camera and collects the image data.
[0309] Input: User prompt and facial image data
[0310] Output: prompt data and face image data
[0311] Step 2:
[0312] The terminal transmits the facial image data to the emotion engine.
[0313] The device sends the collected facial image data to the emotion engine, which uses the DeepFace library to analyze emotions from the facial image. This analysis process identifies the user's emotional state. For example, an emotion such as "curious" may be detected.
[0314] Input: Facial image data
[0315] Output: Emotional information
[0316] Step 3:
[0317] The server analyzes the prompts and generates character actions.
[0318] The device sends the received prompt data to the server, which uses OpenAI GPT-3 to analyze the prompt and understand its meaning. Based on the analysis results, the server determines the character's specific actions. For example, it generates an action such as "take out a map and find the way."
[0319] Input: prompt data
[0320] Output: Character behavior
[0321] Step 4:
[0322] The server generates a story text that reflects the emotional information.
[0323] The server receives the emotion information from the emotion engine. The server then reflects this emotion information in the character's behavior and generates story text using OpenAI GPT-3. For example, if the user is recognized as "curious," the generated behavior is adjusted to "excitedly checks the map," and the story text generated reads, "The character unfolds the map and begins to search for the way with gusto. After walking a little further, he excitedly finds a hidden entrance to the forest."
[0324] Input: Character behavior, emotional information
[0325] Output: Story text
[0326] Step 5:
[0327] The device displays character actions and story text.
[0328] The server sends the generated character actions and story text to the device, which receives this data and displays it on the screen, allowing the user to see in real time how their prompts and emotions are reflected.
[0329] Input: Character actions, story text
[0330] Output: On-screen display
[0331] Through the above processing steps, the system can generate and display dynamically changing character behavior and story text based on the user's prompts and emotional information.
[0332] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0333] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0334] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0335] [Second embodiment]
[0336] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0337] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0338] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0339] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0340] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0341] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0342] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0343] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0344] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0345] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0346] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0347] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0348] The present invention is a system for receiving prompts from a user, analyzing the prompts, generating character actions, and displaying the actions to progress a story. Specific embodiments for implementing this system will be described below.
[0349] System Overview
[0350] This system operates mainly with a server, a terminal, and a user. The server mainly analyzes prompts and generates character actions, while the terminal is responsible for inputting prompts from the user and displaying the received character actions and story text.
[0351] Program processing
[0352] Entering and submitting prompts
[0353] First, the user inputs a prompt into the input field of the terminal and presses the send button, at which point the terminal sends the input prompt to the server.
[0354] Parsing prompts
[0355] The server passes the received prompt to an AI analysis engine, which uses natural language processing technology to understand the meaning of the prompt and determine the character's specific actions.
[0356] Character behavior and story text generation
[0357] Next, the AI analysis engine generates character actions based on the analysis results, and also generates story text related to the character actions. The server compiles this data and sends it to the device.
[0358] Screen display
[0359] The device displays the received character actions and story text on the screen, allowing the user to see in real time how their prompts are reflected.
[0360] Specific examples
[0361] Example 1: Adventure game
[0362] The user types in "Find a way to enter the forest" and presses the send button. The device sends this prompt to the server. The server uses an AI analysis engine to analyze the prompt "Find a way to enter the forest" and determines that the character should take the action of "taking out a map and starting to search for a way." The server then generates the story text, "The character unfolds the map and starts to search for a way. After walking a little further, he finds a hidden entrance into the forest." This data is sent to the device, which displays it on the screen.
[0363] Example 2: Mystery game
[0364] The user types "Find clues" and presses the send button. The device sends this prompt to the server. The server uses an AI analysis engine to analyze the prompt "Find clues" and determines the character's action to "begin to investigate the room in detail." Next, it generates story text that reads, "The character investigates the entire room and finds hidden clues." This data is sent to the device, which displays it on the screen.
[0365] As described above, the system receives and analyzes prompts, generates and displays character actions, and allows users to progress through the story, allowing users to hone their technical skills and creativity through interaction with AI.
[0366] The processing flow will be explained below.
[0367] Step 1:
[0368] The user enters a prompt (e.g., "Find a way into the forest") into the input field of the terminal and presses the send button.
[0369] Step 2:
[0370] The terminal sends the prompts entered by the user to the server, where the input data is converted into packets and sent over the network.
[0371] Step 3:
[0372] The server passes the received prompt to the AI analysis engine. Since the prompt is in text format and cannot be understood by the AI as is, it is analyzed via a natural language processing API.
[0373] Step 4:
[0374] The AI analysis engine analyzes the meaning of the received prompt. Specifically, it goes through processes such as grammatical analysis, keyword extraction, and semantic analysis to understand the user's instruction to "find a way into the forest."
[0375] Step 5:
[0376] The AI analysis engine generates specific character actions based on the prompt (e.g., "take out a map and find the way"), and this generated action data is returned to the server in a data structure such as JSON format.
[0377] Step 6:
[0378] The server receives the character behavior data returned from the AI analysis engine and generates corresponding story text (e.g., "The character unfolds the map and begins to search for a path. After walking a little further, he finds a hidden entrance into the forest.").
[0379] Step 7:
[0380] The server combines the generated character behavior data and story text into a single packet and sends it to the terminal.
[0381] Step 8:
[0382] The device analyzes the data it receives and prepares it for display on the screen, specifically generating character animations and formatting story text.
[0383] Step 9:
[0384] The device displays the character's action scenes and story text on the screen, allowing the user to visually confirm how their prompts have been reflected.
[0385] The above explains the processing content of the program of this system, broken down into specific processing steps. The present invention provides a new type of game in which the actions of an AI character are determined by user prompts and are displayed as a story.
[0386] Example 1
[0387] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0388] Conventional virtual character systems have struggled to respond flexibly and in real time to user input and progress the story. Furthermore, they lack the ability to reflect the user's creative input, resulting in a lack of entertainment value. This limits the user experience and makes it difficult to maintain sustained interest.
[0389] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0390] In this invention, the server includes means for receiving input data from a user, means for analyzing the received input data and generating virtual character behaviors, and means for displaying the generated virtual character behaviors, thereby enabling the server to analyze user prompts in real time and generate appropriate virtual character behaviors and story texts.
[0391] "User" refers to the person who operates the system and provides input data.
[0392] "Input data" refers to text or commands that a user enters into a system.
[0393] "Means for receiving" refers to the functions and processes for obtaining input data provided by a user.
[0394] "Means of analysis" refers to the function or process that understands received input data and processes it to extract meaning.
[0395] "Virtual character behavior" refers to a specific action taken by a virtual character that is generated based on the analysis results.
[0396] "Generating means" refers to the function or process that produces virtual character behavior and narrative text based on the analysis of input data.
[0397] "Display means" refers to a function or process for visually presenting the generated virtual character actions and narrative text to the user.
[0398] "Natural language processing technology" refers to technology that enables computers to understand, interpret, and generate human language.
[0399] "Story text" refers to the text of a story that is generated based on the actions of the generated virtual character.
[0400] A "system" refers to a collection of functions and processes realized by combining the above means.
[0401] The present invention is a system for receiving input data from a user, analyzing the data, and generating and displaying the behavior of a virtual character. Specific embodiments of the present invention will be described in detail below.
[0402] Program processing flow
[0403] This system mainly consists of three components: a server, a terminal, and a user. The server analyzes prompts and generates virtual character actions, while the terminal receives prompts from the user and displays the received virtual character actions and story text.
[0404] Hardware and software used
[0405] server
[0406] The server uses a high-performance computer, and uses software such as a generative AI model (e.g., GPT-3) to analyze prompts and generate virtual character behavior and narrative text.
[0407] Terminal
[0408] The terminals used can be devices capable of input and display, such as personal computers, smartphones, tablets, etc. The user inputs prompts through the terminal and views the results sent from the server.
[0409] Specific explanation of data processing and data calculation
[0410] 1. The user enters the prompt into the input field on the terminal and presses the send button.
[0411] 2. The terminal sends the entered prompt to the server using an HTTP request.
[0412] 3. The server passes the received prompt to a generative AI model (e.g., GPT-3) for analysis.
[0413] 4. The generative AI model uses natural language processing techniques to analyze the input prompt and understand its meaning.
[0414] 5. Based on the analysis results, the virtual character's actions are generated, along with narrative text related to the generated actions.
[0415] 6. The server sends the generated virtual character actions and story text to the device, again using an HTTP request.
[0416] 7. The terminal displays the received data on the screen, allowing the user to see how their prompts were reflected.
[0417] Specific examples
[0418] Example 1: Adventure game
[0419] The user types "Find a way into the forest" and presses the send button. The device sends this prompt to the server. The server uses a generative AI model (e.g., "GPT-3") to analyze the prompt "Find a way into the forest." Based on the analysis results, the server generates an action for the character to "take out a map and start looking for a way." Next, it generates narrative text such as "The character unfolds the map and starts looking for a way. After walking a little further, he finds a hidden entrance into the forest." This data is sent to the device, which displays it on the screen.
[0420] Example 2: Mystery game
[0421] The user types "Find clues" and presses the send button. The device sends this prompt to the server. The server uses a generative AI model (e.g., "GPT-3") to analyze the prompt "Find clues." Based on the analysis results, the server generates an action for the character to "begin to investigate the room in detail." Next, it generates narrative text such as "The character investigates the entire room and finds hidden clues." This data is sent to the device, which displays it on the screen.
[0422] As described above, the present invention is a system that receives and analyzes user prompts, generates and displays the actions of virtual characters, and allows users to progress through their own story, honing their technical skills and creativity through interaction with AI.
[0423] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0424] Step 1:
[0425] The user enters a prompt into the input field of the terminal and presses the send button. The input data is in the form of a specific command or question in text format. For example, "Find a way into the forest." This is the input data. The terminal receives the input data through the user's operation.
[0426] Step 2:
[0427] The terminal generates and sends an HTTP request to send the received input data to the server. The input data is a prompt sentence in text format, and is sent to the server as the payload of the HTTP request. This transfers the input data from the terminal to the server.
[0428] Step 3:
[0429] The server receives the HTTP request sent from the terminal and extracts the prompt text. The received data is the text data included in the payload of the HTTP request. This text data is passed to the next process for analysis.
[0430] Step 4:
[0431] The server passes the received prompt sentence to a generative AI model (for example, GPT-3) for analysis. The input is the text data of the prompt, and the output is the analysis result. The generative AI model uses natural language processing technology to analyze the meaning of the prompt and determine the specific actions of the virtual character. This analysis results in actions such as "taking out a map and starting to look for directions."
[0432] Step 5:
[0433] The server generates narrative text related to the virtual character's actions based on the analysis results obtained by the generative AI model. The input is the analysis results, and the output is the generated character's actions and narrative text. For example, the generated text might read, "The character unfolds the map and begins to search for a path. After walking a little further, he finds a hidden entrance into the forest."
[0434] Step 6:
[0435] The server sends the generated virtual character actions and story text to the terminal as an HTTP response. The input is the generated data, and the output is the HTTP response. This transfers data from the server to the terminal.
[0436] Step 7:
[0437] The device extracts the virtual character's actions and story text from the received HTTP response and displays them on the screen. The input is the received data, and the output is a display format that the user can visually confirm. Specifically, the screen displays the message, "The character unfolds the map and begins to search for a path. After walking a little further, he finds a hidden entrance into the forest." This allows the user to see in real time how the prompts they entered have been reflected.
[0438] As described above, specific data processing and data calculations are carried out at each step, and ultimately the system displays appropriate virtual character behavior and story text to the user.
[0439] (Application example 1)
[0440] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0441] A problem with virtual stores is the lack of a system that responds appropriately to customer requests and questions about products. Another issue is the lack of an interface that allows virtual store clerk characters to make suggestions tailored to the customer's needs. This can lead to a lack of improvement in the customer experience and a decrease in customer satisfaction.
[0442] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0443] In this invention, the server includes means for receiving prompts from a user, means for analyzing the received prompts and generating character actions and suggestions, means for displaying the generated character actions and suggestions, and means for making suggestions in response to requests from customers of the virtual store, thereby enabling the virtual store clerk character to take appropriate actions and make appropriate suggestions in response to requests and questions input by customers.
[0444] The "means for receiving prompts from a user" is an interface for acquiring requests or questions entered by customers of the virtual store via a digital device.
[0445] The "means for analyzing prompts and generating character behavior and proposal content" is a system that uses natural language processing technology to understand the input content from customers and generates appropriate behavior for a virtual store clerk character and product proposals based on that.
[0446] The "means for displaying the generated character's actions and proposals" refers to a display or user interface that visually conveys to the customer the character's actions and proposals generated based on the analysis results.
[0447] The "means for making proposals in response to customer requests in a virtual store" is a system that has the function of allowing a virtual store clerk character to make appropriate product proposals and provide information regarding the products and information that customers are looking for.
[0448] The present invention is a system for generating and displaying character actions and suggestions in response to customer requests and questions in a virtual store. A specific embodiment for implementing this system will be described.
[0449] System configuration
[0450] This system operates mainly with a server, a terminal, and a user. The server mainly analyzes prompts and generates character actions and suggestions, while the terminal is responsible for inputting prompts from the user and displaying character actions, suggestions, and story text received from the server.
[0451] Program processing
[0452] server
[0453] The server has an interface for receiving prompts from the user. It also analyzes the received prompts using natural language processing technology and generates the actions and suggestions of a virtual store clerk character based on the analysis results. A generative AI model is used to generate the character's actions and suggestions. The generated data is sent to the terminal.
[0454] Terminal
[0455] The terminal provides an input field for the user to enter a prompt. When the user enters and submits the prompt, the terminal sends it to the server. Upon receiving the analysis results from the server, the terminal displays the character's actions and suggestions, along with the associated story text.
[0456] User
[0457] The user inputs a request or question into the input field on the device. For example, they can type "Show me shoes that go with this dress" and press the send button. The user can then check the character's actions and suggestions sent from the server on the device screen.
[0458] Technical details
[0459] Hardware and software used
[0460] Hardware: Digital devices such as smartphones, tablets, and PCs
[0461] software:
[0462] Server side: Flask (a Python micro web framework)
[0463] Client-side: Web browser or dedicated application
[0464] Natural Language Processing: Generative AI models such as OpenAI GPT
[0465] Communication: requests module using HTTP protocol
[0466] Data processing and calculation
[0467] 1. The user types a prompt (such as "Show me the shoes that go with this dress") into the device's input field.
[0468] 2. The terminal sends the prompt to the server via an HTTP request.
[0469] 3. The server receives the prompt and passes it to a generative AI model for analysis, using natural language processing techniques.
[0470] 4. As a result of the analysis, character actions and suggested content are generated. For example, a character action such as "I selected three suggested shoes from the catalog and displayed them on the screen" and story text such as "The virtual salesperson selected three shoes from the catalog and explained the features of each pair and how they would match your dress" are generated.
[0471] 5. The server sends the generated data to the terminal as an HTTP response.
[0472] 6. The terminal displays the received data on the user interface, and the user confirms it.
[0473] As described above, the system of the present invention receives user prompts and generates appropriate character actions and suggestions based on the prompts, thereby improving the customer experience in a virtual store.
[0474] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0475] Step 1:
[0476] The user inputs a request or question into the input field of the terminal. For example, the user inputs a prompt sentence such as "Show me the shoes that go with this dress" and presses the send button. The input prompt sentence is provided to the terminal as input.
[0477] Step 2:
[0478] The terminal sends the prompt text entered by the user to the server. Specifically, it creates an HTTP request and sends the prompt text as a payload to the specified endpoint of the server. At this point, the input is the prompt text, and the output is the request data sent to the server.
[0479] Step 3:
[0480] The server analyzes the received prompt. Here, the server passes the prompt to a generative AI model (e.g., OpenAI GPT) for analysis. During this analysis process, the input is the prompt, and the output is the character's actions and suggestions as a result of the analysis.
[0481] Step 4:
[0482] The server generates character actions and suggestions based on the analysis results. For example, the analysis results might generate a character action such as "I selected three suggested shoes from the catalog and displayed them on the screen," and a story text such as "The virtual salesperson selected three shoes from the catalog and explained the features of each pair and how they would match your dress." The input for this step is the analysis results from the generative AI model, and the output is the generated character actions and story text.
[0483] Step 5:
[0484] The server sends the generated character actions and suggestions to the device. Specifically, it returns data created based on the analysis results to the device as an HTTP response. At this point, the input is the generated character actions and story text, and the output is the data sent to the device.
[0485] Step 6:
[0486] The device displays the character actions and suggestions received from the server on the user interface, allowing the user to visually confirm them. In this step, the device analyzes the received data and renders it as a view to be displayed on the screen. The input is the character actions and story text sent from the server, and the output is the visual data displayed on the user interface.
[0487] Through the above processing steps, a system is realized in which appropriate character actions and suggestions are provided to users in real time in response to customer requests in a virtual store.
[0488] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0489] The present invention is a system that receives prompts from a user, analyzes them, generates and displays character actions, and also recognizes the user's emotions and reflects them in the character actions and story text. Specific embodiments for implementing this system will be described below.
[0490] System Overview
[0491] This system mainly consists of a server, a terminal, a user, and an emotion engine. The server mainly analyzes prompts and generates character actions, while the emotion engine recognizes the user's emotions. The terminal is responsible for inputting prompts from the user, acquiring user emotion information, and displaying received character actions and story text.
[0492] Program processing
[0493] Entering and submitting prompts
[0494] First, the user enters a prompt (e.g., "Find a way into the forest") into an input field on the device and presses the send button. Then, the device sends the user's emotional information (e.g., facial recognition and voice tone analysis) to the emotion engine.
[0495] Parsing prompts
[0496] The server passes the received prompt to an AI analysis engine, which uses natural language processing technology to understand the meaning of the prompt and determine the character's specific actions.
[0497] Reflecting emotional information
[0498] The emotion engine analyzes the user's emotions and sends the results to the server. The server uses this emotional information to make fine adjustments to the character's behavior and story text. For example, if the user is excited, the character's behavior may be changed more dynamically.
[0499] Character behavior and story text generation
[0500] Next, the AI analysis engine generates character actions based on the prompt analysis results and emotional information. It also generates story text related to the character actions. The server then compiles this data and sends it to the device.
[0501] Screen display
[0502] The device displays the received character actions and story text on the screen, allowing the user to see in real time how their prompts and emotions are reflected.
[0503] Specific examples
[0504] Example 1: Adventure game
[0505] When the user types "Find a way into the forest" and presses the send button, the device simultaneously sends the user's facial recognition information to the emotion engine. The emotion engine recognizes the user's emotion as "curious" and sends this information to the server. The server uses its AI analysis engine to analyze the prompt, determines the action to "take out a map and find a way," and generates a scene in which the character happily checks the map, reflecting the emotion of "curious." Next, the server generates story text that reads, "The character unfolds the map and begins to search for a way with gusto. After walking a little further, he excitedly finds a hidden entrance to the forest." This data is sent to the device, which displays it on the screen.
[0506] Example 2: Mystery game
[0507] When the user types "find clues" and presses the send button, the device simultaneously sends the user's voice tone information to the emotion engine. The emotion engine recognizes the user's emotion as "anxiety" and sends this information to the server. The server uses its AI analysis engine to analyze the prompt, determines the action to "start investigating the room in detail," and generates a scene of the character moving quickly, reflecting the emotion of "anxiety." Next, story text is generated: "The character investigates the room anxiously and finds hidden clues." This data is sent to the device, which displays it on the screen.
[0508] The above explains the processing content of the program of this system, broken down into specific processing steps. The present invention provides a more advanced and personalized user experience by dynamically changing the behavior and story of the AI character according to the user's prompts and emotions.
[0509] The processing flow will be explained below.
[0510] Step 1:
[0511] The user enters a prompt (e.g., "Find a way into the forest") into the input field of the terminal and presses the send button.
[0512] Step 2:
[0513] The device sends the input prompt to the server, and at the same time, sends the user's facial recognition information and emotion data such as voice tone to the emotion engine.
[0514] Step 3:
[0515] The server receives the prompt and passes it to the AI analysis engine, which uses natural language processing to analyze the prompt and determine the character's actions based on its content.
[0516] Step 4:
[0517] The emotion engine analyzes the user's emotion data and sends the recognized emotional state (e.g., "curious" or "anxious") to the server.
[0518] Step 5:
[0519] The server integrates the character behavior data returned from the AI analysis engine with the emotion data from the emotion engine, and adjusts the character behavior scene to reflect the user's emotions.
[0520] Step 6:
[0521] The server generates story text based on the integrated data. For example, if the recognized emotion is "curious," the generated story text would be, "The character unfolded the map and began to search for the way with gusto."
[0522] Step 7:
[0523] The server transmits the generated character behavior data and the adjusted story text to the terminal.
[0524] Step 8:
[0525] The device analyzes the data it receives and displays the character's action scenes and story text on the screen.
[0526] Step 9:
[0527] Users can visually see how their prompts and emotions are reflected on the device screen.
[0528] In this way, character behaviors and story text are dynamically generated and displayed to the user based on user-entered prompts and recognized emotions. This process occurs in real time, providing the user with a highly personalized, interactive storytelling experience.
[0529] Example 2
[0530] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0531] Conventional character behavior generation systems generate behavior based only on prompts entered by the user, which means they are unable to provide a dynamic user experience that reflects the user's emotions. Furthermore, they are unable to adjust story text or character behavior based on the user's emotions, making it difficult to provide a personalized, highly interactive experience.
[0532] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving text input from a user, means for analyzing the received text input and generating an iconic character behavior, means for displaying the generated iconic character behavior, means for acquiring user emotion information, means for analyzing the acquired emotion information, and means for adjusting the iconic character behavior and the story text using the analyzed emotion information. This allows the character behavior and story text to be dynamically changed based on the user's emotion, enabling a more advanced and personalized user experience.
[0533] "Text input" refers to textual instructions or commands entered by a user using a terminal.
[0534] "Iconic character behavior" refers to the visual representation or animated character behavior that is generated based on the analyzed prompt and emotional information.
[0535] "Display means" refers to a system component for visually presenting the generated iconographic character actions and narrative text on the terminal screen.
[0536] "Emotional information" refers to the user's emotional state analyzed based on the user's facial expression, vocal tone, or other biometric signals.
[0537] "Means for acquiring" refers to a camera, microphone, or other sensors for collecting user emotional information from the terminal.
[0538] The "analysis means" refers to an analysis system or algorithm for processing the acquired emotion information to identify the user's emotion.
[0539] "Narrative" refers to a textual story generated based on the analyzed prompt and emotional information.
[0540] The present invention relates to a system that receives prompts from a user, analyzes them, generates and displays character actions, and also recognizes the user's emotions and reflects them in the character actions and story text. A detailed description of an embodiment of this system will be given below.
[0541] System configuration
[0542] This system mainly consists of a server, a terminal, a user, and an emotion engine. The server mainly analyzes prompts and generates character actions, while the emotion engine recognizes the user's emotions. The terminal is responsible for inputting prompts from the user, acquiring user emotional information, and displaying the received character actions and story text.
[0543] Hardware and software used
[0544] Hardware
[0545] Server: Use high-performance cloud servers (e.g., Amazon Web Services or Google Cloud Platform).
[0546] Device: Smartphone or personal computer (e.g., iPhone, Android device, Windows personal computer)
[0547] Emotion engine: Various input devices of the device, including sensors such as cameras and microphones.
[0548] software
[0549] Natural language processing engine: Generative AI model (e.g., OpenAI GPT-4)
[0550] Facial recognition technology: Microsoft Azure Face API, Google Cloud Vision API, etc.
[0551] Speech recognition technology: Google Speech-to-Text, IBM Watson Speech to Text, etc.
[0552] Specific Examples
[0553] Here is a concrete example of the system's processing flow. First, the user uses the terminal to input a prompt sentence. For example, they might input "Find a way to enter the forest" and press the send button. The terminal then sends this prompt sentence to the server. The terminal also uses the camera and microphone to obtain the user's emotional information and sends it to the emotion engine.
[0554] The emotion engine analyzes the user's emotional information, such as facial expressions and tone of voice, and sends the results to the server. The server receives this and analyzes the prompt using the AI analysis engine (generative AI model). Character behavior is generated based on the prompt analysis results and emotional information.
[0555] The server then fine-tunes the character's behavior based on the emotional information. For example, if the user's emotion is recognized as "curious," the character's behavior will be displayed as being playful. Furthermore, the AI analysis engine generates a narrative sentence such as, "The character unfolded the map and began to search for a path with gusto. After walking a little further, he excitedly found a hidden entrance to the forest."
[0556] The server sends the generated character actions and narrative text to the device, which receives them and displays them on the screen, allowing the user to see in real time how their prompts and emotions are reflected.
[0557] Examples of prompt statements
[0558] "Tell a story that scares the character."
[0559] "Discover the castle's secrets"
[0560] "Wait for further instructions"
[0561] In this way, the system provides a highly personalized interactive experience based on the user's prompts and emotional information.
[0562] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0563] Step 1:
[0564] The user inputs a prompt sentence into the terminal and presses the send button. For example, the user inputs "Find a way to enter the forest" and clicks the send button. The prompt sentence "Find a way to enter the forest" is obtained as input. The terminal sends this prompt sentence to the server. Furthermore, the terminal obtains the user's emotional information (e.g., facial recognition and voice tone) and sends it to the emotion engine. The emotional information is obtained from data provided by the user through the camera and microphone.
[0565] Step 2:
[0566] The server passes the received prompt to the AI analysis engine. The input is the prompt received by the server from the device: "Find a way into the forest." The AI analysis engine analyzes this prompt, understands its meaning, and identifies the user's intention. As a specific operation, the AI analysis engine determines the character's action to "take out a map and find a way." The resulting character action, "take out a map and find a way," is output.
[0567] Step 3:
[0568] The emotion engine receives and analyzes emotion information sent from the device. The input is emotion information such as the user's facial expression analysis and voice tone sent from the device. The emotion engine analyzes this and identifies the user's emotion. In concrete terms, the emotion engine analyzes the user's emotion as "curious." As a result, the analyzed emotion information "curious" is output.
[0569] Step 4:
[0570] The server integrates the character's actions from the AI analysis engine with the emotional information from the emotion engine, and adjusts the character's actions based on that. The inputs are the character's action "take out a map and search for a way" and the emotional information "curious." As a specific action, the server generates an animation of the character happily checking the map. As a result, the animation of the adjusted character's actions is output.
[0571] Step 5:
[0572] The AI analysis engine generates a narrative based on the character's actions and emotional information. The inputs are the character's action of "taking out a map and searching for a way" and emotional information of "curious." As a specific action, the AI analysis engine generates the narrative sentence, "The character unfolded the map and began to search for a way with gusto. After walking a little further, he excitedly found a hidden entrance to the forest." The resulting narrative sentence is output.
[0573] Step 6:
[0574] The server sends the generated animation of the character's actions and the narrative text to the terminal. The input is the animation of the character's actions and the narrative text that the server obtains from the AI analysis engine. The server sends this data to the terminal via the network. As a result, the data sent to the terminal is output.
[0575] Step 7:
[0576] The device displays the received animation of the character's actions and narrative text on the screen. The input is the animation of the character's actions and narrative text sent from the server. The device receives this and presents it visually to the user. In concrete terms, the device displays on the screen an animation of the character happily checking a map and the narrative text: "The character unfolded the map and began to search for a path with gusto. After walking a little further, he excitedly found a hidden entrance to the forest." As a result, the user can see a display that reflects their prompts and emotions in real time.
[0577] (Application example 2)
[0578] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0579] Conventional interactive content systems typically generate character behavior and story text based on simple prompts from the user, but are unable to dynamically change them to take into account the user's emotional information. This results in a lack of personalized user experience and reduced satisfaction. Furthermore, the story development tends to be fixed, making it difficult to provide real-time interactivity that responds to the user's emotions.
[0580] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving prompts from the user, means for analyzing the received prompts and generating character actions, means for displaying the generated character actions, means for acquiring user emotional information, means for analyzing the acquired emotional information, and means for reflecting the emotional information in character actions and story text. This allows the character actions and story text to be dynamically changed based on the user prompts and the emotional information analyzed in real time, enabling a highly personalized user experience.
[0581] "Means for receiving prompts from the user" refers to a function that allows the system to receive text or voice commands entered by the user.
[0582] The "means for analyzing the received prompt and generating character behavior" is a function that analyzes the received prompt using natural language processing technology and determines the specific behavior of the character based on the results.
[0583] The "means for displaying the generated character's actions" is a function for displaying the actions of the generated character on a screen or the like so that the user can visually confirm them.
[0584] The "means for acquiring user emotion information" is a function for collecting data for determining emotions from the user's facial expressions, tone of voice, etc.
[0585] The "means for analyzing acquired emotional information" is a function for analyzing collected emotional data of a user and identifying the emotional state of the user at that time.
[0586] "Means for reflecting emotional information in character behavior and story text" is a function that dynamically modifies character behavior and story text based on analyzed emotional information of the user.
[0587] The present invention is a system that receives prompts from a user, analyzes them, generates and displays character actions, and also recognizes the user's emotions and reflects them in the character actions and story text. Specific embodiments for implementing this system will be described below.
[0588] System Overview
[0589] This system mainly consists of a server, a terminal, a user, and an emotion engine. The server mainly analyzes prompts and generates character actions, while the emotion engine recognizes the user's emotions. The terminal is responsible for inputting prompts from the user, acquiring user emotion information, and displaying received character actions and story text.
[0590] Hardware and software used
[0591] The system configuration uses the following hardware and software:
[0592] Hardware: Smartphone or head-mounted display (with camera)
[0593] software:
[0594] DeepFace: A Python library for analyzing user emotions using facial recognition
[0595] OpenAI GPT-3: An AI model for natural language processing and narrative text generation
[0596] TextBlob: A Python library for simple natural language processing
[0597] Specific processing and data calculations
[0598] 1. Enter and submit the prompt
[0599] The user enters the prompt into the input field on the device and presses the send button, at which point the device captures the user's face with a camera and sends the facial recognition data to the emotion engine.
[0600] 2. Emotional Information Analysis
[0601] The facial recognition data sent by the device is analyzed by an emotion engine to identify the user's emotional state. For example, the DeepFace library can be used to detect emotions such as "curious" or "anxious" from the user's facial expressions.
[0602] 3. Prompt analysis and character behavior generation
[0603] The server analyzes the received prompts using OpenAI GPT-3 and generates specific character actions, such as "taking out a map and finding the way."
[0604] 4. Reflecting Emotional Information and Generating Story Text
[0605] Based on the analysis results of the emotion engine, the server reflects the user's emotional information in the character's behavior. Using OpenAI GPT-3, story text is generated based on the user's emotional state. For example, if the emotion "curious" is recognized, the element "enjoyed" is added to the story text.
[0606] 5. Display of character actions and story text
[0607] The generated character actions and story text are sent to the device and displayed on the device screen, allowing the user to see in real time how their prompts and emotions are reflected.
[0608] Examples of specific examples and prompts
[0609] Specific examples
[0610] When a user types the prompt "Find a way into the forest" into a smartphone app, their face is captured by a camera and the emotion is recognized as "curious." As a result, a scene is generated in which the character happily unfolds a map, and the story text reads, "The character unfolds the map and begins to joyfully search for a way. After walking a little further, he excitedly finds a hidden entrance into the forest."
[0611] Examples of prompt statements
[0612] If the user inputs the prompt "Find a way to save the princess" and the emotion is recognized as "anxiety," the action "The character hurriedly searches around the castle and finds the hidden door" is generated and displayed along with the story text "The character searches around the castle in an anxious manner and discovers the hidden door."
[0613] As described above, this system dynamically changes character behavior and story text based on user prompts and real-time analyzed emotional information, enabling a highly personalized user experience.
[0614] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0615] Step 1:
[0616] The user types a prompt into the terminal.
[0617] The user enters a prompt into the input field of the device and presses the send button. For example, the user enters "Find a way into the forest." The device receives the prompt. The device also captures the user's face using the built-in camera and collects the image data.
[0618] Input: User prompt and facial image data
[0619] Output: prompt data and face image data
[0620] Step 2:
[0621] The terminal transmits the facial image data to the emotion engine.
[0622] The device sends the collected facial image data to the emotion engine, which uses the DeepFace library to analyze emotions from the facial image. This analysis process identifies the user's emotional state. For example, an emotion such as "curious" may be detected.
[0623] Input: Facial image data
[0624] Output: Emotional information
[0625] Step 3:
[0626] The server analyzes the prompts and generates character actions.
[0627] The device sends the received prompt data to the server, which uses OpenAI GPT-3 to analyze the prompt and understand its meaning. Based on the analysis results, the server determines the character's specific actions. For example, it generates an action such as "take out a map and find the way."
[0628] Input: prompt data
[0629] Output: Character behavior
[0630] Step 4:
[0631] The server generates a story text that reflects the emotional information.
[0632] The server receives the emotion information from the emotion engine. The server then reflects this emotion information in the character's behavior and generates story text using OpenAI GPT-3. For example, if the user is recognized as "curious," the generated behavior is adjusted to "excitedly checks the map," and the story text generated reads, "The character unfolds the map and begins to search for the way with gusto. After walking a little further, he excitedly finds a hidden entrance to the forest."
[0633] Input: Character behavior, emotional information
[0634] Output: Story text
[0635] Step 5:
[0636] The device displays character actions and story text.
[0637] The server sends the generated character actions and story text to the device, which receives this data and displays it on the screen, allowing the user to see in real time how their prompts and emotions are reflected.
[0638] Input: Character actions, story text
[0639] Output: On-screen display
[0640] Through the above processing steps, the system can generate and display dynamically changing character behavior and story text based on the user's prompts and emotional information.
[0641] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0642] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0643] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0644] [Third embodiment]
[0645] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0646] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0647] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0648] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0649] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0650] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0651] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0652] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0653] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0654] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0655] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0656] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0657] The present invention is a system for receiving prompts from a user, analyzing the prompts, generating character actions, and displaying the actions to progress a story. Specific embodiments for implementing this system will be described below.
[0658] System Overview
[0659] This system operates mainly with a server, a terminal, and a user. The server mainly analyzes prompts and generates character actions, while the terminal is responsible for inputting prompts from the user and displaying the received character actions and story text.
[0660] Program processing
[0661] Entering and submitting prompts
[0662] First, the user inputs a prompt into the input field of the terminal and presses the send button, at which point the terminal sends the input prompt to the server.
[0663] Parsing prompts
[0664] The server passes the received prompt to an AI analysis engine, which uses natural language processing technology to understand the meaning of the prompt and determine the character's specific actions.
[0665] Character behavior and story text generation
[0666] Next, the AI analysis engine generates character actions based on the analysis results, and also generates story text related to the character actions. The server compiles this data and sends it to the device.
[0667] Screen display
[0668] The device displays the received character actions and story text on the screen, allowing the user to see in real time how their prompts are reflected.
[0669] Specific examples
[0670] Example 1: Adventure game
[0671] The user types in "Find a way to enter the forest" and presses the send button. The device sends this prompt to the server. The server uses an AI analysis engine to analyze the prompt "Find a way to enter the forest" and determines that the character should take the action of "taking out a map and starting to search for a way." The server then generates the story text, "The character unfolds the map and starts to search for a way. After walking a little further, he finds a hidden entrance into the forest." This data is sent to the device, which displays it on the screen.
[0672] Example 2: Mystery game
[0673] The user types "Find clues" and presses the send button. The device sends this prompt to the server. The server uses an AI analysis engine to analyze the prompt "Find clues" and determines the character's action to "begin to investigate the room in detail." Next, it generates story text that reads, "The character investigates the entire room and finds hidden clues." This data is sent to the device, which displays it on the screen.
[0674] As described above, the system receives and analyzes prompts, generates and displays character actions, and allows users to progress through the story, allowing users to hone their technical skills and creativity through interaction with AI.
[0675] The processing flow will be explained below.
[0676] Step 1:
[0677] The user enters a prompt (e.g., "Find a way into the forest") into the input field of the terminal and presses the send button.
[0678] Step 2:
[0679] The terminal sends the prompts entered by the user to the server, where the input data is converted into packets and sent over the network.
[0680] Step 3:
[0681] The server passes the received prompt to the AI analysis engine. Since the prompt is in text format and cannot be understood by the AI as is, it is analyzed via a natural language processing API.
[0682] Step 4:
[0683] The AI analysis engine analyzes the meaning of the received prompt. Specifically, it goes through processes such as grammatical analysis, keyword extraction, and semantic analysis to understand the user's instruction to "find a way into the forest."
[0684] Step 5:
[0685] The AI analysis engine generates specific character actions based on the prompt (e.g., "take out a map and find the way"), and this generated action data is returned to the server in a data structure such as JSON format.
[0686] Step 6:
[0687] The server receives the character behavior data returned from the AI analysis engine and generates corresponding story text (e.g., "The character unfolds the map and begins to search for a path. After walking a little further, he finds a hidden entrance into the forest.").
[0688] Step 7:
[0689] The server combines the generated character behavior data and story text into a single packet and sends it to the terminal.
[0690] Step 8:
[0691] The device analyzes the data it receives and prepares it for display on the screen, specifically generating character animations and formatting story text.
[0692] Step 9:
[0693] The device displays the character's action scenes and story text on the screen, allowing the user to visually confirm how their prompts have been reflected.
[0694] The above explains the processing content of the program of this system, broken down into specific processing steps. The present invention provides a new type of game in which the actions of an AI character are determined by user prompts and are displayed as a story.
[0695] Example 1
[0696] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0697] Conventional virtual character systems have struggled to respond flexibly and in real time to user input and progress the story. Furthermore, they lack the ability to reflect the user's creative input, resulting in a lack of entertainment value. This limits the user experience and makes it difficult to maintain sustained interest.
[0698] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0699] In this invention, the server includes means for receiving input data from a user, means for analyzing the received input data and generating virtual character behaviors, and means for displaying the generated virtual character behaviors, thereby enabling the server to analyze user prompts in real time and generate appropriate virtual character behaviors and story texts.
[0700] "User" refers to the person who operates the system and provides input data.
[0701] "Input data" refers to text or commands that a user enters into a system.
[0702] "Means for receiving" refers to the functions and processes for obtaining input data provided by a user.
[0703] "Means of analysis" refers to the function or process that understands received input data and processes it to extract meaning.
[0704] "Virtual character behavior" refers to a specific action taken by a virtual character that is generated based on the analysis results.
[0705] "Generating means" refers to the function or process that produces virtual character behavior and narrative text based on the analysis of input data.
[0706] "Display means" refers to a function or process for visually presenting the generated virtual character actions and narrative text to the user.
[0707] "Natural language processing technology" refers to technology that enables computers to understand, interpret, and generate human language.
[0708] "Story text" refers to the text of a story that is generated based on the actions of the generated virtual character.
[0709] A "system" refers to a collection of functions and processes realized by combining the above means.
[0710] The present invention is a system for receiving input data from a user, analyzing the data, and generating and displaying the behavior of a virtual character. Specific embodiments of the present invention will be described in detail below.
[0711] Program processing flow
[0712] This system mainly consists of three components: a server, a terminal, and a user. The server analyzes prompts and generates virtual character actions, while the terminal receives prompts from the user and displays the received virtual character actions and story text.
[0713] Hardware and software used
[0714] server
[0715] The server uses a high-performance computer, and uses software such as a generative AI model (e.g., GPT-3) to analyze prompts and generate virtual character behavior and narrative text.
[0716] Terminal
[0717] The terminals used can be devices capable of input and display, such as personal computers, smartphones, tablets, etc. The user inputs prompts through the terminal and views the results sent from the server.
[0718] Specific explanation of data processing and data calculation
[0719] 1. The user enters the prompt into the input field on the terminal and presses the send button.
[0720] 2. The terminal sends the entered prompt to the server using an HTTP request.
[0721] 3. The server passes the received prompt to a generative AI model (e.g., GPT-3) for analysis.
[0722] 4. The generative AI model uses natural language processing techniques to analyze the input prompt and understand its meaning.
[0723] 5. Based on the analysis results, the virtual character's actions are generated, along with narrative text related to the generated actions.
[0724] 6. The server sends the generated virtual character actions and story text to the device, again using an HTTP request.
[0725] 7. The terminal displays the received data on the screen, allowing the user to see how their prompts were reflected.
[0726] Specific examples
[0727] Example 1: Adventure game
[0728] The user types "Find a way into the forest" and presses the send button. The device sends this prompt to the server. The server uses a generative AI model (e.g., "GPT-3") to analyze the prompt "Find a way into the forest." Based on the analysis results, the server generates an action for the character to "take out a map and start looking for a way." Next, it generates narrative text such as "The character unfolds the map and starts looking for a way. After walking a little further, he finds a hidden entrance into the forest." This data is sent to the device, which displays it on the screen.
[0729] Example 2: Mystery game
[0730] The user types "Find clues" and presses the send button. The device sends this prompt to the server. The server uses a generative AI model (e.g., "GPT-3") to analyze the prompt "Find clues." Based on the analysis results, the server generates an action for the character to "begin to investigate the room in detail." Next, it generates narrative text such as "The character investigates the entire room and finds hidden clues." This data is sent to the device, which displays it on the screen.
[0731] As described above, the present invention is a system that receives and analyzes user prompts, generates and displays the actions of virtual characters, and allows users to progress through their own story, honing their technical skills and creativity through interaction with AI.
[0732] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0733] Step 1:
[0734] The user enters a prompt into the input field of the terminal and presses the send button. The input data is in the form of a specific command or question in text format. For example, "Find a way into the forest." This is the input data. The terminal receives the input data through the user's operation.
[0735] Step 2:
[0736] The terminal generates and sends an HTTP request to send the received input data to the server. The input data is a prompt sentence in text format, and is sent to the server as the payload of the HTTP request. This transfers the input data from the terminal to the server.
[0737] Step 3:
[0738] The server receives the HTTP request sent from the terminal and extracts the prompt text. The received data is the text data included in the payload of the HTTP request. This text data is passed to the next process for analysis.
[0739] Step 4:
[0740] The server passes the received prompt sentence to a generative AI model (for example, GPT-3) for analysis. The input is the text data of the prompt, and the output is the analysis result. The generative AI model uses natural language processing technology to analyze the meaning of the prompt and determine the specific actions of the virtual character. This analysis results in actions such as "taking out a map and starting to look for directions."
[0741] Step 5:
[0742] The server generates narrative text related to the virtual character's actions based on the analysis results obtained by the generative AI model. The input is the analysis results, and the output is the generated character's actions and narrative text. For example, the generated text might read, "The character unfolds the map and begins to search for a path. After walking a little further, he finds a hidden entrance into the forest."
[0743] Step 6:
[0744] The server sends the generated virtual character actions and story text to the terminal as an HTTP response. The input is the generated data, and the output is the HTTP response. This transfers data from the server to the terminal.
[0745] Step 7:
[0746] The device extracts the virtual character's actions and story text from the received HTTP response and displays them on the screen. The input is the received data, and the output is a display format that the user can visually confirm. Specifically, the screen displays the message, "The character unfolds the map and begins to search for a path. After walking a little further, he finds a hidden entrance into the forest." This allows the user to see in real time how the prompts they entered have been reflected.
[0747] As described above, specific data processing and data calculations are carried out at each step, and ultimately the system displays appropriate virtual character behavior and story text to the user.
[0748] (Application example 1)
[0749] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0750] A problem with virtual stores is the lack of a system that responds appropriately to customer requests and questions about products. Another issue is the lack of an interface that allows virtual store clerk characters to make suggestions tailored to the customer's needs. This can lead to a lack of improvement in the customer experience and a decrease in customer satisfaction.
[0751] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0752] In this invention, the server includes means for receiving prompts from a user, means for analyzing the received prompts and generating character actions and suggestions, means for displaying the generated character actions and suggestions, and means for making suggestions in response to requests from customers of the virtual store, thereby enabling the virtual store clerk character to take appropriate actions and make appropriate suggestions in response to requests and questions input by customers.
[0753] The "means for receiving prompts from a user" is an interface for acquiring requests or questions entered by customers of the virtual store via a digital device.
[0754] The "means for analyzing prompts and generating character behavior and proposal content" is a system that uses natural language processing technology to understand the input content from customers and generates appropriate virtual store clerk character behavior and product proposals based on that.
[0755] The "means for displaying the generated character's actions and proposals" refers to a display or user interface that visually conveys to the customer the character's actions and proposals generated based on the analysis results.
[0756] The "means for making proposals in response to customer requests in a virtual store" is a system that has the function of allowing a virtual store clerk character to make appropriate product proposals and provide information regarding the products and information that customers are looking for.
[0757] The present invention is a system for generating and displaying character actions and suggestions in response to customer requests and questions in a virtual store. A specific embodiment for implementing this system will be described.
[0758] System configuration
[0759] This system operates mainly with a server, a terminal, and a user. The server mainly analyzes prompts and generates character actions and suggestions, while the terminal is responsible for inputting prompts from the user and displaying character actions, suggestions, and story text received from the server.
[0760] Program processing
[0761] server
[0762] The server has an interface for receiving prompts from the user. It also analyzes the received prompts using natural language processing technology and generates the actions and suggestions of a virtual store clerk character based on the analysis results. A generative AI model is used to generate the character's actions and suggestions. The generated data is sent to the terminal.
[0763] Terminal
[0764] The terminal provides an input field for the user to enter a prompt. When the user enters and submits the prompt, the terminal sends it to the server. After receiving the analysis results from the server, the terminal displays the character's actions and suggestions, along with the associated story text.
[0765] User
[0766] The user inputs a request or question into the device's input field. For example, they can type "Show me shoes that go with this dress" and press the send button. The user can then check the character's actions and suggestions sent from the server on the device screen.
[0767] Technical details
[0768] Hardware and software used
[0769] Hardware: Digital devices such as smartphones, tablets, and PCs
[0770] software:
[0771] Server side: Flask (a Python micro web framework)
[0772] Client-side: Web browser or dedicated application
[0773] Natural Language Processing: Generative AI models such as OpenAI GPT
[0774] Communication: requests module using HTTP protocol
[0775] Data processing and calculation
[0776] 1. The user types a prompt (such as "Show me the shoes that go with this dress") into the device's input field.
[0777] 2. The terminal sends the prompt to the server via an HTTP request.
[0778] 3. The server receives the prompt and passes it to a generative AI model for analysis, using natural language processing techniques.
[0779] 4. As a result of the analysis, character actions and suggested content are generated. For example, a character action such as "I selected three suggested shoes from the catalog and displayed them on the screen" and story text such as "The virtual salesperson selected three shoes from the catalog and explained the features of each pair and how they would match your dress" are generated.
[0780] 5. The server sends the generated data to the terminal as an HTTP response.
[0781] 6. The terminal displays the received data on the user interface, and the user confirms it.
[0782] As described above, the system of the present invention receives user prompts and generates appropriate character actions and suggestions based on the prompts, thereby improving the customer experience in a virtual store.
[0783] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0784] Step 1:
[0785] The user inputs a request or question into the input field of the terminal. For example, the user inputs a prompt sentence such as "Show me the shoes that go with this dress" and presses the send button. The input prompt sentence is provided to the terminal as input.
[0786] Step 2:
[0787] The terminal sends the prompt text entered by the user to the server. Specifically, it creates an HTTP request and sends the prompt text as a payload to the specified endpoint of the server. At this point, the input is the prompt text, and the output is the request data sent to the server.
[0788] Step 3:
[0789] The server analyzes the received prompt. Here, the server passes the prompt to a generative AI model (e.g., OpenAI GPT) for analysis. During this analysis process, the input is the prompt, and the output is the character's actions and suggestions as a result of the analysis.
[0790] Step 4:
[0791] The server generates character actions and suggestions based on the analysis results. For example, the analysis results might generate a character action such as "I selected three suggested shoes from the catalog and displayed them on the screen," and a story text such as "The virtual salesperson selected three shoes from the catalog and explained the features of each pair and how they would match your dress." The input for this step is the analysis results from the generative AI model, and the output is the generated character actions and story text.
[0792] Step 5:
[0793] The server sends the generated character actions and suggestions to the device. Specifically, it returns data created based on the analysis results to the device as an HTTP response. At this point, the input is the generated character actions and story text, and the output is the data sent to the device.
[0794] Step 6:
[0795] The device displays the character actions and suggestions received from the server on the user interface, allowing the user to visually confirm them. In this step, the device analyzes the received data and renders it as a view to be displayed on the screen. The input is the character actions and story text sent from the server, and the output is the visual data displayed on the user interface.
[0796] Through the above processing steps, a system is realized in which appropriate character actions and suggestions are provided to users in real time in response to customer requests in a virtual store.
[0797] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0798] The present invention is a system that receives prompts from a user, analyzes them, generates and displays character actions, and also recognizes the user's emotions and reflects them in the character actions and story text. Specific embodiments for implementing this system will be described below.
[0799] System Overview
[0800] This system mainly consists of a server, a terminal, a user, and an emotion engine. The server mainly analyzes prompts and generates character actions, while the emotion engine recognizes the user's emotions. The terminal is responsible for inputting prompts from the user, acquiring user emotion information, and displaying received character actions and story text.
[0801] Program processing
[0802] Entering and submitting prompts
[0803] First, the user enters a prompt (e.g., "Find a way to enter the forest") into an input field on the device and presses the send button. Then, the device sends the user's emotion information (e.g., facial recognition and voice tone analysis) to the emotion engine.
[0804] Parsing prompts
[0805] The server passes the received prompt to an AI analysis engine, which uses natural language processing technology to understand the meaning of the prompt and determine the character's specific actions.
[0806] Reflecting emotional information
[0807] The emotion engine analyzes the user's emotions and sends the results to the server. The server uses this emotional information to make fine adjustments to the character's behavior and story text. For example, if the user is excited, the character's behavior may be changed more dynamically.
[0808] Character behavior and story text generation
[0809] Next, the AI analysis engine generates character actions based on the prompt analysis results and emotional information. It also generates story text related to the character actions. The server then compiles this data and sends it to the device.
[0810] Screen display
[0811] The device displays the received character actions and story text on the screen, allowing the user to see in real time how their prompts and emotions are reflected.
[0812] Specific examples
[0813] Example 1: Adventure game
[0814] When the user types "Find a way into the forest" and presses the send button, the device simultaneously sends the user's facial recognition information to the emotion engine. The emotion engine recognizes the user's emotion as "curious" and sends this information to the server. The server uses its AI analysis engine to analyze the prompt, determines the action to "take out a map and find a way," and generates a scene in which the character happily checks the map, reflecting the emotion of "curious." Next, the server generates story text that reads, "The character unfolds the map and begins to search for a way with gusto. After walking a little further, he excitedly finds a hidden entrance to the forest." This data is sent to the device, which displays it on the screen.
[0815] Example 2: Mystery game
[0816] When the user types "find clues" and presses the send button, the device simultaneously sends the user's voice tone information to the emotion engine. The emotion engine recognizes the user's emotion as "anxiety" and sends this information to the server. The server uses its AI analysis engine to analyze the prompt, determines the action to "start investigating the room in detail," and generates a scene of the character moving quickly, reflecting the emotion of "anxiety." Next, story text is generated: "The character investigates the room anxiously and finds hidden clues." This data is sent to the device, which displays it on the screen.
[0817] The above explains the processing content of the program of this system, broken down into specific processing steps. The present invention provides a more advanced and personalized user experience by dynamically changing the behavior and story of the AI character according to the user's prompts and emotions.
[0818] The processing flow will be explained below.
[0819] Step 1:
[0820] The user enters a prompt (e.g., "Find a way into the forest") into the input field of the terminal and presses the send button.
[0821] Step 2:
[0822] The device sends the input prompt to the server, and at the same time, sends the user's facial recognition information and emotion data such as voice tone to the emotion engine.
[0823] Step 3:
[0824] The server receives the prompt and passes it to the AI analysis engine, which uses natural language processing to analyze the prompt and determine the character's behavior based on its content.
[0825] Step 4:
[0826] The emotion engine analyzes the user's emotion data and sends the recognized emotional state (e.g., "curious" or "anxious") to the server.
[0827] Step 5:
[0828] The server integrates the character behavior data returned from the AI analysis engine with the emotion data from the emotion engine, and adjusts the character behavior scene to reflect the user's emotions.
[0829] Step 6:
[0830] The server generates story text based on the integrated data. For example, if the recognized emotion is "curious," the generated story text would be, "The character unfolded the map and began to search for the way with gusto."
[0831] Step 7:
[0832] The server transmits the generated character behavior data and the adjusted story text to the terminal.
[0833] Step 8:
[0834] The device analyzes the data it receives and displays the character's action scenes and story text on the screen.
[0835] Step 9:
[0836] Users can visually see how their prompts and emotions are reflected on the device screen.
[0837] In this way, character behaviors and story text are dynamically generated and displayed to the user based on user-entered prompts and recognized emotions. This process occurs in real time, providing the user with a highly personalized, interactive storytelling experience.
[0838] Example 2
[0839] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0840] Conventional character behavior generation systems generate behavior based only on prompts entered by the user, which means they are unable to provide a dynamic user experience that reflects the user's emotions. Furthermore, they are unable to adjust story text or character behavior based on the user's emotions, making it difficult to provide a personalized, highly interactive experience.
[0841] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving text input from a user, means for analyzing the received text input and generating an iconic character behavior, means for displaying the generated iconic character behavior, means for acquiring user emotion information, means for analyzing the acquired emotion information, and means for adjusting the iconic character behavior and the story text using the analyzed emotion information. This allows the character behavior and story text to be dynamically changed based on the user's emotion, enabling a more advanced and personalized user experience.
[0842] "Text input" refers to textual instructions or commands entered by a user using a terminal.
[0843] "Iconic character behavior" refers to the visual representation and animated character behavior generated based on the analyzed prompt and emotional information.
[0844] "Display means" refers to a system component for visually presenting the generated iconographic character actions and narrative text on the terminal screen.
[0845] "Emotional information" refers to the user's emotional state analyzed based on the user's facial expression, vocal tone, or other bio-signals.
[0846] "Means for acquiring" refers to a camera, microphone, or other sensors for collecting user emotional information from the terminal.
[0847] The "analysis means" refers to an analysis system or algorithm for processing the acquired emotion information to identify the user's emotion.
[0848] "Narrative" refers to a textual story generated based on the analyzed prompt and emotional information.
[0849] The present invention relates to a system that receives prompts from a user, analyzes them, generates and displays character actions, and also recognizes the user's emotions and reflects them in the character actions and story text. A detailed description of an embodiment of this system will be given below.
[0850] System configuration
[0851] This system mainly consists of a server, a terminal, a user, and an emotion engine. The server mainly analyzes prompts and generates character actions, while the emotion engine recognizes the user's emotions. The terminal is responsible for inputting prompts from the user, acquiring user emotional information, and displaying the received character actions and story text.
[0852] Hardware and software used
[0853] Hardware
[0854] Server: Use high-performance cloud servers (e.g., Amazon Web Services or Google Cloud Platform).
[0855] Device: Smartphone or personal computer (e.g., iPhone, Android device, Windows personal computer)
[0856] Emotion engine: Various input devices of the device, including sensors such as cameras and microphones.
[0857] software
[0858] Natural language processing engine: Generative AI model (e.g., OpenAI GPT-4)
[0859] Facial recognition technology: Microsoft Azure Face API, Google Cloud Vision API, etc.
[0860] Speech recognition technology: Google Speech-to-Text, IBM Watson Speech to Text, etc.
[0861] Specific Examples
[0862] Here is a concrete example of the system's processing flow. First, the user uses the terminal to input a prompt sentence. For example, they might input "Find a way to enter the forest" and press the send button. The terminal then sends this prompt sentence to the server. The terminal also uses the camera and microphone to obtain the user's emotional information and sends it to the emotion engine.
[0863] The emotion engine analyzes the user's emotional information, such as facial expressions and tone of voice, and sends the results to the server. The server receives this and analyzes the prompt using the AI analysis engine (generative AI model). Character behavior is generated based on the prompt analysis results and emotional information.
[0864] The server then fine-tunes the character's behavior based on the emotional information. For example, if the user's emotion is recognized as "curious," the character's behavior will be displayed as being playful. Furthermore, the AI analysis engine generates a narrative sentence such as, "The character unfolded the map and began to search for a path with gusto. After walking a little further, he excitedly found a hidden entrance to the forest."
[0865] The server sends the generated character actions and narrative text to the device, which receives them and displays them on the screen, allowing the user to see in real time how their prompts and emotions are reflected.
[0866] Examples of prompt statements
[0867] "Tell a story that scares the character."
[0868] "Discover the castle's secrets"
[0869] "Wait for further instructions"
[0870] In this way, the system provides a highly personalized interactive experience based on the user's prompts and emotional information.
[0871] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0872] Step 1:
[0873] The user inputs a prompt sentence into the terminal and presses the send button. For example, the user inputs "Find a way to enter the forest" and clicks the send button. The prompt sentence "Find a way to enter the forest" is obtained as input. The terminal sends this prompt sentence to the server. Furthermore, the terminal obtains the user's emotional information (e.g., facial recognition and voice tone) and sends it to the emotion engine. The emotional information is obtained from data provided by the user through the camera and microphone.
[0874] Step 2:
[0875] The server passes the received prompt to the AI analysis engine. The input is the prompt received by the server from the device: "Find a way into the forest." The AI analysis engine analyzes this prompt, understands its meaning, and identifies the user's intention. As a specific operation, the AI analysis engine determines the character's action to "take out a map and find a way." The resulting character action, "take out a map and find a way," is output.
[0876] Step 3:
[0877] The emotion engine receives and analyzes emotion information sent from the device. The input is emotion information such as the user's facial expression analysis and voice tone sent from the device. The emotion engine analyzes this and identifies the user's emotion. In concrete terms, the emotion engine analyzes the user's emotion as "curious." As a result, the analyzed emotion information "curious" is output.
[0878] Step 4:
[0879] The server integrates the character's actions from the AI analysis engine with the emotional information from the emotion engine, and adjusts the character's actions based on that. The inputs are the character's action "take out a map and search for a way" and the emotional information "curious." As a specific action, the server generates an animation of the character happily checking the map. As a result, the animation of the adjusted character's actions is output.
[0880] Step 5:
[0881] The AI analysis engine generates a narrative based on the character's actions and emotional information. The inputs are the character's action of "taking out a map and searching for a way" and emotional information of "curious." As a specific action, the AI analysis engine generates the narrative sentence, "The character unfolded the map and began to search for a way with gusto. After walking a little further, he excitedly found a hidden entrance to the forest." The resulting narrative sentence is output.
[0882] Step 6:
[0883] The server sends the generated animation of the character's actions and the narrative text to the terminal. The input is the animation of the character's actions and the narrative text that the server obtains from the AI analysis engine. The server sends this data to the terminal via the network. As a result, the data sent to the terminal is output.
[0884] Step 7:
[0885] The device displays the received animation of the character's actions and narrative text on the screen. The input is the animation of the character's actions and narrative text sent from the server. The device receives this and presents it visually to the user. In concrete terms, the device displays on the screen an animation of the character happily checking a map and the narrative text: "The character unfolded the map and began to search for a path with gusto. After walking a little further, he excitedly found a hidden entrance to the forest." As a result, the user can see a display that reflects their prompts and emotions in real time.
[0886] (Application example 2)
[0887] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0888] Conventional interactive content systems typically generate character behavior and story text based on simple prompts from the user, but are unable to dynamically change them to take into account the user's emotional information. This results in a lack of personalized user experience and reduced satisfaction. Furthermore, the story development tends to be fixed, making it difficult to provide real-time interactivity that responds to the user's emotions.
[0889] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving prompts from the user, means for analyzing the received prompts and generating character actions, means for displaying the generated character actions, means for acquiring user emotional information, means for analyzing the acquired emotional information, and means for reflecting the emotional information in character actions and story text. This allows the character actions and story text to be dynamically changed based on the user prompts and the emotional information analyzed in real time, enabling a highly personalized user experience.
[0890] "Means for receiving prompts from the user" refers to a function that allows the system to receive text or voice commands entered by the user.
[0891] The "means for analyzing the received prompt and generating character behavior" is a function that analyzes the received prompt using natural language processing technology and determines the specific behavior of the character based on the results.
[0892] The "means for displaying the generated character's actions" is a function for displaying the actions of the generated character on a screen or the like so that the user can visually confirm them.
[0893] The "means for acquiring user emotion information" is a function for collecting data for determining emotions from the user's facial expressions, tone of voice, etc.
[0894] The "means for analyzing acquired emotional information" is a function for analyzing collected emotional data of a user and identifying the emotional state of the user at that time.
[0895] "Means for reflecting emotional information in character behavior and story text" is a function that dynamically modifies character behavior and story text based on analyzed emotional information of the user.
[0896] The present invention is a system that receives prompts from a user, analyzes them, generates and displays character actions, and also recognizes the user's emotions and reflects them in the character actions and story text. Specific embodiments for implementing this system will be described below.
[0897] System Overview
[0898] This system mainly consists of a server, a terminal, a user, and an emotion engine. The server mainly analyzes prompts and generates character actions, while the emotion engine recognizes the user's emotions. The terminal is responsible for inputting prompts from the user, acquiring user emotion information, and displaying received character actions and story text.
[0899] Hardware and software used
[0900] The system configuration uses the following hardware and software:
[0901] Hardware: Smartphone or head-mounted display (with camera)
[0902] software:
[0903] DeepFace: A Python library for analyzing user emotions using facial recognition
[0904] OpenAI GPT-3: An AI model for natural language processing and narrative text generation
[0905] TextBlob: A Python library for simple natural language processing
[0906] Specific processing and data calculations
[0907] 1. Enter and submit the prompt
[0908] The user enters the prompt into the input field on the device and presses the send button, at which point the device captures the user's face with a camera and sends the facial recognition data to the emotion engine.
[0909] 2. Emotional Information Analysis
[0910] The facial recognition data sent by the device is analyzed by an emotion engine to identify the user's emotional state. For example, the DeepFace library can be used to detect emotions such as "curious" or "anxious" from the user's facial expressions.
[0911] 3. Prompt analysis and character behavior generation
[0912] The server analyzes the received prompts using OpenAI GPT-3 and generates specific character actions, such as "taking out a map and finding the way."
[0913] 4. Reflecting Emotional Information and Generating Story Text
[0914] Based on the analysis results of the emotion engine, the server reflects the user's emotional information in the character's behavior. Using OpenAI GPT-3, story text is generated based on the user's emotional state. For example, if the emotion "curious" is recognized, the element "enjoyed" is added to the story text.
[0915] 5. Display of character actions and story text
[0916] The generated character actions and story text are sent to the device and displayed on the device screen, allowing the user to see in real time how their prompts and emotions are reflected.
[0917] Examples of specific examples and prompts
[0918] Specific examples
[0919] When a user types the prompt "Find a way into the forest" into a smartphone app, their face is captured by a camera and the emotion is recognized as "curious." As a result, a scene is generated in which the character happily unfolds a map, and the story text reads, "The character unfolds the map and begins to joyfully search for a way. After walking a little further, he excitedly finds a hidden entrance into the forest."
[0920] Examples of prompt statements
[0921] If the user inputs the prompt "Find a way to save the princess" and the emotion is recognized as "anxiety," the action "The character hurriedly searches around the castle and finds the hidden door" is generated and displayed along with the story text "The character searches around the castle in an anxious manner and discovers the hidden door."
[0922] As described above, this system dynamically changes character behavior and story text based on user prompts and real-time analyzed emotional information, enabling a highly personalized user experience.
[0923] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0924] Step 1:
[0925] The user types a prompt into the terminal.
[0926] The user enters a prompt into the input field of the device and presses the send button. For example, the user enters "Find a way into the forest." The device receives the prompt. The device also captures the user's face using the built-in camera and collects the image data.
[0927] Input: User prompt and facial image data
[0928] Output: prompt data and face image data
[0929] Step 2:
[0930] The terminal transmits the facial image data to the emotion engine.
[0931] The device sends the collected facial image data to the emotion engine, which uses the DeepFace library to analyze emotions from the facial image. This analysis process identifies the user's emotional state. For example, an emotion such as "curious" may be detected.
[0932] Input: Facial image data
[0933] Output: Emotional information
[0934] Step 3:
[0935] The server analyzes the prompts and generates character actions.
[0936] The device sends the received prompt data to the server, which uses OpenAI GPT-3 to analyze the prompt and understand its meaning. Based on the analysis results, the server determines the character's specific actions. For example, it generates an action such as "take out a map and find the way."
[0937] Input: prompt data
[0938] Output: Character behavior
[0939] Step 4:
[0940] The server generates a story text that reflects the emotional information.
[0941] The server receives the emotion information from the emotion engine. The server then reflects this emotion information in the character's behavior and generates story text using OpenAI GPT-3. For example, if the user is recognized as "curious," the generated behavior is adjusted to "excitedly checks the map," and the story text generated reads, "The character unfolds the map and begins to search for the way with gusto. After walking a little further, he excitedly finds a hidden entrance to the forest."
[0942] Input: Character behavior, emotional information
[0943] Output: Story text
[0944] Step 5:
[0945] The device displays character actions and story text.
[0946] The server sends the generated character actions and story text to the device, which receives this data and displays it on the screen, allowing the user to see in real time how their prompts and emotions are reflected.
[0947] Input: Character actions, story text
[0948] Output: On-screen display
[0949] Through the above processing steps, the system can generate and display dynamically changing character behavior and story text based on the user's prompts and emotional information.
[0950] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0951] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0952] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[0953] [Fourth embodiment]
[0954] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[0955] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0956] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0957] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[0958] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0959] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0960] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0961] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[0962] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0963] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0964] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0965] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0966] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0967] The present invention is a system for receiving prompts from a user, analyzing the prompts, generating character actions, and displaying the actions to progress a story. Specific embodiments for implementing this system will be described below.
[0968] System Overview
[0969] This system operates mainly with a server, a terminal, and a user. The server mainly analyzes prompts and generates character actions, while the terminal is responsible for inputting prompts from the user and displaying the received character actions and story text.
[0970] Program processing
[0971] Entering and submitting prompts
[0972] First, the user inputs a prompt into the input field of the terminal and presses the send button, at which point the terminal sends the input prompt to the server.
[0973] Parsing prompts
[0974] The server passes the received prompt to an AI analysis engine, which uses natural language processing technology to understand the meaning of the prompt and determine the character's specific actions.
[0975] Character behavior and story text generation
[0976] Next, the AI analysis engine generates character actions based on the analysis results, and also generates story text related to the character actions. The server compiles this data and sends it to the device.
[0977] Screen display
[0978] The device displays the received character actions and story text on the screen, allowing the user to see in real time how their prompts are reflected.
[0979] Specific examples
[0980] Example 1: Adventure game
[0981] The user types in "Find a way to enter the forest" and presses the send button. The device sends this prompt to the server. The server uses an AI analysis engine to analyze the prompt "Find a way to enter the forest" and determines that the character should take the action of "taking out a map and starting to search for a way." The server then generates the story text, "The character unfolds the map and starts to search for a way. After walking a little further, he finds a hidden entrance into the forest." This data is sent to the device, which displays it on the screen.
[0982] Example 2: Mystery game
[0983] The user types "Find clues" and presses the send button. The device sends this prompt to the server. The server uses an AI analysis engine to analyze the prompt "Find clues" and determines the character's action to "begin to investigate the room in detail." Next, it generates story text that reads, "The character investigates the entire room and finds hidden clues." This data is sent to the device, which displays it on the screen.
[0984] As described above, the system receives and analyzes prompts, generates and displays character actions, and allows users to progress through the story, allowing users to hone their technical skills and creativity through interaction with AI.
[0985] The processing flow will be explained below.
[0986] Step 1:
[0987] The user enters a prompt (e.g., "Find a way into the forest") into the input field of the terminal and presses the send button.
[0988] Step 2:
[0989] The terminal sends the prompts entered by the user to the server, where the input data is converted into packets and sent over the network.
[0990] Step 3:
[0991] The server passes the received prompt to the AI analysis engine. Since the prompt is in text format and cannot be understood by the AI as is, it is analyzed via a natural language processing API.
[0992] Step 4:
[0993] The AI analysis engine analyzes the meaning of the received prompt. Specifically, it goes through processes such as grammatical analysis, keyword extraction, and semantic analysis to understand the user's instruction to "find a way into the forest."
[0994] Step 5:
[0995] The AI analysis engine generates specific character actions based on the prompt (e.g., "take out a map and find the way"), and this generated action data is returned to the server in a data structure such as JSON format.
[0996] Step 6:
[0997] The server receives the character behavior data returned from the AI analysis engine and generates corresponding story text (e.g., "The character unfolds the map and begins to search for a path. After walking a little further, he finds a hidden entrance into the forest.").
[0998] Step 7:
[0999] The server combines the generated character behavior data and story text into a single packet and sends it to the terminal.
[1000] Step 8:
[1001] The device analyzes the data it receives and prepares it for display on the screen, specifically generating character animations and formatting story text.
[1002] Step 9:
[1003] The device displays the character's action scenes and story text on the screen, allowing the user to visually confirm how their prompts have been reflected.
[1004] The above explains the processing content of the program of this system, broken down into specific processing steps. The present invention provides a new type of game in which the actions of an AI character are determined by user prompts and are displayed as a story.
[1005] Example 1
[1006] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1007] Conventional virtual character systems have struggled to respond flexibly and in real time to user input and progress the story. Furthermore, they lack the ability to reflect the user's creative input, resulting in a lack of entertainment value. This limits the user experience and makes it difficult to maintain sustained interest.
[1008] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1009] In this invention, the server includes means for receiving input data from a user, means for analyzing the received input data and generating virtual character behaviors, and means for displaying the generated virtual character behaviors, thereby enabling the server to analyze user prompts in real time and generate appropriate virtual character behaviors and story texts.
[1010] "User" refers to the person who operates the system and provides input data.
[1011] "Input data" refers to text or commands that a user enters into a system.
[1012] "Means for receiving" refers to the functions and processes for obtaining input data provided by a user.
[1013] "Means of analysis" refers to the function or process that understands received input data and processes it to extract meaning.
[1014] "Virtual character behavior" refers to a specific action taken by a virtual character that is generated based on the analysis results.
[1015] "Generating means" refers to the function or process that produces virtual character behavior and narrative text based on the analysis of input data.
[1016] "Display means" refers to a function or process for visually presenting the generated virtual character actions and narrative text to the user.
[1017] "Natural language processing technology" refers to technology that enables computers to understand, interpret, and generate human language.
[1018] "Story text" refers to the text of a story that is generated based on the actions of the generated virtual character.
[1019] A "system" refers to a collection of functions and processes realized by combining the above means.
[1020] The present invention is a system for receiving input data from a user, analyzing the data, and generating and displaying the behavior of a virtual character. Specific embodiments of the present invention will be described in detail below.
[1021] Program processing flow
[1022] This system mainly consists of three components: a server, a terminal, and a user. The server analyzes prompts and generates virtual character actions, while the terminal receives prompts from the user and displays the received virtual character actions and story text.
[1023] Hardware and software used
[1024] server
[1025] The server uses a high-performance computer, and uses software such as a generative AI model (e.g., GPT-3) to analyze prompts and generate virtual character behavior and narrative text.
[1026] Terminal
[1027] The terminals used can be devices capable of input and display, such as personal computers, smartphones, tablets, etc. The user inputs prompts through the terminal and views the results sent from the server.
[1028] Specific explanation of data processing and data calculation
[1029] 1. The user enters the prompt into the input field on the terminal and presses the send button.
[1030] 2. The terminal sends the entered prompt to the server using an HTTP request.
[1031] 3. The server passes the received prompt to a generative AI model (e.g., GPT-3) for analysis.
[1032] 4. The generative AI model uses natural language processing techniques to analyze the input prompt and understand its meaning.
[1033] 5. Based on the analysis results, the virtual character's actions are generated, along with narrative text related to the generated actions.
[1034] 6. The server sends the generated virtual character actions and story text to the device, again using an HTTP request.
[1035] 7. The terminal displays the received data on the screen, allowing the user to see how their prompts were reflected.
[1036] Specific examples
[1037] Example 1: Adventure game
[1038] The user types "Find a way into the forest" and presses the send button. The device sends this prompt to the server. The server uses a generative AI model (e.g., "GPT-3") to analyze the prompt "Find a way into the forest." Based on the analysis results, the server generates an action for the character to "take out a map and start looking for a way." Next, it generates narrative text such as "The character unfolds the map and starts looking for a way. After walking a little further, he finds a hidden entrance into the forest." This data is sent to the device, which displays it on the screen.
[1039] Example 2: Mystery game
[1040] The user types "Find clues" and presses the send button. The device sends this prompt to the server. The server uses a generative AI model (e.g., "GPT-3") to analyze the prompt "Find clues." Based on the analysis results, the server generates an action for the character to "begin to investigate the room in detail." Next, it generates narrative text such as "The character investigates the entire room and finds hidden clues." This data is sent to the device, which displays it on the screen.
[1041] As described above, the present invention is a system that receives and analyzes user prompts, generates and displays the actions of virtual characters, and allows users to progress through their own story, honing their technical skills and creativity through interaction with AI.
[1042] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1043] Step 1:
[1044] The user enters a prompt into the input field of the terminal and presses the send button. The input data is in the form of a specific command or question in text format. For example, "Find a way into the forest." This is the input data. The terminal receives the input data through the user's operation.
[1045] Step 2:
[1046] The terminal generates and sends an HTTP request to send the received input data to the server. The input data is a prompt sentence in text format, and is sent to the server as the payload of the HTTP request. This transfers the input data from the terminal to the server.
[1047] Step 3:
[1048] The server receives the HTTP request sent from the terminal and extracts the prompt text. The received data is the text data included in the payload of the HTTP request. This text data is passed to the next process for analysis.
[1049] Step 4:
[1050] The server passes the received prompt sentence to a generative AI model (for example, GPT-3) for analysis. The input is the text data of the prompt, and the output is the analysis result. The generative AI model uses natural language processing technology to analyze the meaning of the prompt and determine the specific actions of the virtual character. This analysis results in actions such as "taking out a map and starting to look for directions."
[1051] Step 5:
[1052] The server generates narrative text related to the virtual character's actions based on the analysis results obtained by the generative AI model. The input is the analysis results, and the output is the generated character's actions and narrative text. For example, the generated text might read, "The character unfolds the map and begins to search for a path. After walking a little further, he finds a hidden entrance into the forest."
[1053] Step 6:
[1054] The server sends the generated virtual character actions and story text to the terminal as an HTTP response. The input is the generated data, and the output is the HTTP response. This transfers data from the server to the terminal.
[1055] Step 7:
[1056] The device extracts the virtual character's actions and story text from the received HTTP response and displays them on the screen. The input is the received data, and the output is a display format that the user can visually confirm. Specifically, the screen displays the message, "The character unfolds the map and begins to search for a path. After walking a little further, he finds a hidden entrance into the forest." This allows the user to see in real time how the prompts they entered have been reflected.
[1057] As described above, specific data processing and data calculations are carried out at each step, and ultimately the system displays appropriate virtual character behavior and story text to the user.
[1058] (Application example 1)
[1059] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1060] A problem with virtual stores is the lack of a system that responds appropriately to customer requests and questions about products. Another issue is the lack of an interface that allows virtual store clerk characters to make suggestions tailored to the customer's needs. This can lead to a lack of improvement in the customer experience and a decrease in customer satisfaction.
[1061] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1062] In this invention, the server includes means for receiving prompts from a user, means for analyzing the received prompts and generating character actions and suggestions, means for displaying the generated character actions and suggestions, and means for making suggestions in response to requests from customers of the virtual store, thereby enabling the virtual store clerk character to take appropriate actions and make appropriate suggestions in response to requests and questions input by customers.
[1063] The "means for receiving prompts from a user" is an interface for acquiring requests or questions entered by customers of the virtual store via a digital device.
[1064] The "means for analyzing prompts and generating character behavior and proposal content" is a system that uses natural language processing technology to understand the input content from customers and generates appropriate virtual store clerk character behavior and product proposals based on that.
[1065] The "means for displaying the generated character's actions and proposals" refers to a display or user interface that visually conveys to the customer the character's actions and proposals generated based on the analysis results.
[1066] The "means for making proposals in response to customer requests in a virtual store" is a system that has the function of allowing a virtual store clerk character to make appropriate product proposals and provide information regarding the products and information that customers are looking for.
[1067] The present invention is a system for generating and displaying character actions and suggestions in response to customer requests and questions in a virtual store. A specific embodiment for implementing this system will be described.
[1068] System configuration
[1069] This system operates mainly with a server, a terminal, and a user. The server mainly analyzes prompts and generates character actions and suggestions, while the terminal is responsible for inputting prompts from the user and displaying character actions, suggestions, and story text received from the server.
[1070] Program processing
[1071] server
[1072] The server has an interface for receiving prompts from the user. It also analyzes the received prompts using natural language processing technology and generates the actions and suggestions of a virtual store clerk character based on the analysis results. A generative AI model is used to generate the character's actions and suggestions. The generated data is sent to the terminal.
[1073] Terminal
[1074] The terminal provides an input field for the user to enter a prompt. When the user enters and submits the prompt, the terminal sends it to the server. After receiving the analysis results from the server, the terminal displays the character's actions and suggestions, along with the associated story text.
[1075] User
[1076] The user inputs a request or question into the device's input field. For example, they can type "Show me shoes that go with this dress" and press the send button. The user can then check the character's actions and suggestions sent from the server on the device screen.
[1077] Technical details
[1078] Hardware and software used
[1079] Hardware: Digital devices such as smartphones, tablets, and PCs
[1080] software:
[1081] Server side: Flask (a Python micro web framework)
[1082] Client-side: Web browser or dedicated application
[1083] Natural Language Processing: Generative AI models such as OpenAI GPT
[1084] Communication: requests module using HTTP protocol
[1085] Data processing and calculation
[1086] 1. The user types a prompt (such as "Show me the shoes that go with this dress") into the device's input field.
[1087] 2. The terminal sends the prompt to the server via an HTTP request.
[1088] 3. The server receives the prompt and passes it to a generative AI model for analysis, using natural language processing techniques.
[1089] 4. As a result of the analysis, character actions and suggested content are generated. For example, a character action such as "I selected three suggested shoes from the catalog and displayed them on the screen" and story text such as "The virtual salesperson selected three shoes from the catalog and explained the features of each pair and how they would match your dress" are generated.
[1090] 5. The server sends the generated data to the terminal as an HTTP response.
[1091] 6. The terminal displays the received data on the user interface, and the user confirms it.
[1092] As described above, the system of the present invention receives user prompts and generates appropriate character actions and suggestions based on the prompts, thereby improving the customer experience in a virtual store.
[1093] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1094] Step 1:
[1095] The user inputs a request or question into the input field of the terminal. For example, the user inputs a prompt sentence such as "Show me the shoes that go with this dress" and presses the send button. The input prompt sentence is provided to the terminal as input.
[1096] Step 2:
[1097] The terminal sends the prompt text entered by the user to the server. Specifically, it creates an HTTP request and sends the prompt text as a payload to the specified endpoint of the server. At this point, the input is the prompt text, and the output is the request data sent to the server.
[1098] Step 3:
[1099] The server analyzes the received prompt. Here, the server passes the prompt to a generative AI model (e.g., OpenAI GPT) for analysis. During this analysis process, the input is the prompt, and the output is the character's actions and suggestions as a result of the analysis.
[1100] Step 4:
[1101] The server generates character actions and suggestions based on the analysis results. For example, the analysis results might generate a character action such as "I selected three suggested shoes from the catalog and displayed them on the screen," and a story text such as "The virtual salesperson selected three shoes from the catalog and explained the features of each pair and how they would match your dress." The input for this step is the analysis results from the generative AI model, and the output is the generated character actions and story text.
[1102] Step 5:
[1103] The server sends the generated character actions and suggestions to the device. Specifically, it returns data created based on the analysis results to the device as an HTTP response. At this point, the input is the generated character actions and story text, and the output is the data sent to the device.
[1104] Step 6:
[1105] The device displays the character actions and suggestions received from the server on the user interface, allowing the user to visually confirm them. In this step, the device analyzes the received data and renders it as a view to be displayed on the screen. The input is the character actions and story text sent from the server, and the output is the visual data displayed on the user interface.
[1106] Through the above processing steps, a system is realized in which appropriate character actions and suggestions are provided to users in real time in response to customer requests in a virtual store.
[1107] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1108] The present invention is a system that receives prompts from a user, analyzes them, generates and displays character actions, and also recognizes the user's emotions and reflects them in the character actions and story text. Specific embodiments for implementing this system will be described below.
[1109] System Overview
[1110] This system mainly consists of a server, a terminal, a user, and an emotion engine. The server mainly analyzes prompts and generates character actions, while the emotion engine recognizes the user's emotions. The terminal is responsible for inputting prompts from the user, acquiring user emotion information, and displaying received character actions and story text.
[1111] Program processing
[1112] Entering and submitting prompts
[1113] First, the user enters a prompt (e.g., "Find a way to enter the forest") into an input field on the device and presses the send button. Then, the device sends the user's emotion information (e.g., facial recognition and voice tone analysis) to the emotion engine.
[1114] Parsing prompts
[1115] The server passes the received prompt to an AI analysis engine, which uses natural language processing technology to understand the meaning of the prompt and determine the character's specific actions.
[1116] Reflecting emotional information
[1117] The emotion engine analyzes the user's emotions and sends the results to the server. The server uses this emotional information to make fine adjustments to the character's behavior and story text. For example, if the user is excited, the character's behavior may be changed more dynamically.
[1118] Character behavior and story text generation
[1119] Next, the AI analysis engine generates character actions based on the prompt analysis results and emotional information. It also generates story text related to the character actions. The server then compiles this data and sends it to the device.
[1120] Screen display
[1121] The device displays the received character actions and story text on the screen, allowing the user to see in real time how their prompts and emotions are reflected.
[1122] Specific examples
[1123] Example 1: Adventure game
[1124] When the user types "Find a way into the forest" and presses the send button, the device simultaneously sends the user's facial recognition information to the emotion engine. The emotion engine recognizes the user's emotion as "curious" and sends this information to the server. The server uses its AI analysis engine to analyze the prompt, determines the action to "take out a map and find a way," and generates a scene in which the character happily checks the map, reflecting the emotion of "curious." Next, the server generates story text that reads, "The character unfolds the map and begins to search for a way with gusto. After walking a little further, he excitedly finds a hidden entrance to the forest." This data is sent to the device, which displays it on the screen.
[1125] Example 2: Mystery game
[1126] When the user types "find clues" and presses the send button, the device simultaneously sends the user's voice tone information to the emotion engine. The emotion engine recognizes the user's emotion as "anxiety" and sends this information to the server. The server uses its AI analysis engine to analyze the prompt, determines the action to "start investigating the room in detail," and generates a scene of the character moving quickly, reflecting the emotion of "anxiety." Next, story text is generated: "The character investigates the room anxiously and finds hidden clues." This data is sent to the device, which displays it on the screen.
[1127] The above explains the processing content of the program of this system, broken down into specific processing steps. The present invention provides a more advanced and personalized user experience by dynamically changing the behavior and story of the AI character according to the user's prompts and emotions.
[1128] The processing flow will be explained below.
[1129] Step 1:
[1130] The user enters a prompt (e.g., "Find a way into the forest") into the input field of the terminal and presses the send button.
[1131] Step 2:
[1132] The device sends the input prompt to the server, and at the same time, sends the user's facial recognition information and emotion data such as voice tone to the emotion engine.
[1133] Step 3:
[1134] The server receives the prompt and passes it to the AI analysis engine, which uses natural language processing to analyze the prompt and determine the character's behavior based on its content.
[1135] Step 4:
[1136] The emotion engine analyzes the user's emotion data and sends the recognized emotional state (e.g., "curious" or "anxious") to the server.
[1137] Step 5:
[1138] The server integrates the character behavior data returned from the AI analysis engine with the emotion data from the emotion engine, and adjusts the character behavior scene to reflect the user's emotions.
[1139] Step 6:
[1140] The server generates story text based on the integrated data. For example, if the recognized emotion is "curious," the generated story text would be, "The character unfolded the map and began to search for the way with gusto."
[1141] Step 7:
[1142] The server transmits the generated character behavior data and the adjusted story text to the terminal.
[1143] Step 8:
[1144] The device analyzes the data it receives and displays the character's action scenes and story text on the screen.
[1145] Step 9:
[1146] Users can visually see how their prompts and emotions are reflected on the device screen.
[1147] In this way, character behaviors and story text are dynamically generated and displayed to the user based on user-entered prompts and recognized emotions. This process occurs in real time, providing the user with a highly personalized, interactive storytelling experience.
[1148] Example 2
[1149] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1150] Conventional character behavior generation systems generate behavior based only on prompts entered by the user, which means they are unable to provide a dynamic user experience that reflects the user's emotions. Furthermore, they are unable to adjust story text or character behavior based on the user's emotions, making it difficult to provide a personalized, highly interactive experience.
[1151] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving text input from a user, means for analyzing the received text input and generating an iconic character behavior, means for displaying the generated iconic character behavior, means for acquiring user emotion information, means for analyzing the acquired emotion information, and means for adjusting the iconic character behavior and the story text using the analyzed emotion information. This allows the character behavior and story text to be dynamically changed based on the user's emotion, enabling a more advanced and personalized user experience.
[1152] "Text input" refers to textual instructions or commands entered by a user using a terminal.
[1153] "Iconic character behavior" refers to the visual representation and animated character behavior generated based on the analyzed prompt and emotional information.
[1154] "Display means" refers to a system component for visually presenting the generated iconographic character actions and narrative text on the terminal screen.
[1155] "Emotional information" refers to the user's emotional state analyzed based on the user's facial expression, vocal tone, or other bio-signals.
[1156] "Means for acquiring" refers to a camera, microphone, or other sensors for collecting user emotional information from the terminal.
[1157] The "analysis means" refers to an analysis system or algorithm for processing the acquired emotion information to identify the user's emotion.
[1158] "Narrative" refers to a textual story generated based on the analyzed prompt and emotional information.
[1159] The present invention relates to a system that receives prompts from a user, analyzes them, generates and displays character actions, and also recognizes the user's emotions and reflects them in the character actions and story text. A detailed description of an embodiment of this system will be given below.
[1160] System configuration
[1161] This system mainly consists of a server, a terminal, a user, and an emotion engine. The server mainly analyzes prompts and generates character actions, while the emotion engine recognizes the user's emotions. The terminal is responsible for inputting prompts from the user, acquiring user emotional information, and displaying the received character actions and story text.
[1162] Hardware and software used
[1163] Hardware
[1164] Server: Use high-performance cloud servers (e.g., Amazon Web Services or Google Cloud Platform).
[1165] Device: Smartphone or personal computer (e.g., iPhone, Android device, Windows personal computer)
[1166] Emotion engine: Various input devices of the device, including sensors such as cameras and microphones.
[1167] software
[1168] Natural language processing engine: Generative AI model (e.g., OpenAI GPT-4)
[1169] Facial recognition technology: Microsoft Azure Face API, Google Cloud Vision API, etc.
[1170] Speech recognition technology: Google Speech-to-Text, IBM Watson Speech to Text, etc.
[1171] Specific Examples
[1172] Here is a concrete example of the system's processing flow. First, the user uses the terminal to input a prompt sentence. For example, they might input "Find a way to enter the forest" and press the send button. The terminal then sends this prompt sentence to the server. The terminal also uses the camera and microphone to obtain the user's emotional information and sends it to the emotion engine.
[1173] The emotion engine analyzes the user's emotional information, such as facial expressions and tone of voice, and sends the results to the server. The server receives this and analyzes the prompt using the AI analysis engine (generative AI model). Character behavior is generated based on the prompt analysis results and emotional information.
[1174] The server then fine-tunes the character's behavior based on the emotional information. For example, if the user's emotion is recognized as "curious," the character's behavior will be displayed as being playful. Furthermore, the AI analysis engine generates a narrative sentence such as, "The character unfolded the map and began to search for a path with gusto. After walking a little further, he excitedly found a hidden entrance to the forest."
[1175] The server sends the generated character actions and story text to the device, which receives them and displays them on the screen, allowing the user to see in real time how their prompts and emotions are reflected.
[1176] Examples of prompt statements
[1177] "Tell a story that scares the character."
[1178] "Discover the castle's secrets"
[1179] "Wait for further instructions"
[1180] In this way, the system provides a highly personalized interactive experience based on the user's prompts and emotional information.
[1181] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1182] Step 1:
[1183] The user inputs a prompt sentence into the terminal and presses the send button. For example, the user inputs "Find a way to enter the forest" and clicks the send button. The prompt sentence "Find a way to enter the forest" is obtained as input. The terminal sends this prompt sentence to the server. Furthermore, the terminal obtains the user's emotional information (e.g., facial recognition and voice tone) and sends it to the emotion engine. The emotional information is obtained from data provided by the user through the camera and microphone.
[1184] Step 2:
[1185] The server passes the received prompt to the AI analysis engine. The input is the prompt received by the server from the device: "Find a way into the forest." The AI analysis engine analyzes this prompt, understands its meaning, and identifies the user's intention. As a specific operation, the AI analysis engine determines the character's action to "take out a map and find a way." The resulting character action, "take out a map and find a way," is output.
[1186] Step 3:
[1187] The emotion engine receives and analyzes emotion information sent from the device. The input is emotion information such as the user's facial expression analysis and voice tone sent from the device. The emotion engine analyzes this and identifies the user's emotion. In concrete terms, the emotion engine analyzes the user's emotion as "curious." As a result, the analyzed emotion information "curious" is output.
[1188] Step 4:
[1189] The server integrates the character's actions from the AI analysis engine with the emotional information from the emotion engine, and adjusts the character's actions based on that. The inputs are the character's action "take out a map and search for a way" and the emotional information "curious." As a specific action, the server generates an animation of the character happily checking the map. As a result, the animation of the adjusted character's actions is output.
[1190] Step 5:
[1191] The AI analysis engine generates a narrative based on the character's actions and emotional information. The inputs are the character's action of "taking out a map and searching for a way" and emotional information of "curious." As a specific action, the AI analysis engine generates the narrative sentence, "The character unfolded the map and began to search for a way with gusto. After walking a little further, he excitedly found a hidden entrance to the forest." The resulting narrative sentence is output.
[1192] Step 6:
[1193] The server sends the generated animation of the character's actions and the narrative text to the terminal. The input is the animation of the character's actions and the narrative text that the server obtains from the AI analysis engine. The server sends this data to the terminal via the network. As a result, the data sent to the terminal is output.
[1194] Step 7:
[1195] The device displays the received animation of the character's actions and narrative text on the screen. The input is the animation of the character's actions and narrative text sent from the server. The device receives this and presents it visually to the user. In concrete terms, the device displays on the screen an animation of the character happily checking a map and the narrative text: "The character unfolded the map and began to search for a path with gusto. After walking a little further, he excitedly found a hidden entrance to the forest." As a result, the user can see a display that reflects their prompts and emotions in real time.
[1196] (Application example 2)
[1197] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1198] Conventional interactive content systems typically generate character behavior and story text based on simple prompts from the user, but are unable to dynamically change them to take into account the user's emotional information. This results in a lack of personalized user experience and reduced satisfaction. Furthermore, the story development tends to be fixed, making it difficult to provide real-time interactivity that responds to the user's emotions.
[1199] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving a prompt from the user, means for analyzing the received prompt and generating a character behavior, means for displaying the generated character behavior, means for acquiring user emotional information, means for analyzing the acquired emotional information, and means for reflecting the emotional information in the character behavior and story text. This allows the character behavior and story text to be dynamically changed based on the user prompt and the emotional information analyzed in real time, enabling a highly personalized user experience.
[1200] "Means for receiving prompts from the user" refers to a function that allows the system to receive text or voice commands entered by the user.
[1201] The "means for analyzing the received prompt and generating character behavior" is a function that analyzes the received prompt using natural language processing technology and determines the specific behavior of the character based on the results.
[1202] The "means for displaying the generated character's actions" is a function for displaying the actions of the generated character on a screen or the like so that the user can visually confirm them.
[1203] The "means for acquiring user emotion information" is a function for collecting data for determining emotions from the user's facial expressions, tone of voice, etc.
[1204] The "means for analyzing acquired emotional information" is a function for analyzing collected emotional data of a user and identifying the emotional state of the user at that time.
[1205] "Means for reflecting emotional information in character behavior and story text" is a function that dynamically modifies character behavior and story text based on analyzed emotional information of the user.
[1206] The present invention is a system that receives prompts from a user, analyzes them, generates and displays character actions, and also recognizes the user's emotions and reflects them in the character actions and story text. Specific embodiments for implementing this system will be described below.
[1207] System Overview
[1208] This system mainly consists of a server, a terminal, a user, and an emotion engine. The server mainly analyzes prompts and generates character actions, while the emotion engine recognizes the user's emotions. The terminal is responsible for inputting prompts from the user, acquiring user emotion information, and displaying received character actions and story text.
[1209] Hardware and software used
[1210] The system configuration uses the following hardware and software:
[1211] Hardware: Smartphone or head-mounted display (with camera)
[1212] software:
[1213] DeepFace: A Python library for analyzing user emotions using facial recognition
[1214] OpenAI GPT-3: An AI model for natural language processing and narrative text generation
[1215] TextBlob: A Python library for simple natural language processing
[1216] Specific processing and data calculations
[1217] 1. Enter and submit the prompt
[1218] The user enters the prompt into the input field on the device and presses the send button, at which point the device captures the user's face with a camera and sends the facial recognition data to the emotion engine.
[1219] 2. Emotional Information Analysis
[1220] The facial recognition data sent by the device is analyzed by an emotion engine to identify the user's emotional state. For example, the DeepFace library can be used to detect emotions such as "curious" or "anxious" from the user's facial expressions.
[1221] 3. Prompt analysis and character behavior generation
[1222] The server analyzes the received prompts using OpenAI GPT-3 and generates specific character actions, such as "taking out a map and finding the way."
[1223] 4. Reflecting Emotional Information and Generating Story Text
[1224] Based on the analysis results of the emotion engine, the server reflects the user's emotional information in the character's behavior. Using OpenAI GPT-3, story text is generated based on the user's emotional state. For example, if the emotion "curious" is recognized, the element "enjoyed" is added to the story text.
[1225] 5. Display of character actions and story text
[1226] The generated character actions and story text are sent to the device and displayed on the device screen, allowing the user to see in real time how their prompts and emotions are reflected.
[1227] Examples of concrete examples and prompts
[1228] Specific examples
[1229] When a user types the prompt "Find a way into the forest" into a smartphone app, their face is captured by a camera and the emotion is recognized as "curious." As a result, a scene is generated in which the character happily unfolds a map, and the story text reads, "The character unfolds the map and begins to joyfully search for a way. After walking a little further, he excitedly finds a hidden entrance into the forest."
[1230] Examples of prompt statements
[1231] If the user inputs the prompt "Find a way to save the princess" and the emotion is recognized as "anxiety," the action "The character hurriedly searches around the castle and finds the hidden door" is generated and displayed along with the story text "The character searches around the castle in an anxious manner and discovers the hidden door."
[1232] As described above, this system dynamically changes character behavior and story text based on user prompts and real-time analyzed emotional information, enabling a highly personalized user experience.
[1233] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1234] Step 1:
[1235] The user types a prompt into the terminal.
[1236] The user enters a prompt into the input field of the device and presses the send button. For example, the user enters "Find a way into the forest." The device receives the prompt. The device also captures the user's face using the built-in camera and collects the image data.
[1237] Input: User prompt and facial image data
[1238] Output: prompt data and face image data
[1239] Step 2:
[1240] The terminal transmits the facial image data to the emotion engine.
[1241] The device sends the collected facial image data to the emotion engine, which uses the DeepFace library to analyze emotions from the facial image. This analysis process identifies the user's emotional state. For example, an emotion such as "curious" may be detected.
[1242] Input: Facial image data
[1243] Output: Emotional information
[1244] Step 3:
[1245] The server analyzes the prompts and generates character actions.
[1246] The device sends the received prompt data to the server, which uses OpenAI GPT-3 to analyze the prompt and understand its meaning. Based on the analysis results, the server determines the character's specific actions. For example, it might generate an action such as "take out a map and find the way."
[1247] Input: prompt data
[1248] Output: Character behavior
[1249] Step 4:
[1250] The server generates a story text that reflects the emotional information.
[1251] The server receives the emotion information from the emotion engine. The server then reflects this emotion information in the character's behavior and generates story text using OpenAI GPT-3. For example, if the user is recognized as "curious," the generated behavior is adjusted to "check the map with excitement," and the story text generated reads, "The character unfolds the map and begins to search for the way with excitement. After walking a little further, he excitedly finds a hidden entrance to the forest."
[1252] Input: Character behavior, emotional information
[1253] Output: Story text
[1254] Step 5:
[1255] The device displays character actions and story text.
[1256] The server sends the generated character actions and story text to the device, which receives this data and displays it on the screen, allowing the user to see in real time how their prompts and emotions are reflected.
[1257] Input: Character actions, story text
[1258] Output: On-screen display
[1259] Through the above processing steps, the system can generate and display dynamically changing character behavior and story text based on the user's prompts and emotional information.
[1260] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1261] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1262] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1263] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1264] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1265] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1266] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1267] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1268] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1269] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1270] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1271] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1272] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1273] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1274] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1275] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1276] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1277] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1278] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1279] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1280] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1281] The following is further disclosed regarding the above embodiment.
[1282] (Claim 1)
[1283] means for receiving a prompt from a user;
[1284] means for analyzing the received prompts and generating character actions;
[1285] a means for displaying the generated character behavior;
[1286] A system including:
[1287] (Claim 2)
[1288] 10. The system of claim 1, further comprising means for using natural language processing to parse the prompts.
[1289] (Claim 3)
[1290] 10. The system of claim 1, further comprising means for generating story text based on the generated character actions.
[1291] (Claim 4)
[1292] 10. The system of claim 1, further comprising means for varying the complexity of the generated story in response to the user's prompting skills.
[1293] (Claim 5)
[1294] 2. The system according to claim 1, further comprising a network connection means for transmitting and receiving data between the server and the user terminal.
[1295] (Claim 6)
[1296] 10. The system of claim 1, further comprising means for storing a user's prompt history and using it for analysis.
[1297] "Example 1"
[1298] (Claim 1)
[1299] means for receiving input data from a user;
[1300] means for analyzing the received input data and generating virtual character behavior;
[1301] means for displaying the generated virtual character behavior;
[1302] A system including:
[1303] (Claim 2)
[1304] 10. The system of claim 1, further comprising means for using natural language processing techniques for the analysis.
[1305] (Claim 3)
[1306] 10. The system of claim 1, further comprising means for generating narrative text based on the generated virtual character behavior.
[1307] "Application Example 1"
[1308] (Claim 1)
[1309] means for receiving a prompt from a user;
[1310] means for analyzing the received prompts and generating character actions and suggestions;
[1311] A means for displaying the generated character behavior and suggestions;
[1312] A means for making suggestions according to requests from customers of the virtual store;
[1313] A system including:
[1314] (Claim 2)
[1315] 10. The system of claim 1, further comprising means for using natural language processing to parse the prompts.
[1316] (Claim 3)
[1317] 10. The system of claim 1, further comprising means for generating story text based on the generated character actions and suggestions.
[1318] "Example 2: Combining Emotion Engines"
[1319] (Claim 1)
[1320] means for receiving text input from a user;
[1321] means for analyzing the received text input and generating an image character behavior;
[1322] a means for displaying the generated iconic character behavior;
[1323] A means for acquiring user emotion information;
[1324] A means for analyzing the acquired emotional information;
[1325] a means for adjusting the illustrated character's behavior and the narrative text using the analyzed emotional information;
[1326] A system including:
[1327] (Claim 2)
[1328] 10. The system of claim 1, further comprising means for using natural language processing techniques to analyze the received text input.
[1329] (Claim 3)
[1330] 10. The system of claim 1, further comprising means for generating a narrative based on the generated iconic character behavior.
[1331] "Application example 2 when combining emotion engines"
[1332] (Claim 1)
[1333] means for receiving a prompt from a user;
[1334] means for analyzing the received prompts and generating character actions;
[1335] a means for displaying the generated character behavior;
[1336] A means for acquiring user emotion information;
[1337] A means for analyzing the acquired emotional information;
[1338] A means for reflecting emotional information in character actions and story text;
[1339] A system including:
[1340] (Claim 2)
[1341] 10. The system of claim 1, further comprising means for using natural language processing to parse the prompts.
[1342] (Claim 3)
[1343] 10. The system of claim 1, further comprising means for generating story text based on the generated character actions. [Explanation of symbols]
[1344] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for receiving a prompt from a user; means for analyzing the received prompts and generating character actions; a means for displaying the generated character behavior; A system including:
2. 10. The system of claim 1, further comprising means for using natural language processing to parse the prompts.
3. The system of claim 1 further comprising means for generating story text based on the generated character actions.
4. 10. The system of claim 1, further comprising means for varying the complexity of the story generated in response to the user's prompting skill.
5. 2. The system according to claim 1, further comprising a network connection means for transmitting and receiving data between the server and the user terminal.
6. 10. The system of claim 1, further comprising means for storing a user's prompt history and using it for analysis.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A