System
The system addresses static storylines in games by integrating multimodal user inputs and real-time generative responses, offering a dynamic and immersive experience that simplifies game development.
Patent Information
- Application Number
- JP2024120541
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-25
- Publication Date
- 2026-02-05
AI Technical Summary
Conventional interactive games have static storylines, limiting immersion and placing a significant burden on game developers due to complex scenario design, and lack the ability to integrate and respond to multimodal user inputs effectively.
A system that acquires multimodal user input, standardizes it, and uses a generative model to dynamically generate a story and real-time responses, allowing users to interact with characters and environments based on their choices and emotions.
Provides a highly personalized and immersive gaming experience while reducing the complexity of scenario design, enhancing development flexibility and creativity.
Smart Images

Figure 2026019132000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional interactive games have a static storyline and offer limited options, resulting in a lack of immersion. Furthermore, complex scenario design places a significant burden on game developers. This invention aims to solve these issues and provide a highly personalized gaming experience by providing a system that dynamically generates a story in real time based on the user's choices and actions. [Means for solving the problem]
[0005] The present invention provides a system that includes a means for acquiring multimodal input from a user, integrating and standardizing the input, inputting the standardized data into a generative model, and generating a story in real time. It also provides a system that includes a means for generating character and environmental responses in real time based on the generated story and presenting them to the user. This system dynamically changes a traditional static storyline, allowing for natural responses based on the user's choices and actions.
[0006] "Multimodal input" refers to data in multiple formats, such as text, audio, images, and video.
[0007] "Standardization" refers to the process of converting data of different formats into a consistent form and making them compatible.
[0008] A "generative model" refers to an artificial intelligence model that generates output based on input data.
[0009] "Real-time" refers to a process that produces output immediately after an input is made, with little delay.
[0010] "Story" refers to a narrative that depicts a sequence of characters and events.
[0011] "Character" refers to any character or creature with which the user interacts within the game.
[0012] "Environment" refers to the physical and visual background or setting within which characters operate within a game.
[0013] "Response" refers to the system's reaction or reply to user input.
[0014] "Presenting to the user" refers to allowing the user to see, hear, or experience information or responses generated by the system. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0023] [First embodiment]
[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0036] This invention is an interactive storytelling system that generates stories in real time based on the user's multimodal input and presents them to the user. This system mainly consists of the following components: a user input processing module, a multimodal data integration module, a generative model module, a real-time response generation module, and an output display module.
[0037] User Input Processing Module
[0038] Users wear a VR headset and interact with the game using voice and gestures, which are then captured by the device through a microphone, camera, and gesture sensors. For example, a user can speak to a dragon by waving their hand or issuing a voice command.
[0039] Multimodal Data Integration Module
[0040] The voice data collected by the device is sent to a server, where it is converted to text using speech recognition technology. Gesture data is also sent to the server, where the user's movements are analyzed using a motion analysis algorithm. These multiple inputs are then combined to create a standardized data set.
[0041] Generative Model Module
[0042] The standardized dataset created above is input into a generative model, which is then run by a server to generate a storyline based on the user's choices and actions. This generative model can be, for example, GPT or another generative AI model, and generates a story in real time in response to the user's actions.
[0043] Real-time response generation module
[0044] Based on the generated storyline, the server generates responses for characters (e.g., dragons) and environments (e.g., the caves in which the dragons live), including the dragon's dialogue, actions, and changes to the surrounding environment. For example, the dragon may start speaking to answer the user's questions.
[0045] Output Display Module
[0046] The data generated by the server is sent to the device, which then displays it in the appropriate format for the user: audio data is played through the speakers, and text, images, and animations are displayed on the screen. For example, a scene in which a dragon moves and talks to the user in a realistic manner is projected on the user's VR headset.
[0047] Specific examples
[0048] When a user issues the voice command "talk to dragon", the following occurs:
[0049] 1. The user says "talk to dragon" aloud.
[0050] 2. The device captures the user's hand-waving gesture along with the audio.
[0051] 3. The device sends the captured data to the server.
[0052] 4. The server converts the speech to text and analyzes the gestures.
[0053] 5. The integrated dataset is fed into a generative model to generate a storyline.
[0054] 6. The server generates the dragon's response (e.g., "Hello, adventurer. What do you want?").
[0055] 7. The server sends the generated response to the terminal, which presents it to the user in real time.
[0056] This system allows users to have a highly personalized and immersive gaming experience, while also allowing game developers to omit complex scenario design, improving development flexibility and creativity.
[0057] The processing flow will be explained below.
[0058] Step 1:
[0059] The user puts on the VR headset and says, "Talk to the dragon." The device captures this voice through the microphone. At the same time, the device captures the user's hand gestures through the camera and sensors.
[0060] Step 2:
[0061] The device sends the captured voice data to the server, and also sends the captured gesture data to the server along with the voice data.
[0062] Step 3:
[0063] The server converts the received voice data into text using voice recognition technology. The voice recognition engine analyzes the voice waveform and generates the corresponding text data.
[0064] Step 4:
[0065] The server analyzes the gesture data and uses a motion analysis algorithm to identify the user's hand gestures and analyze their meaning. For example, it recognizes the "waving" gesture.
[0066] Step 5:
[0067] The server merges the speech and gesture data. The merger process combines the speech text and gesture data into a single standardized data set.
[0068] Step 6:
[0069] The server inputs the combined dataset into a generative model (e.g., GPT-3), which then runs the model and generates a storyline based on the user's input in real time.
[0070] Step 7:
[0071] The server analyzes the generated storyline and generates the character (dragon)'s reactions. The generation process includes specific actions such as the dragon nodding and talking.
[0072] Step 8:
[0073] The server sends generated data, including the dragon's reactions, to the device. This data includes text, audio, and animation data.
[0074] Step 9:
[0075] The device processes the received data and presents it to the user in real time, playing the dragon's voice through the speaker and showing the dragon's movements on the display.
[0076] Step 10:
[0077] The device presents the user with new options, such as "Ask the dragon" or "Fight the dragon."
[0078] Step 11:
[0079] The user selects their next action. Once a new selection is made, the device again sends this information to the server, and the process begins again at step 1.
[0080] Through this series of steps, the system dynamically generates a story based on user input, providing an interactive and immersive gaming experience.
[0081] Example 1
[0082] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0083] Conventional interactive entertainment systems have struggled to respond to user input in real time or generate complex storylines. Furthermore, they lacked technology to integrate and standardize multiple input modes (voice, gestures, etc.), making it difficult to provide an immersive experience that takes advantage of diverse user interactions. Therefore, there is a need for systems that allow users to enjoy more advanced interactive experiences and increase the flexibility and creativity of game developers.
[0084] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0085] In this invention, the server includes means for acquiring multimodal input from a user, means for integrating and standardizing the multimodal input, means for inputting the standardized data into a generative model and generating a story in real time, means for generating character and environmental responses in real time based on the generated story, means for presenting the responses to the user, means for the user to wear a head-mounted display and interact with the user through voice and gestures, means for converting voice data into text and analyzing gesture data, means for generating a storyline based on the user's actions using an artificial intelligence model as a generative model, and means for outputting the generated character's actions and environmental changes in real time, thereby enabling the user to have a highly personalized and immersive interactive experience.
[0086] "Multimodal input" refers to data input by users in multiple different modes, such as voice, gestures, images, and video.
[0087] A "head-mounted display" is a device worn by the user that displays visual information directly in front of the user's eyes.
[0088] A "generative model" refers to an artificial intelligence model that generates new content or stories based on user actions and input.
[0089] A "voice recognition module" is a piece of software that converts voice data into text data.
[0090] A "gesture sensor" is a device that detects a user's movements and captures their action data.
[0091] "Motion analysis algorithm" refers to an algorithm that analyzes captured gesture data to identify specific user movements.
[0092] An "integrated dataset" is a set of data that has been integrated and standardized from data input from multiple different modes.
[0093] "Real-time response generation" refers to the process of generating instant character and environmental reactions based on user input.
[0094] The "output display module" is a module for presenting data generated by the server to the user in an appropriate format.
[0095] The present invention relates to an interactive entertainment system that generates a story in real time based on a user's multimodal input and presents it to the user. The system consists of a user input processing module, a multimodal data integration module, a generative model module, a real-time response generation module, and an output display module.
[0096] User Input Processing Module
[0097] The user wears a head-mounted display and interacts with it using voice and gestures. The device captures these voices and gestures using a microphone, camera, and gesture sensor. For example, if a user says "talk to the dragon" and waves their hand, this data is captured.
[0098] Multimodal Data Integration Module
[0099] The voice data captured by the device is sent to a server where it is converted into text using speech recognition technology. Similarly, gesture data is sent to a server where a motion analysis algorithm analyzes the user's specific movements. This process standardizes the voice and gesture data and compiles them into a unified data set.
[0100] Generative Model Module
[0101] The server inputs the combined dataset into a generative model, such as an advanced generative AI model like GPT-3, which then generates a storyline in real time based on the user's actions.
[0102] Real-time response generation module
[0103] Based on the generated storyline, the server generates responses from characters and the environment in real time. For example, a scene is generated in which a dragon responds to a user's speech by saying, "Hello, adventurer. What do you want?" Changes to the dragon's movements and the surrounding environment (for example, the dragon moving forward or the lights in a cave turning on) are also generated in real time.
[0104] Output Display Module
[0105] The server sends the generated voice, text, image, and animation data to the device. The device then displays the data on the user's head-mounted display, and the audio is played through the speakers. For example, a scene in which a dragon moves and talks to you in a realistic way, or a scene in which lights in a cave light up, may be displayed on the user's head-mounted display.
[0106] Specific examples
[0107] When a user says "talk to dragon" and waves their hand, the following happens:
[0108] 1. The user says "talk to the dragon" and waves their right hand.
[0109] 2. The device captures these voices and gestures and sends them to the server.
[0110] 3. The server converts the voice data into text and analyzes the gesture data.
[0111] 4. The server inputs the integrated dataset into the generative model and generates a storyline.
[0112] 5. The server generates the dragon's responses based on the generated storyline and generates the dragon's movements in real time.
[0113] 6. The server sends the generated response and action to the terminal, which presents it to the user.
[0114] The system allows users to have a highly personalized, immersive, and interactive experience, and an example prompt using the generative AI model is "Generate a scene where the user talks to the dragon and asks what to do next."
[0115] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0116] Step 1: User Input Capture
[0117] The user wears a head-mounted display and interacts with it using voice and gestures. Specifically, the user says "talk to the dragon" and waves their right hand at the same time. The device captures voice with a microphone and gestures with a camera and gesture sensor. Voice data and gesture data are acquired as input. The output is the captured voice data and gesture data.
[0118] Step 2: Sending data
[0119] The device sends the captured voice and gesture data to the server. Specifically, the data is transferred to the server via the Internet in real time. The input is the captured voice and gesture data. The output is the data sent to the server.
[0120] Step 3: Speech and gesture analysis
[0121] The server converts the received voice data into text data using a voice recognition module. For example, the text "talk to the dragon" is acquired. At the same time, the server analyzes the gesture data using a motion analysis algorithm to identify the user's specific actions. The input is voice data and gesture data, and the output is text data and analyzed motion data.
[0122] Step 4: Data integration and standardization
[0123] The server integrates the text data from speech recognition and the gesture analysis data to create a standardized dataset. This is to unify data from different modes. Specifically, the server standardizes the data format and generates the integrated dataset. The input is text data and analyzed gesture data, and the output is the integrated dataset.
[0124] Step 5: Story Generation
[0125] The server inputs the standardized dataset into a generative model. GPT-3 or other models are used as generative models, generating a storyline based on user behavior. Specifically, the server runs the generative model and generates the next development in the story. The input is the integrated dataset, and the output is the generated storyline.
[0126] Step 6: Real-time response generation
[0127] The server determines the response of the character (for example, a dragon) based on the generated storyline. Environmental changes (for example, turning on the lights in a cave) are also generated. As specific actions, the server generates the dragon's lines and actions. For example, a response such as "Hello, adventurer. What do you want?" is generated. The input is the generated storyline, and the output is the character's response and environmental change data.
[0128] Step 7: Sending Responses and Actions
[0129] The server sends the generated response and action data to the terminal. Specifically, the server sends data to the terminal in real time via the Internet. The input is the character's response and environmental change data, and the output is the data sent to the terminal.
[0130] Step 8: Presenting the output
[0131] The device receives audio, text, images, and animation data from the server and presents it to the user. Specific operations include displaying a scene of a dragon speaking to the user on the head-mounted display and lighting up the cave, and playing audio through the speaker. The input is data sent from the server, and the output is the visuals and audio presented to the user.
[0132] (Application example 1)
[0133] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0134] In conventional storytelling systems, user interactions are fixed, making it difficult to generate stories in real time based on individual actions and utterances. Furthermore, the limited means of interaction mean that users lack a sense of immersion. Furthermore, the inability to integrate and utilize multi-modal data, such as gestures and voice, limits the user experience.
[0135] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0136] In this invention, the server includes: means for acquiring multimodal input from a user; means for integrating and standardizing the multimodal input; means for inputting the standardized data into a generative model to generate a story in real time; means for generating character and environmental responses in real time based on the generated story; means for presenting the responses to the user through a smartphone or tablet display; means for capturing and recognizing user gestures with a camera; means for combining the gestures and voice input to form story generation prompts; and means for playing the generated story aloud. This enables rich interactions that combine the user's voice and gestures, and makes it possible to provide a personalized story experience according to individual actions and utterances.
[0137] "Means for obtaining multimodal input from the user" refers to a function for capturing data in multiple forms, such as text, audio, images, and video, generated by the user.
[0138] "Means to integrate and standardize multimodal input" refers to a function that converts acquired data in multiple formats into a single unified dataset for easier processing.
[0139] "Means of inputting data into a generative model and generating a story in real time" refers to a function that allows standardized data to be input into a generative model (e.g., a generative AI model) to create a dynamic story on the spot.
[0140] "Means for generating character and environmental responses in real time" is a function that allows the environment, such as characters and backgrounds, to react in real time based on the generated story, creating appropriate actions and situations.
[0141] "Means of presenting to the user" refers to the function for showing or listening to the generated story or response to the user through a smartphone or tablet display, speaker, etc.
[0142] "Means for capturing and recognizing user gestures with a camera" refers to a function that uses a camera to capture the user's physical movements (e.g., hand gestures) and analyzes and understands them.
[0143] "Means for combining gestures and voice input to form story generation prompts" is a function for generating instructions (prompts) to be given to a generative AI model based on data that combines the user's gestures and voice input.
[0144] "Means for playing the generated story aloud" is a function for converting the story created by the generative model into audio and playing it through a speaker.
[0145] To implement this invention, it is necessary to acquire user voice and gesture input, integrate the data, and input it into a generative AI model to generate a story in real time. A specific embodiment of this is shown below.
[0146] Hardware and software used
[0147] 1. Hardware
[0148] Smartphone or tablet: Used to provide the user interface.
[0149] Camera: Used to capture user gestures.
[0150] Microphone: Used to collect the user's voice.
[0151] Speaker: Used to play the generated audio back to the user.
[0152] 2. Software
[0153] Google Cloud Speech-to-Text API: Used to convert user speech into text.
[0154] OpenCV: Used to recognize user gestures.
[0155] OpenAI GPT (Generative AI Model): Used to generate stories in real time based on user input.
[0156] gTTS (Google Text-to-Speech) and playsound: Used to play the audio of the generated story.
[0157] System Operation Overview
[0158] 1. Acquiring Multimodal Input
[0159] The user speaks into the smartphone or tablet and simultaneously gestures in front of the camera.
[0160] Smartphones capture audio through their microphones and gestures through their cameras.
[0161] 2. Data integration and standardization
[0162] The captured audio data is converted to text using the Google Cloud Speech-to-Text API.
[0163] Similarly, the captured gesture data is analyzed using OpenCV to create a standardized dataset.
[0164] 3. Story Generation
[0165] The combined dataset is fed into a generative AI model (OpenAI GPT) to generate a story in real time.
[0166] Generative AI models create appropriate character and environmental responses based on user actions and speech.
[0167] 4. Real-time response generation and presentation
[0168] The generated story is converted into audio data using gTTS and played back to the user through the smartphone's speakers.
[0169] At the same time, the display shows related images and text.
[0170] Specific examples
[0171] For example, if a user issues the voice command "talk to the wizard" and makes a hand waving gesture, the system provides the following prompt sentence to the generative AI model:
[0172] text
[0173] User Voice: Talk to the Wizard
[0174] User gesture: wave
[0175] Generate the following story:
[0176] Based on this prompt, the generative AI model generates a story like this:
[0177] text
[0178] "Hello, traveller. My name is Eldritch, how may I help you?" the wizard replied in a kind voice.
[0179] This generated story is then converted into audio by gTTS and played back to the user through the smartphone's speakers, with related images and text also appearing on the display, creating a more immersive experience for the user.
[0180] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0181] Step 1: The user speaks a command (e.g., "Talk to the wizard") into a smartphone or tablet. The smartphone captures this voice through its microphone and saves it as audio data. At the same time, the camera captures the user's hand-waving gesture and saves it as video data. This provides audio and video data as input.
[0182] Step 2: The device converts the captured voice data to text data using the Google Cloud Speech-to-Text API. When converting voice data to text data, the input is the voice data, and the output is the converted text data. For example, the utterance "Talk to the wizard" becomes the text "Talk to the wizard."
[0183] Step 3: The device analyzes the captured video data using OpenCV and recognizes the user's gesture (e.g., waving). The input is the video data, and the output is the recognized gesture information. For example, the user's hand wave is recognized as a "wave."
[0184] Step 4: The device combines the results of converting the voice data to text and the recognized gesture data to generate a standardized dataset. The input is text data and gesture data, and the output is the combined dataset. For example, the dataset might have the format "Voice: Talk to the Wizard" and "Gesture: Wave."
[0185] Step 5: The server receives the integrated dataset and inputs it into the generative AI model (OpenAI GPT). The input is the integrated dataset, and the output is a story text based on the generative AI model. For example, based on the prompts "User voice: Talk to the wizard" and "User gesture: Wave", the story "Hello, traveler. My name is Eldritch. How may I help you?" is generated.
[0186] Step 6: The server converts the generated story text into audio data using gTTS (Google Text-to-Speech). The input is the generated story text, and the output is the audio data of the story. This audio data is sent to the smartphone.
[0187] Step 7: The device plays the received audio data to the user through the speaker. The input is audio data, and the output is audio playback. At the same time, related text and images are displayed on the screen, providing the user with an immersive experience.
[0188] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0189] The present invention is an interactive storytelling system that generates a story in real time based on the user's multimodal input and presents it to the user, and further incorporates an emotion engine to recognize the user's emotions and adjust the story and character responses based on those emotions. The system of the present invention mainly consists of the following components: a user input processing module, an emotion engine, a multimodal data integration module, a generative model module, a real-time response generation module, and an output display module.
[0190] User Input Processing Module
[0191] Users wear a VR headset and interact with the game using voice and gestures. The device captures these interactions through a microphone, camera, and gesture sensors. For example, a user can speak to a dragon by issuing a voice command or waving their hand.
[0192] Emotion Engine
[0193] The server operates an emotion engine using voice data and image data acquired from the user. The emotion engine recognizes the user's emotions using voice analysis and detects and identifies the user's facial expressions using image analysis. This allows the engine to determine the user's current emotion (e.g., joy, anger, sadness) from the tone and facial expression of the user when speaking.
[0194] Multimodal Data Integration Module
[0195] The voice and gesture data collected by the device is sent to a server, where it is converted into text using speech recognition technology. Gesture data is also sent to the server, where the user's movements are analyzed using a motion analysis algorithm. These multiple input data are integrated to create a standardized dataset, which also includes emotional information recognized by an emotion engine.
[0196] Generative Model Module
[0197] The standardized dataset created above is input into a generative model. The server runs the generative model (e.g., GPT-3) to generate a storyline based on the user's input in real time. The generative model also uses the user's emotional information to adjust the story and generate a more convincing response.
[0198] Real-time response generation module
[0199] Based on the generated storyline, the server generates responses for characters (e.g., dragons) and environments (e.g., the caves where dragons live). The generation process includes specific actions such as the dragon nodding and talking. Reactions based on the user's emotions are also incorporated. For example, if the user is angry, a scenario will be generated in which the dragon speaks words to calm them down.
[0200] Output Display Module
[0201] The data generated by the server is sent to the device, which then displays it in an appropriate format for the user. Audio data is played through the speaker, and text, images, and animations are displayed on the screen. For example, a dragon may talk to the user while moving realistically, presenting new options to the user (e.g., "Ask the dragon" or "Fight the dragon").
[0202] Specific examples
[0203] Consider a scenario where a user issues the voice command "talk to the dragon" and simultaneously smiles and waves:
[0204] 1. The user says "talk to the dragon" and smiles and waves.
[0205] 2. The device captures voice, gesture, and facial expression data and sends it to the server.
[0206] 3. The server converts the speech into text and analyzes gestures and facial expressions.
[0207] 4. The emotion engine recognizes the user's smile and determines it to be a positive emotion (joy).
[0208] 5. The integrated dataset is fed into a generative model to generate a storyline.
[0209] 6. The server generates a dragon response based on the positive emotion (e.g., "You're in a good mood today, adventurer. How can I help you?").
[0210] 7. The server sends the generated response to the terminal, which presents it to the user in real time.
[0211] This system allows users to have a highly personalized, emotionally immersive gaming experience, while also allowing game developers to eliminate the need for complex scenario design, improving development flexibility and creativity.
[0212] The processing flow will be explained below.
[0213] Step 1:
[0214] The user puts on the VR headset and says "talk to the dragon." The device captures this voice through the microphone. At the same time, the device captures the user's hand gesture through the camera and sensors. The device also captures the user's facial expressions (e.g., smiling).
[0215] Step 2:
[0216] The terminal transmits the captured voice data, gesture data, and facial expression data to a server.
[0217] Step 3:
[0218] The server converts the received voice data into text using voice recognition technology. The voice recognition engine analyzes the voice waveform and generates the corresponding text data.
[0219] Step 4:
[0220] The server analyzes the gesture data and uses a motion analysis algorithm to identify the user's hand gestures and analyze their meaning. For example, it recognizes the "waving" gesture.
[0221] Step 5:
[0222] The server analyzes the facial expression data. It uses an expression analysis algorithm to recognize the user's facial expressions and identify their emotional state (e.g., joy, anger, sadness, or happiness). For example, it recognizes the emotion of "joy" from the user's smile.
[0223] Step 6:
[0224] The server integrates the speech, gesture, and facial expression data. The integration process brings together speech text, gesture data, and emotion information into a single standardized dataset.
[0225] Step 7:
[0226] The server inputs the combined dataset into a generative model (e.g., GPT-3), which then runs the model and generates a storyline based on the user's input in real time.
[0227] Step 8:
[0228] The server analyzes the generated storyline and generates a character (e.g., a dragon)'s reaction based on the user's emotional information. For example, if the user has the emotion "joy," the server generates a scenario in which the dragon responds in a friendly manner.
[0229] Step 9:
[0230] The server sends generated data, including character reactions, to the device. This data includes text, audio, and animation data.
[0231] Step 10:
[0232] The device processes the received data and presents it to the user in real time, playing the dragon's voice through the speaker and showing the dragon's movements on the display.
[0233] Step 11:
[0234] The device presents the user with new options, such as "Ask the dragon" or "Fight the dragon."
[0235] Step 12:
[0236] The user selects their next action. Once a new selection is made, the device again sends this information to the server, and the process begins again at step 1.
[0237] This series of steps enables the system to dynamically generate a story while taking into account the user's emotions, providing an interactive and immersive gaming experience.
[0238] Example 2
[0239] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0240] Previous interactive storytelling systems lacked the ability to recognize users' emotions in real time and adjust the story and character responses based on those emotions. As a result, users often did not receive responses that reflected their own emotions and actions, resulting in a less immersive experience. Furthermore, there was a lack of technical means to provide flexible and advanced interactions.
[0241] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0242] In this invention, the server includes means for acquiring multimodal input from a user, means for integrating and standardizing the multimodal input, means for inputting the standardized data into a generative model to generate a story in real time, means for recognizing the user's emotions, means for reflecting the recognized emotional information in a real-time response, and means for presenting the response to the user. This enables real-time responses that are in line with the user's emotions and behavior, providing a more immersive interactive experience.
[0243] "Multimodal input" refers to input from a user that includes data in multiple formats, such as text, audio, images, and video.
[0244] "Means of integration and standardization" refers to the process of converting data in multiple different formats into a consistent format and compiling it into a single data set.
[0245] A "generative model" refers to an artificial intelligence model that generates new text or storylines based on given input data.
[0246] A "means for generating stories in real time" is a means for instantly responding to input from the user and generating new stories on the spot.
[0247] "Means for generating character and environmental responses in real time" refers to means for instantly generating character actions and environmental changes based on the generated story.
[0248] "Means for recognizing user emotions" refers to means for identifying a user's emotional state in real time using voice analysis or image analysis.
[0249] "Means for reflecting emotional information in real-time responses" refers to means for adjusting the responses and stories generated based on the recognized emotions of the user.
[0250] "Means for presenting responses to users" refers to the means by which the generated responses or stories are conveyed to users in the form of audio, text, animation, etc.
[0251] This invention is an interactive storytelling system that generates a story in real time based on the user's multimodal input and presents it to the user. It also incorporates an emotion engine, which recognizes the user's emotions and adjusts the story and character responses based on those emotions. This system mainly consists of the following components: a user input processing module, an emotion engine, a multimodal data integration module, a generative model module, a real-time response generation module, and an output display module.
[0252] User Input Processing Module
[0253] Users wear a VR headset and interact with the game using voice and gestures. The device captures these interactions through a microphone, camera, and gesture sensor. For example, a user can issue a voice command to "talk to the dragon" and wave their hand.
[0254] Emotion Engine
[0255] The server runs an emotion engine using voice and image data acquired from the user. The emotion engine recognizes the user's emotions using voice analysis and detects and identifies the user's facial expressions using image analysis. This allows the engine to determine the user's current emotion (e.g., joy, anger, sadness) from the tone and facial expression of the user when speaking.
[0256] Multimodal Data Integration Module
[0257] The device sends collected voice and gesture data to a server, which converts the data into text using speech recognition technology. Gesture data is also sent to the server, where the user's movements are analyzed using a motion analysis algorithm. These multiple inputs are combined to create a standardized dataset, which also includes emotional information recognized by an emotion engine.
[0258] Generative Model Module
[0259] The standardized dataset is then fed into a generative model. The server runs the generative model (e.g., a large-scale language model) to generate a storyline based on the user's input in real time. The generative model also uses the user's emotional information to adjust the story and generate a more convincing response.
[0260] Real-time response generation module
[0261] Based on the generated storyline, the server generates responses for characters (e.g., dragons) and environments (e.g., the caves where the dragons live). The generation process includes specific actions such as the dragon nodding and talking. Reactions based on the user's emotions can also be incorporated. For example, if the user is angry, a scenario will be generated in which the dragon speaks words to calm them down.
[0262] Output Display Module
[0263] The data generated by the server is sent to the device, which then displays it in an appropriate format for the user. Audio data is played through the speaker, and text, images, and animations are displayed on the screen. For example, a dragon moves realistically and speaks to the user, presenting new options to the user.
[0264] Specific examples
[0265] Consider a scenario where the user gives the voice command "talk to the dragon" and smiles and waves:
[0266] 1. The user says "talk to the dragon" and smiles and waves.
[0267] 2. The device captures voice, gesture, and facial expression data and sends it to the server.
[0268] 3. The server converts the speech into text and analyzes gestures and facial expressions.
[0269] 4. The emotion engine recognizes the user's smile and determines it to be a positive emotion (joy).
[0270] 5. The integrated dataset is fed into a generative model to generate a storyline.
[0271] 6. The server generates a dragon response based on the positive emotion (e.g., "You're in a good mood today, adventurer. How can I help you?").
[0272] 7. The server sends the generated response to the terminal, which presents it to the user in real time.
[0273] Prompt Sentence Examples
[0274] "The user smiles and says the voice command 'talk to the dragon.' Generate a response from the dragon."
[0275] "If the user is angry, write a line for the dragon that will calm them down."
[0276] This system allows users to have a highly personalized, emotionally immersive experience, while also allowing game developers to eliminate the need for complex scenario design, improving development flexibility and creativity.
[0277] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0278] Step 1:
[0279] The user puts on the VR headset and begins interacting with it using voice and gestures. For example, they might say, "Talk to the dragon," and then wave their hand. The device's microphone, camera, and gesture sensors capture this voice, image, and movement data.
[0280] Input: Audio, video, gesture data
[0281] Output: Raw captured data
[0282] Step 2:
[0283] The voice data and gesture data captured by the device are sent to the server via the network, along with the voice data captured by the microphone and the gesture data captured by the camera and gesture sensor.
[0284] Input: Raw captured data
[0285] Output: Voice and gesture data sent to the server
[0286] Step 3:
[0287] The server uses voice recognition technology (e.g., a voice-to-text conversion system) to convert the captured voice data into text, and uses a motion analysis algorithm (e.g., a motion detection algorithm) to analyze the gesture data and recognize the user's movements.
[0288] Input: Voice data, gesture data
[0289] Data processing: Converts voice data into text, and gesture data into motion recognition information
[0290] Output: Text data, motion recognition data
[0291] Step 4:
[0292] The server runs an emotion engine (e.g., a voice emotion analysis system, an expression analysis system) to analyze the voice tone and facial expressions, thereby determining whether the user is expressing an emotion such as joy or anger.
[0293] Input: Audio data, video data
[0294] Data processing: voice tone analysis, facial expression analysis
[0295] Output: Emotion recognition data
[0296] Step 5:
[0297] The server combines the text data, action recognition data, and emotion recognition data to create a standardized dataset, which is then input into the subsequent generative model.
[0298] Input: Text data, action recognition data, emotion recognition data
[0299] Data processing: data integration and standardization
[0300] Output: Standardized dataset
[0301] Step 6:
[0302] The server runs a generative AI model (e.g., a large-scale language model) to generate storylines in real time based on the integrated dataset, for example, generating positive stories based on what makes the user happy.
[0303] Input: Standardized dataset
[0304] Data processing: Story generation using generative AI models
[0305] Output: Generated storyline
[0306] Step 7:
[0307] The server generates character actions and conversations based on the generated storyline. For example, it generates lines and actions for a dragon to speak to the user. The response changes depending on the user's emotions.
[0308] Input: Storyline, emotion recognition data
[0309] Data processing: Character response generation
[0310] Output: Character movement data, character conversation data
[0311] Step 8:
[0312] The server sends the generated character movement data and conversation data to the device, which then presents this to the user in real time. Audio is played through the speaker and animation is shown on the display. For example, a dragon could move realistically and say, "You're in a good mood today, adventurer. Is there anything I can help you with?"
[0313] Input: Character movement data, character conversation data
[0314] Output: Real-time display to the user
[0315] (Application example 2)
[0316] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0317] Conventional interactive storytelling systems do not take into account the user's emotional state, resulting in a lack of individualized optimization of the user's experience and difficulty in achieving emotional empathy. Furthermore, it is difficult to effectively integrate the user's diverse inputs (voice, facial expressions, gestures), making it difficult to generate appropriate responses in real time. Furthermore, it has been challenging to realize a system that is not dependent on a specific device and uses general-purpose hardware.
[0318] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0319] In this invention, the server includes: means for acquiring multimodal input from a user; means for integrating and standardizing the multimodal input; means for inputting the standardized data into a generative model to generate a story in real time; means for generating responses of characters and the environment in real time based on the generated story; means for presenting the responses to the user; means for analyzing the user's emotions and adjusting the generated story and character responses based on the emotions; means for acquiring the user's voice, facial expressions, and gestures using a smartphone camera and microphone; and means for displaying the story generated in real time on a display device according to the user's emotional state, thereby enabling interactive storytelling that is individually optimized according to the user's emotional state.
[0320] "Means for acquiring multimodal input from a user" refers to a device or system for simultaneously acquiring multiple forms of input, such as a user's voice, facial expressions, and gestures.
[0321] The "means for integrating and standardizing the multimodal input" refers to a device or system that processes multiple acquired input formats in a unified manner and converts them into a format that is easy to analyze.
[0322] "Means for inputting the standardized data into a generative model and generating a story in real time" refers to a device or system that generates a story in real time using an artificial intelligence model based on the integrated and standardized data.
[0323] "Means for generating character and environmental responses in real time based on the generated story" refers to a device or system that generates character and environmental responses in real time according to a story created by a generative model.
[0324] "Means for presenting the response to the user" refers to a device or system that displays or outputs the response of the generated character or environment to the user.
[0325] "Means for analyzing the user's emotions and adjusting the generated story and character responses based on said emotions" refers to a device or system that analyzes the user's emotional state and adjusts the generated story and character responses based on that data.
[0326] "Means for capturing a user's voice, facial expressions, and gestures using a smartphone's camera and microphone" refers to a device or system that captures a user's voice, facial expressions, and gestures using a smartphone's built-in camera and microphone.
[0327] "Means for displaying a story generated in real time on a display device according to the user's emotional state" refers to a device or system for appropriately displaying to a user a story generated in real time based on the user's emotional state.
[0328] This invention uses a system that analyzes various user inputs (voice, facial expressions, gestures) in real time and generates an interactive story based on the analysis. This system mainly uses a smartphone and its built-in camera and microphone.
[0329] Hardware and software used
[0330] 1. Smartphone: Uses a camera and microphone to capture the user's voice, facial expressions, and gestures.
[0331] 2. Server: Data processing, sentiment analysis, story generation, character response generation. The main software used here is as follows:
[0332] OpenCV: A library used for camera capture and image analysis.
[0333] speech_recognition: The library to use for speech recognition.
[0334] transformers (GPT-2): A library used as a generative model.
[0335] TextBlob: A library used for sentiment analysis.
[0336] Processing flow
[0337] 1. User Input Processing:
[0338] The user speaks and gestures into the smartphone's camera, which then captures the audio and image data.
[0339] 2. Multimodal Data Integration:
[0340] The smartphone captures audio and image data and sends it to the server. The audio data is converted to text using the speech_recognition library. The image data is analyzed for gestures and facial expressions using OpenCV.
[0341] 3. Emotion analysis:
[0342] The converted voice data and image analysis results are sent to the server, where the TextBlob library analyzes the user's emotions, classifying them as positive, negative, or neutral.
[0343] 4. Story generation:
[0344] The data combined with the sentiment analysis results is fed into a GPT-2 model from the transformers library to generate an interactive story based on the user's emotional state.
[0345] 5. Character response generation:
[0346] Based on the story created by the generative model, specific responses from the character are generated. For example, when a dragon speaks to the user, it responds appropriately based on the user's emotions.
[0347] 6. Real-time display:
[0348] The generated story and character responses are displayed on the smartphone screen and played back aloud.
[0349] Specific examples
[0350] Consider a scenario where the user says "talk to the dragon" and smiles and waves:
[0351] 1. The smartphone's microphone and camera capture the user's voice, facial expressions, and gestures.
[0352] 2. The captured data is sent to a server, where the voice data is converted into text and the image data is analyzed for gestures and facial expressions.
[0353] 3. The TextBlob library parses the user's smile as a positive emotion.
[0354] 4. The combined data is fed into the GPT-2 model to generate prompts for the user to talk to the dragon in a positive manner.
[0355] 5. An interactive story is generated based on the generated prompts.
[0356] 6. As part of the story, the dragon responds to the user by saying, "You're in a good mood today, adventurer. How can I help you?"
[0357] Prompt Sentence Examples
[0358] An example prompt that describes a situation in which the user says "talk to the dragon" and smiles and waves is:
[0359] The user says "talk to the dragon," smiles, and waves. The user is pleased. How does the dragon respond?
[0360] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0361] Step 1:
[0362] The user speaks and gestures into the smartphone camera. The smartphone camera and microphone capture audio and image data. In this step, the user's voice input and gesture input are collected. The input data are audio files and image files.
[0363] Step 2:
[0364] The device sends captured voice and image data to the server. The voice data is converted to text using the speech_recognition library. The image data is analyzed for gestures and facial expressions using OpenCV. The input is the captured voice data and image data, and the output is text data and analyzed gesture and facial expression data.
[0365] Step 3:
[0366] The server inputs the converted voice data and image analysis results into the TextBlob library to analyze the user's emotions. The emotion analysis results are classified as positive, negative, or neutral. The input is text data and image analysis data, and the output is the user's emotional state.
[0367] Step 4:
[0368] The integrated data (text data, gesture analysis data, and emotion analysis data) is input to the GPT-2 model in the transformers library on the server to generate an interactive story based on the user's emotional state. The input is the integrated data, and the output is the generated story.
[0369] Step 5:
[0370] The server generates specific responses for the characters based on the generated story. For example, when a dragon speaks to a user, it responds appropriately based on the user's emotions. The input is the generated story, and the output is the character's response.
[0371] Step 6:
[0372] The server sends the generated story and character responses to the device, which displays them on the smartphone screen and plays them audibly. The input is the character responses and the generated story, and the output is an interactive story experience that is displayed and played back to the user.
[0373] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0374] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0375] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0376] [Second embodiment]
[0377] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0378] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0379] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0380] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0381] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0382] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0383] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0384] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0385] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0386] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0387] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0388] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0389] This invention is an interactive storytelling system that generates stories in real time based on the user's multimodal input and presents them to the user. This system mainly consists of the following components: a user input processing module, a multimodal data integration module, a generative model module, a real-time response generation module, and an output display module.
[0390] User Input Processing Module
[0391] Users wear a VR headset and interact with the game using voice and gestures, which are then captured by the device through a microphone, camera, and gesture sensors. For example, a user can speak to a dragon by waving their hand or issuing a voice command.
[0392] Multimodal Data Integration Module
[0393] The voice data collected by the device is sent to a server, where it is converted to text using speech recognition technology. Gesture data is also sent to the server, where the user's movements are analyzed using a motion analysis algorithm. These multiple inputs are then combined to create a standardized data set.
[0394] Generative Model Module
[0395] The standardized dataset created above is input into a generative model, which is then run by a server to generate a storyline based on the user's choices and actions. This generative model can be, for example, GPT or another generative AI model, and generates a story in real time in response to the user's actions.
[0396] Real-time response generation module
[0397] Based on the generated storyline, the server generates responses for characters (e.g., dragons) and environments (e.g., the caves in which the dragons live), including the dragon's dialogue, actions, and changes to the surrounding environment. For example, the dragon may start speaking to answer the user's questions.
[0398] Output Display Module
[0399] The data generated by the server is sent to the device, which then displays it in the appropriate format for the user: audio data is played through the speakers, and text, images, and animations are displayed on the screen. For example, a scene in which a dragon moves and talks to the user in a realistic manner is projected on the user's VR headset.
[0400] Specific examples
[0401] When a user issues the voice command "talk to dragon", the following occurs:
[0402] 1. The user says "talk to dragon" aloud.
[0403] 2. The device captures the user's hand-waving gesture along with the audio.
[0404] 3. The device sends the captured data to the server.
[0405] 4. The server converts the speech to text and analyzes the gestures.
[0406] 5. The integrated dataset is fed into a generative model to generate a storyline.
[0407] 6. The server generates the dragon's response (e.g., "Hello, adventurer. What do you want?").
[0408] 7. The server sends the generated response to the terminal, which presents it to the user in real time.
[0409] This system allows users to have a highly personalized and immersive gaming experience, while also allowing game developers to omit complex scenario design, improving development flexibility and creativity.
[0410] The processing flow will be explained below.
[0411] Step 1:
[0412] The user puts on the VR headset and says, "Talk to the dragon." The device captures this voice through the microphone. At the same time, the device captures the user's hand gestures through the camera and sensors.
[0413] Step 2:
[0414] The device sends the captured voice data to the server, and also sends the captured gesture data to the server along with the voice data.
[0415] Step 3:
[0416] The server converts the received voice data into text using voice recognition technology. The voice recognition engine analyzes the voice waveform and generates the corresponding text data.
[0417] Step 4:
[0418] The server analyzes the gesture data and uses a motion analysis algorithm to identify the user's hand gestures and analyze their meaning. For example, it recognizes the "waving" gesture.
[0419] Step 5:
[0420] The server merges the speech and gesture data. The merger process combines the speech text and gesture data into a single standardized data set.
[0421] Step 6:
[0422] The server inputs the combined dataset into a generative model (e.g., GPT-3), which then runs the model and generates a storyline based on the user's input in real time.
[0423] Step 7:
[0424] The server analyzes the generated storyline and generates the character (dragon)'s reactions. The generation process includes specific actions such as the dragon nodding and talking.
[0425] Step 8:
[0426] The server sends generated data, including the dragon's reactions, to the device. This data includes text, audio, and animation data.
[0427] Step 9:
[0428] The device processes the received data and presents it to the user in real time, playing the dragon's voice through the speaker and showing the dragon's movements on the display.
[0429] Step 10:
[0430] The device presents the user with new options, such as "Ask the dragon" or "Fight the dragon."
[0431] Step 11:
[0432] The user selects their next action. Once a new selection is made, the device again sends this information to the server, and the process begins again at step 1.
[0433] Through this series of steps, the system dynamically generates a story based on user input, providing an interactive and immersive gaming experience.
[0434] Example 1
[0435] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0436] Conventional interactive entertainment systems have struggled to respond to user input in real time or generate complex storylines. Furthermore, they lacked technology to integrate and standardize multiple input modes (voice, gestures, etc.), making it difficult to provide an immersive experience that takes advantage of diverse user interactions. Therefore, there is a need for systems that allow users to enjoy more advanced interactive experiences and increase the flexibility and creativity of game developers.
[0437] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0438] In this invention, the server includes means for acquiring multimodal input from a user, means for integrating and standardizing the multimodal input, means for inputting the standardized data into a generative model and generating a story in real time, means for generating character and environmental responses in real time based on the generated story, means for presenting the responses to the user, means for the user to wear a head-mounted display and interact with the user through voice and gestures, means for converting voice data into text and analyzing gesture data, means for generating a storyline based on the user's actions using an artificial intelligence model as a generative model, and means for outputting the generated character's actions and environmental changes in real time, thereby enabling the user to have a highly personalized and immersive interactive experience.
[0439] "Multimodal input" refers to data input by users in multiple different modes, such as voice, gestures, images, and video.
[0440] A "head-mounted display" is a device worn by the user that displays visual information directly in front of the user's eyes.
[0441] A "generative model" refers to an artificial intelligence model that generates new content or stories based on user actions and input.
[0442] A "voice recognition module" is a piece of software that converts voice data into text data.
[0443] A "gesture sensor" is a device that detects a user's movements and captures their action data.
[0444] "Motion analysis algorithm" refers to an algorithm that analyzes captured gesture data to identify specific user movements.
[0445] An "integrated dataset" is a set of data that has been integrated and standardized from data input from multiple different modes.
[0446] "Real-time response generation" refers to the process of generating instant character and environmental reactions based on user input.
[0447] The "output display module" is a module for presenting data generated by the server to the user in an appropriate format.
[0448] The present invention relates to an interactive entertainment system that generates a story in real time based on a user's multimodal input and presents it to the user. The system consists of a user input processing module, a multimodal data integration module, a generative model module, a real-time response generation module, and an output display module.
[0449] User Input Processing Module
[0450] The user wears a head-mounted display and interacts with it using voice and gestures. The device captures these voices and gestures using a microphone, camera, and gesture sensor. For example, if a user says "talk to the dragon" and waves their hand, this data is captured.
[0451] Multimodal Data Integration Module
[0452] The voice data captured by the device is sent to a server where it is converted into text using speech recognition technology. Similarly, gesture data is sent to a server where a motion analysis algorithm analyzes the user's specific movements. This process standardizes the voice and gesture data and compiles them into a unified data set.
[0453] Generative Model Module
[0454] The server inputs the combined dataset into a generative model, such as an advanced generative AI model like GPT-3, which then generates a storyline in real time based on the user's actions.
[0455] Real-time response generation module
[0456] Based on the generated storyline, the server generates responses from characters and the environment in real time. For example, a scene is generated in which a dragon responds to a user's speech by saying, "Hello, adventurer. What do you want?" Changes to the dragon's movements and the surrounding environment (for example, the dragon moving forward or the lights in a cave turning on) are also generated in real time.
[0457] Output Display Module
[0458] The server sends the generated voice, text, image, and animation data to the device. The device then displays the data on the user's head-mounted display, and the audio is played through the speakers. For example, a scene in which a dragon moves and talks to you in a realistic way, or a scene in which lights in a cave light up, may be displayed on the user's head-mounted display.
[0459] Specific examples
[0460] When a user says "talk to dragon" and waves their hand, the following happens:
[0461] 1. The user says "talk to the dragon" and waves their right hand.
[0462] 2. The device captures these voices and gestures and sends them to the server.
[0463] 3. The server converts the voice data into text and analyzes the gesture data.
[0464] 4. The server inputs the integrated dataset into the generative model and generates a storyline.
[0465] 5. The server generates the dragon's responses based on the generated storyline and generates the dragon's movements in real time.
[0466] 6. The server sends the generated response and action to the terminal, which presents it to the user.
[0467] The system allows users to have a highly personalized, immersive, and interactive experience, and an example prompt using the generative AI model is "Generate a scene where the user talks to the dragon and asks what to do next."
[0468] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0469] Step 1: User Input Capture
[0470] The user wears a head-mounted display and interacts with it using voice and gestures. Specifically, the user says "talk to the dragon" and waves their right hand at the same time. The device captures voice with a microphone and gestures with a camera and gesture sensor. Voice data and gesture data are acquired as input. The output is the captured voice data and gesture data.
[0471] Step 2: Sending data
[0472] The device sends the captured voice and gesture data to the server. Specifically, the data is transferred to the server via the Internet in real time. The input is the captured voice and gesture data. The output is the data sent to the server.
[0473] Step 3: Speech and gesture analysis
[0474] The server converts the received voice data into text data using a voice recognition module. For example, the text "talk to the dragon" is acquired. At the same time, the server analyzes the gesture data using a motion analysis algorithm to identify the user's specific actions. The input is voice data and gesture data, and the output is text data and analyzed motion data.
[0475] Step 4: Data integration and standardization
[0476] The server integrates the text data from speech recognition and the gesture analysis data to create a standardized dataset. This is to unify data from different modes. Specifically, the server standardizes the data format and generates the integrated dataset. The input is text data and analyzed gesture data, and the output is the integrated dataset.
[0477] Step 5: Story Generation
[0478] The server inputs the standardized dataset into a generative model. GPT-3 or other models are used as generative models, generating a storyline based on user behavior. Specifically, the server runs the generative model and generates the next development in the story. The input is the integrated dataset, and the output is the generated storyline.
[0479] Step 6: Real-time response generation
[0480] The server determines the response of the character (for example, a dragon) based on the generated storyline. Environmental changes (for example, turning on the lights in a cave) are also generated. As specific actions, the server generates the dragon's lines and actions. For example, a response such as "Hello, adventurer. What do you want?" is generated. The input is the generated storyline, and the output is the character's response and environmental change data.
[0481] Step 7: Sending Responses and Actions
[0482] The server sends the generated response and action data to the terminal. Specifically, the server sends data to the terminal in real time via the Internet. The input is the character's response and environmental change data, and the output is the data sent to the terminal.
[0483] Step 8: Presenting the output
[0484] The device receives audio, text, images, and animation data from the server and presents it to the user. Specific operations include displaying a scene of a dragon speaking to the user on the head-mounted display and lighting up the cave, and playing audio through the speaker. The input is data sent from the server, and the output is the visuals and audio presented to the user.
[0485] (Application example 1)
[0486] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0487] In conventional storytelling systems, user interactions are fixed, making it difficult to generate stories in real time based on individual actions and utterances. Furthermore, the limited means of interaction mean that users lack a sense of immersion. Furthermore, the inability to integrate and utilize multi-modal data, such as gestures and voice, limits the user experience.
[0488] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0489] In this invention, the server includes: means for acquiring multimodal input from a user; means for integrating and standardizing the multimodal input; means for inputting the standardized data into a generative model to generate a story in real time; means for generating character and environmental responses in real time based on the generated story; means for presenting the responses to the user through a smartphone or tablet display; means for capturing and recognizing user gestures with a camera; means for combining the gestures and voice input to form story generation prompts; and means for playing the generated story aloud. This enables rich interactions that combine the user's voice and gestures, and makes it possible to provide a personalized story experience according to individual actions and utterances.
[0490] "Means for obtaining multimodal input from the user" refers to a function for capturing data in multiple forms, such as text, audio, images, and video, generated by the user.
[0491] "Means to integrate and standardize multimodal input" refers to a function that converts acquired data in multiple formats into a single unified dataset for easier processing.
[0492] "Means of inputting data into a generative model and generating a story in real time" refers to a function that allows standardized data to be input into a generative model (e.g., a generative AI model) to create a dynamic story on the spot.
[0493] "Means for generating character and environmental responses in real time" is a function that allows the environment, such as characters and backgrounds, to react in real time based on the generated story, creating appropriate actions and situations.
[0494] "Means of presenting to the user" refers to the function for showing or listening to the generated story or response to the user through a smartphone or tablet display, speaker, etc.
[0495] "Means for capturing and recognizing user gestures with a camera" refers to a function that uses a camera to capture the user's physical movements (e.g., hand gestures) and analyzes and understands them.
[0496] "Means for combining gestures and voice input to form story generation prompts" is a function for generating instructions (prompts) to be given to a generative AI model based on data that combines the user's gestures and voice input.
[0497] "Means for playing the generated story aloud" is a function for converting the story created by the generative model into audio and playing it through a speaker.
[0498] To implement this invention, it is necessary to acquire user voice and gesture input, integrate the data, and input it into a generative AI model to generate a story in real time. A specific embodiment of this is shown below.
[0499] Hardware and software used
[0500] 1. Hardware
[0501] Smartphone or tablet: Used to provide the user interface.
[0502] Camera: Used to capture user gestures.
[0503] Microphone: Used to collect the user's voice.
[0504] Speaker: Used to play the generated audio back to the user.
[0505] 2. Software
[0506] Google Cloud Speech-to-Text API: Used to convert user speech into text.
[0507] OpenCV: Used to recognize user gestures.
[0508] OpenAI GPT (Generative AI Model): Used to generate stories in real time based on user input.
[0509] gTTS (Google Text-to-Speech) and playsound: Used to play the audio of the generated story.
[0510] System Operation Overview
[0511] 1. Acquiring Multimodal Input
[0512] The user speaks into the smartphone or tablet and simultaneously gestures in front of the camera.
[0513] Smartphones capture audio through their microphones and gestures through their cameras.
[0514] 2. Data integration and standardization
[0515] The captured audio data is converted to text using the Google Cloud Speech-to-Text API.
[0516] Similarly, the captured gesture data is analyzed using OpenCV to create a standardized dataset.
[0517] 3. Story Generation
[0518] The combined dataset is fed into a generative AI model (OpenAI GPT) to generate a story in real time.
[0519] Generative AI models create appropriate character and environmental responses based on user actions and speech.
[0520] 4. Real-time response generation and presentation
[0521] The generated story is converted into audio data using gTTS and played back to the user through the smartphone's speakers.
[0522] At the same time, the display shows related images and text.
[0523] Specific examples
[0524] For example, if a user issues the voice command "talk to the wizard" and makes a hand waving gesture, the system provides the following prompt sentence to the generative AI model:
[0525] text
[0526] User Voice: Talk to the Wizard
[0527] User gesture: wave
[0528] Generate the following story:
[0529] Based on this prompt, the generative AI model generates a story like this:
[0530] text
[0531] "Hello, traveller. My name is Eldritch, how may I help you?" the wizard replied in a kind voice.
[0532] This generated story is then converted into audio by gTTS and played back to the user through the smartphone's speakers, with related images and text also appearing on the display, creating a more immersive experience for the user.
[0533] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0534] Step 1: The user speaks a command (e.g., "Talk to the wizard") into a smartphone or tablet. The smartphone captures this voice through its microphone and saves it as audio data. At the same time, the camera captures the user's hand-waving gesture and saves it as video data. This provides audio and video data as input.
[0535] Step 2: The device converts the captured voice data to text data using the Google Cloud Speech-to-Text API. When converting voice data to text data, the input is the voice data, and the output is the converted text data. For example, the utterance "Talk to the wizard" becomes the text "Talk to the wizard."
[0536] Step 3: The device analyzes the captured video data using OpenCV and recognizes the user's gesture (e.g., waving). The input is the video data, and the output is the recognized gesture information. For example, the user's hand wave is recognized as a "wave."
[0537] Step 4: The device combines the results of converting the voice data to text and the recognized gesture data to generate a standardized dataset. The input is text data and gesture data, and the output is the combined dataset. For example, the dataset might have the format "Voice: Talk to the Wizard" and "Gesture: Wave."
[0538] Step 5: The server receives the integrated dataset and inputs it into the generative AI model (OpenAI GPT). The input is the integrated dataset, and the output is a story text based on the generative AI model. For example, based on the prompts "User voice: Talk to the wizard" and "User gesture: Wave", the story "Hello, traveler. My name is Eldritch. How may I help you?" is generated.
[0539] Step 6: The server converts the generated story text into audio data using gTTS (Google Text-to-Speech). The input is the generated story text, and the output is the audio data of the story. This audio data is sent to the smartphone.
[0540] Step 7: The device plays the received audio data to the user through the speaker. The input is audio data, and the output is audio playback. At the same time, related text and images are displayed on the screen, providing the user with an immersive experience.
[0541] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0542] The present invention is an interactive storytelling system that generates a story in real time based on the user's multimodal input and presents it to the user, and further incorporates an emotion engine to recognize the user's emotions and adjust the story and character responses based on those emotions. The system of the present invention mainly consists of the following components: a user input processing module, an emotion engine, a multimodal data integration module, a generative model module, a real-time response generation module, and an output display module.
[0543] User Input Processing Module
[0544] Users wear a VR headset and interact with the game using voice and gestures. The device captures these interactions through a microphone, camera, and gesture sensors. For example, a user can speak to a dragon by issuing a voice command or waving their hand.
[0545] Emotion Engine
[0546] The server operates an emotion engine using voice data and image data acquired from the user. The emotion engine recognizes the user's emotions using voice analysis and detects and identifies the user's facial expressions using image analysis. This allows the engine to determine the user's current emotion (e.g., joy, anger, sadness) from the tone and facial expression of the user when speaking.
[0547] Multimodal Data Integration Module
[0548] The voice and gesture data collected by the device is sent to a server, where it is converted into text using speech recognition technology. Gesture data is also sent to the server, where the user's movements are analyzed using a motion analysis algorithm. These multiple input data are integrated to create a standardized dataset, which also includes emotional information recognized by an emotion engine.
[0549] Generative Model Module
[0550] The standardized dataset created above is input into a generative model. The server runs the generative model (e.g., GPT-3) to generate a storyline based on the user's input in real time. The generative model also uses the user's emotional information to adjust the story and generate a more convincing response.
[0551] Real-time response generation module
[0552] Based on the generated storyline, the server generates responses for characters (e.g., dragons) and environments (e.g., the caves where dragons live). The generation process includes specific actions such as the dragon nodding and talking. Reactions based on the user's emotions are also incorporated. For example, if the user is angry, a scenario will be generated in which the dragon speaks words to calm them down.
[0553] Output Display Module
[0554] The data generated by the server is sent to the device, which then displays it in an appropriate format for the user. Audio data is played through the speaker, and text, images, and animations are displayed on the screen. For example, a dragon may talk to the user while moving realistically, presenting new options to the user (e.g., "Ask the dragon" or "Fight the dragon").
[0555] Specific examples
[0556] Consider a scenario where a user issues the voice command "talk to the dragon" and simultaneously smiles and waves:
[0557] 1. The user says "talk to the dragon" and smiles and waves.
[0558] 2. The device captures voice, gesture, and facial expression data and sends it to the server.
[0559] 3. The server converts the speech into text and analyzes gestures and facial expressions.
[0560] 4. The emotion engine recognizes the user's smile and determines it to be a positive emotion (joy).
[0561] 5. The integrated dataset is fed into a generative model to generate a storyline.
[0562] 6. The server generates a dragon response based on the positive emotion (e.g., "You're in a good mood today, adventurer. How can I help you?").
[0563] 7. The server sends the generated response to the terminal, which presents it to the user in real time.
[0564] This system allows users to have a highly personalized, emotionally immersive gaming experience, while also allowing game developers to eliminate the need for complex scenario design, improving development flexibility and creativity.
[0565] The processing flow will be explained below.
[0566] Step 1:
[0567] The user puts on the VR headset and says "talk to the dragon." The device captures this voice through the microphone. At the same time, the device captures the user's hand gesture through the camera and sensors. The device also captures the user's facial expressions (e.g., smiling).
[0568] Step 2:
[0569] The terminal transmits the captured voice data, gesture data, and facial expression data to a server.
[0570] Step 3:
[0571] The server converts the received voice data into text using voice recognition technology. The voice recognition engine analyzes the voice waveform and generates the corresponding text data.
[0572] Step 4:
[0573] The server analyzes the gesture data and uses a motion analysis algorithm to identify the user's hand gestures and analyze their meaning. For example, it recognizes the "waving" gesture.
[0574] Step 5:
[0575] The server analyzes the facial expression data. It uses an expression analysis algorithm to recognize the user's facial expressions and identify their emotional state (e.g., joy, anger, sadness, or happiness). For example, it recognizes the emotion of "joy" from the user's smile.
[0576] Step 6:
[0577] The server integrates the speech, gesture, and facial expression data. The integration process brings together speech text, gesture data, and emotion information into a single standardized dataset.
[0578] Step 7:
[0579] The server inputs the combined dataset into a generative model (e.g., GPT-3), which then runs the model and generates a storyline based on the user's input in real time.
[0580] Step 8:
[0581] The server analyzes the generated storyline and generates a character (e.g., a dragon)'s reaction based on the user's emotional information. For example, if the user has the emotion "joy," the server generates a scenario in which the dragon responds in a friendly manner.
[0582] Step 9:
[0583] The server sends generated data, including character reactions, to the device. This data includes text, audio, and animation data.
[0584] Step 10:
[0585] The device processes the received data and presents it to the user in real time, playing the dragon's voice through the speaker and showing the dragon's movements on the display.
[0586] Step 11:
[0587] The device presents the user with new options, such as "Ask the dragon" or "Fight the dragon."
[0588] Step 12:
[0589] The user selects their next action. Once a new selection is made, the device again sends this information to the server, and the process begins again at step 1.
[0590] This series of steps enables the system to dynamically generate a story while taking into account the user's emotions, providing an interactive and immersive gaming experience.
[0591] Example 2
[0592] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0593] Previous interactive storytelling systems lacked the ability to recognize users' emotions in real time and adjust the story and character responses based on those emotions. As a result, users often did not receive responses that reflected their own emotions and actions, resulting in a less immersive experience. Furthermore, there was a lack of technical means to provide flexible and advanced interactions.
[0594] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0595] In this invention, the server includes means for acquiring multimodal input from a user, means for integrating and standardizing the multimodal input, means for inputting the standardized data into a generative model to generate a story in real time, means for recognizing the user's emotions, means for reflecting the recognized emotional information in a real-time response, and means for presenting the response to the user. This enables real-time responses that are in line with the user's emotions and behavior, providing a more immersive interactive experience.
[0596] "Multimodal input" refers to input from a user that includes data in multiple formats, such as text, audio, images, and video.
[0597] "Means of integration and standardization" refers to the process of converting data in multiple different formats into a consistent format and compiling it into a single data set.
[0598] A "generative model" refers to an artificial intelligence model that generates new text or storylines based on given input data.
[0599] A "means for generating stories in real time" is a means for instantly responding to input from the user and generating new stories on the spot.
[0600] "Means for generating character and environmental responses in real time" refers to means for instantly generating character actions and environmental changes based on the generated story.
[0601] "Means for recognizing user emotions" refers to means for identifying a user's emotional state in real time using voice analysis or image analysis.
[0602] "Means for reflecting emotional information in real-time responses" refers to means for adjusting the responses and stories generated based on the recognized emotions of the user.
[0603] "Means for presenting responses to users" refers to the means by which the generated responses or stories are conveyed to users in the form of audio, text, animation, etc.
[0604] This invention is an interactive storytelling system that generates a story in real time based on the user's multimodal input and presents it to the user. It also incorporates an emotion engine, which recognizes the user's emotions and adjusts the story and character responses based on those emotions. This system mainly consists of the following components: a user input processing module, an emotion engine, a multimodal data integration module, a generative model module, a real-time response generation module, and an output display module.
[0605] User Input Processing Module
[0606] Users wear a VR headset and interact with the game using voice and gestures. The device captures these interactions through a microphone, camera, and gesture sensor. For example, a user can issue a voice command to "talk to the dragon" and wave their hand.
[0607] Emotion Engine
[0608] The server runs an emotion engine using voice and image data acquired from the user. The emotion engine recognizes the user's emotions using voice analysis and detects and identifies the user's facial expressions using image analysis. This allows the engine to determine the user's current emotion (e.g., joy, anger, sadness) from the tone and facial expression of the user when speaking.
[0609] Multimodal Data Integration Module
[0610] The device sends collected voice and gesture data to a server, which converts the data into text using speech recognition technology. Gesture data is also sent to the server, where the user's movements are analyzed using a motion analysis algorithm. These multiple inputs are combined to create a standardized dataset, which also includes emotional information recognized by an emotion engine.
[0611] Generative Model Module
[0612] The standardized dataset is then fed into a generative model. The server runs the generative model (e.g., a large-scale language model) to generate a storyline based on the user's input in real time. The generative model also uses the user's emotional information to adjust the story and generate a more convincing response.
[0613] Real-time response generation module
[0614] Based on the generated storyline, the server generates responses for characters (e.g., dragons) and environments (e.g., the caves where the dragons live). The generation process includes specific actions such as the dragon nodding and talking. Reactions based on the user's emotions can also be incorporated. For example, if the user is angry, a scenario will be generated in which the dragon speaks words to calm them down.
[0615] Output Display Module
[0616] The data generated by the server is sent to the device, which then displays it in an appropriate format for the user. Audio data is played through the speaker, and text, images, and animations are displayed on the screen. For example, a dragon moves realistically and speaks to the user, presenting new options to the user.
[0617] Specific examples
[0618] Consider a scenario where the user gives the voice command "talk to the dragon" and smiles and waves:
[0619] 1. The user says "talk to the dragon" and smiles and waves.
[0620] 2. The device captures voice, gesture, and facial expression data and sends it to the server.
[0621] 3. The server converts the speech into text and analyzes gestures and facial expressions.
[0622] 4. The emotion engine recognizes the user's smile and determines it to be a positive emotion (joy).
[0623] 5. The integrated dataset is fed into a generative model to generate a storyline.
[0624] 6. The server generates a dragon response based on the positive emotion (e.g., "You're in a good mood today, adventurer. How can I help you?").
[0625] 7. The server sends the generated response to the terminal, which presents it to the user in real time.
[0626] Prompt Sentence Examples
[0627] "The user smiles and says the voice command 'talk to the dragon.' Generate a response from the dragon."
[0628] "If the user is angry, write a line for the dragon that will calm them down."
[0629] This system allows users to have a highly personalized, emotionally immersive experience, while also allowing game developers to eliminate the need for complex scenario design, improving development flexibility and creativity.
[0630] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0631] Step 1:
[0632] The user puts on the VR headset and begins interacting with it using voice and gestures. For example, they might say, "Talk to the dragon," and then wave their hand. The device's microphone, camera, and gesture sensors capture this voice, image, and movement data.
[0633] Input: Audio, video, gesture data
[0634] Output: Raw captured data
[0635] Step 2:
[0636] The voice data and gesture data captured by the device are sent to the server via the network, along with the voice data captured by the microphone and the gesture data captured by the camera and gesture sensor.
[0637] Input: Raw captured data
[0638] Output: Voice and gesture data sent to the server
[0639] Step 3:
[0640] The server uses voice recognition technology (e.g., a voice-to-text conversion system) to convert the captured voice data into text, and uses a motion analysis algorithm (e.g., a motion detection algorithm) to analyze the gesture data and recognize the user's movements.
[0641] Input: Voice data, gesture data
[0642] Data processing: Converts voice data into text, and gesture data into motion recognition information
[0643] Output: Text data, motion recognition data
[0644] Step 4:
[0645] The server runs an emotion engine (e.g., a voice emotion analysis system, an expression analysis system) to analyze the voice tone and facial expressions, thereby determining whether the user is expressing an emotion such as joy or anger.
[0646] Input: Audio data, video data
[0647] Data processing: voice tone analysis, facial expression analysis
[0648] Output: Emotion recognition data
[0649] Step 5:
[0650] The server combines the text data, action recognition data, and emotion recognition data to create a standardized dataset, which is then input into the subsequent generative model.
[0651] Input: Text data, action recognition data, emotion recognition data
[0652] Data processing: data integration and standardization
[0653] Output: Standardized dataset
[0654] Step 6:
[0655] The server runs a generative AI model (e.g., a large-scale language model) to generate storylines in real time based on the integrated dataset, for example, generating positive stories based on what makes the user happy.
[0656] Input: Standardized dataset
[0657] Data processing: Story generation using generative AI models
[0658] Output: Generated storyline
[0659] Step 7:
[0660] The server generates character actions and conversations based on the generated storyline. For example, it generates lines and actions for a dragon to speak to the user. The response changes depending on the user's emotions.
[0661] Input: Storyline, emotion recognition data
[0662] Data processing: Character response generation
[0663] Output: Character movement data, character conversation data
[0664] Step 8:
[0665] The server sends the generated character movement data and conversation data to the device, which then presents this to the user in real time. Audio is played through the speaker and animation is shown on the display. For example, a dragon could move realistically and say, "You're in a good mood today, adventurer. Is there anything I can help you with?"
[0666] Input: Character movement data, character conversation data
[0667] Output: Real-time display to the user
[0668] (Application example 2)
[0669] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0670] Conventional interactive storytelling systems do not take into account the user's emotional state, resulting in a lack of individualized optimization of the user's experience and difficulty in achieving emotional empathy. Furthermore, it is difficult to effectively integrate the user's diverse inputs (voice, facial expressions, gestures), making it difficult to generate appropriate responses in real time. Furthermore, it has been challenging to realize a system that is not dependent on a specific device and uses general-purpose hardware.
[0671] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0672] In this invention, the server includes: means for acquiring multimodal input from a user; means for integrating and standardizing the multimodal input; means for inputting the standardized data into a generative model to generate a story in real time; means for generating responses of characters and the environment in real time based on the generated story; means for presenting the responses to the user; means for analyzing the user's emotions and adjusting the generated story and character responses based on the emotions; means for acquiring the user's voice, facial expressions, and gestures using a smartphone camera and microphone; and means for displaying the story generated in real time on a display device according to the user's emotional state, thereby enabling interactive storytelling that is individually optimized according to the user's emotional state.
[0673] "Means for acquiring multimodal input from a user" refers to a device or system for simultaneously acquiring multiple forms of input, such as a user's voice, facial expressions, and gestures.
[0674] The "means for integrating and standardizing the multimodal input" refers to a device or system that processes multiple acquired input formats in a unified manner and converts them into a format that is easy to analyze.
[0675] "Means for inputting the standardized data into a generative model and generating a story in real time" refers to a device or system that generates a story in real time using an artificial intelligence model based on the integrated and standardized data.
[0676] "Means for generating character and environmental responses in real time based on the generated story" refers to a device or system that generates character and environmental responses in real time according to a story created by a generative model.
[0677] "Means for presenting the response to the user" refers to a device or system that displays or outputs the response of the generated character or environment to the user.
[0678] "Means for analyzing the user's emotions and adjusting the generated story and character responses based on said emotions" refers to a device or system that analyzes the user's emotional state and adjusts the generated story and character responses based on that data.
[0679] "Means for capturing a user's voice, facial expressions, and gestures using a smartphone's camera and microphone" refers to a device or system that captures a user's voice, facial expressions, and gestures using a smartphone's built-in camera and microphone.
[0680] "Means for displaying a story generated in real time on a display device according to the user's emotional state" refers to a device or system for appropriately displaying to a user a story generated in real time based on the user's emotional state.
[0681] This invention uses a system that analyzes various user inputs (voice, facial expressions, gestures) in real time and generates an interactive story based on the analysis. This system mainly uses a smartphone and its built-in camera and microphone.
[0682] Hardware and software used
[0683] 1. Smartphone: Uses a camera and microphone to capture the user's voice, facial expressions, and gestures.
[0684] 2. Server: Data processing, sentiment analysis, story generation, character response generation. The main software used here is as follows:
[0685] OpenCV: A library used for camera capture and image analysis.
[0686] speech_recognition: The library to use for speech recognition.
[0687] transformers (GPT-2): A library used as a generative model.
[0688] TextBlob: A library used for sentiment analysis.
[0689] Processing flow
[0690] 1. User Input Processing:
[0691] The user speaks and gestures into the smartphone's camera, which then captures the audio and image data.
[0692] 2. Multimodal Data Integration:
[0693] The smartphone captures audio and image data and sends it to the server. The audio data is converted to text using the speech_recognition library. The image data is analyzed for gestures and facial expressions using OpenCV.
[0694] 3. Emotion analysis:
[0695] The converted voice data and image analysis results are sent to the server, where the TextBlob library analyzes the user's emotions, classifying them as positive, negative, or neutral.
[0696] 4. Story generation:
[0697] The data combined with the sentiment analysis results is fed into a GPT-2 model from the transformers library to generate an interactive story based on the user's emotional state.
[0698] 5. Character response generation:
[0699] Based on the story created by the generative model, specific responses from the character are generated. For example, when a dragon speaks to the user, it responds appropriately based on the user's emotions.
[0700] 6. Real-time display:
[0701] The generated story and character responses are displayed on the smartphone screen and played back aloud.
[0702] Specific examples
[0703] Consider a scenario where the user says "talk to the dragon" and smiles and waves:
[0704] 1. The smartphone's microphone and camera capture the user's voice, facial expressions, and gestures.
[0705] 2. The captured data is sent to a server, where the voice data is converted into text and the image data is analyzed for gestures and facial expressions.
[0706] 3. The TextBlob library parses the user's smile as a positive emotion.
[0707] 4. The combined data is fed into the GPT-2 model to generate prompts for the user to talk to the dragon in a positive manner.
[0708] 5. An interactive story is generated based on the generated prompts.
[0709] 6. As part of the story, the dragon responds to the user by saying, "You're in a good mood today, adventurer. How can I help you?"
[0710] Prompt Sentence Examples
[0711] An example prompt that describes a situation in which the user says "talk to the dragon" and smiles and waves is:
[0712] The user says "talk to the dragon," smiles, and waves. The user is pleased. How does the dragon respond?
[0713] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0714] Step 1:
[0715] The user speaks and gestures into the smartphone camera. The smartphone camera and microphone capture audio and image data. In this step, the user's voice input and gesture input are collected. The input data are audio files and image files.
[0716] Step 2:
[0717] The device sends captured voice and image data to the server. The voice data is converted to text using the speech_recognition library. The image data is analyzed for gestures and facial expressions using OpenCV. The input is the captured voice data and image data, and the output is text data and analyzed gesture and facial expression data.
[0718] Step 3:
[0719] The server inputs the converted voice data and image analysis results into the TextBlob library to analyze the user's emotions. The emotion analysis results are classified as positive, negative, or neutral. The input is text data and image analysis data, and the output is the user's emotional state.
[0720] Step 4:
[0721] The integrated data (text data, gesture analysis data, and emotion analysis data) is input to the GPT-2 model in the transformers library on the server to generate an interactive story based on the user's emotional state. The input is the integrated data, and the output is the generated story.
[0722] Step 5:
[0723] The server generates specific responses for the characters based on the generated story. For example, when a dragon speaks to a user, it responds appropriately based on the user's emotions. The input is the generated story, and the output is the character's response.
[0724] Step 6:
[0725] The server sends the generated story and character responses to the device, which displays them on the smartphone screen and plays them audibly. The input is the character responses and the generated story, and the output is an interactive story experience that is displayed and played back to the user.
[0726] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0727] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0728] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0729] [Third embodiment]
[0730] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0731] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0732] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0733] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0734] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0735] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0736] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0737] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0738] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0739] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0740] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0741] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0742] This invention is an interactive storytelling system that generates stories in real time based on the user's multimodal input and presents them to the user. This system mainly consists of the following components: a user input processing module, a multimodal data integration module, a generative model module, a real-time response generation module, and an output display module.
[0743] User Input Processing Module
[0744] Users wear a VR headset and interact with the game using voice and gestures, which are then captured by the device through a microphone, camera, and gesture sensors. For example, a user can speak to a dragon by waving their hand or issuing a voice command.
[0745] Multimodal Data Integration Module
[0746] The voice data collected by the device is sent to a server, where it is converted to text using speech recognition technology. Gesture data is also sent to the server, where the user's movements are analyzed using a motion analysis algorithm. These multiple inputs are then combined to create a standardized data set.
[0747] Generative Model Module
[0748] The standardized dataset created above is input into a generative model, which is then run by a server to generate a storyline based on the user's choices and actions. This generative model can be, for example, GPT or another generative AI model, and generates a story in real time in response to the user's actions.
[0749] Real-time response generation module
[0750] Based on the generated storyline, the server generates responses for characters (e.g., dragons) and environments (e.g., the caves in which the dragons live), including the dragon's dialogue, actions, and changes to the surrounding environment. For example, the dragon may start speaking to answer the user's questions.
[0751] Output Display Module
[0752] The data generated by the server is sent to the device, which then displays it in the appropriate format for the user: audio data is played through the speakers, and text, images, and animations are displayed on the screen. For example, a scene in which a dragon moves and talks to the user in a realistic manner is projected on the user's VR headset.
[0753] Specific examples
[0754] When a user issues the voice command "talk to dragon", the following occurs:
[0755] 1. The user says "talk to dragon" aloud.
[0756] 2. The device captures the user's hand-waving gesture along with the audio.
[0757] 3. The device sends the captured data to the server.
[0758] 4. The server converts the speech to text and analyzes the gestures.
[0759] 5. The integrated dataset is fed into a generative model to generate a storyline.
[0760] 6. The server generates the dragon's response (e.g., "Hello, adventurer. What do you want?").
[0761] 7. The server sends the generated response to the terminal, which presents it to the user in real time.
[0762] This system allows users to have a highly personalized and immersive gaming experience, while also allowing game developers to omit complex scenario design, improving development flexibility and creativity.
[0763] The processing flow will be explained below.
[0764] Step 1:
[0765] The user puts on the VR headset and says, "Talk to the dragon." The device captures this voice through the microphone. At the same time, the device captures the user's hand gestures through the camera and sensors.
[0766] Step 2:
[0767] The device sends the captured voice data to the server, and also sends the captured gesture data to the server along with the voice data.
[0768] Step 3:
[0769] The server converts the received voice data into text using voice recognition technology. The voice recognition engine analyzes the voice waveform and generates the corresponding text data.
[0770] Step 4:
[0771] The server analyzes the gesture data and uses a motion analysis algorithm to identify the user's hand gestures and analyze their meaning. For example, it recognizes the "waving" gesture.
[0772] Step 5:
[0773] The server merges the speech and gesture data. The merger process combines the speech text and gesture data into a single standardized data set.
[0774] Step 6:
[0775] The server inputs the combined dataset into a generative model (e.g., GPT-3), which then runs the model and generates a storyline based on the user's input in real time.
[0776] Step 7:
[0777] The server analyzes the generated storyline and generates the character (dragon)'s reactions. The generation process includes specific actions such as the dragon nodding and talking.
[0778] Step 8:
[0779] The server sends generated data, including the dragon's reactions, to the device. This data includes text, audio, and animation data.
[0780] Step 9:
[0781] The device processes the received data and presents it to the user in real time, playing the dragon's voice through the speaker and showing the dragon's movements on the display.
[0782] Step 10:
[0783] The device presents the user with new options, such as "Ask the dragon" or "Fight the dragon."
[0784] Step 11:
[0785] The user selects their next action. Once a new selection is made, the device again sends this information to the server, and the process begins again at step 1.
[0786] Through this series of steps, the system dynamically generates a story based on user input, providing an interactive and immersive gaming experience.
[0787] Example 1
[0788] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0789] Conventional interactive entertainment systems have struggled to respond to user input in real time or generate complex storylines. Furthermore, they lacked technology to integrate and standardize multiple input modes (voice, gestures, etc.), making it difficult to provide an immersive experience that takes advantage of diverse user interactions. Therefore, there is a need for systems that allow users to enjoy more advanced interactive experiences and increase the flexibility and creativity of game developers.
[0790] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0791] In this invention, the server includes means for acquiring multimodal input from a user, means for integrating and standardizing the multimodal input, means for inputting the standardized data into a generative model and generating a story in real time, means for generating character and environmental responses in real time based on the generated story, means for presenting the responses to the user, means for the user to wear a head-mounted display and interact with the user through voice and gestures, means for converting voice data into text and analyzing gesture data, means for generating a storyline based on the user's actions using an artificial intelligence model as a generative model, and means for outputting the generated character's actions and environmental changes in real time, thereby enabling the user to have a highly personalized and immersive interactive experience.
[0792] "Multimodal input" refers to data input by users in multiple different modes, such as voice, gestures, images, and video.
[0793] A "head-mounted display" is a device worn by the user that displays visual information directly in front of the user's eyes.
[0794] A "generative model" refers to an artificial intelligence model that generates new content or stories based on user actions and input.
[0795] A "voice recognition module" is a piece of software that converts voice data into text data.
[0796] A "gesture sensor" is a device that detects a user's movements and captures their action data.
[0797] "Motion analysis algorithm" refers to an algorithm that analyzes captured gesture data to identify specific user movements.
[0798] An "integrated dataset" is a set of data that has been integrated and standardized from data input from multiple different modes.
[0799] "Real-time response generation" refers to the process of generating instant character and environmental reactions based on user input.
[0800] The "output display module" is a module for presenting data generated by the server to the user in an appropriate format.
[0801] The present invention relates to an interactive entertainment system that generates a story in real time based on a user's multimodal input and presents it to the user. The system consists of a user input processing module, a multimodal data integration module, a generative model module, a real-time response generation module, and an output display module.
[0802] User Input Processing Module
[0803] The user wears a head-mounted display and interacts with it using voice and gestures. The device captures these voices and gestures using a microphone, camera, and gesture sensor. For example, if a user says "talk to the dragon" and waves their hand, this data is captured.
[0804] Multimodal Data Integration Module
[0805] The voice data captured by the device is sent to a server where it is converted into text using speech recognition technology. Similarly, gesture data is sent to a server where a motion analysis algorithm analyzes the user's specific movements. This process standardizes the voice and gesture data and compiles them into a unified data set.
[0806] Generative Model Module
[0807] The server inputs the combined dataset into a generative model, such as an advanced generative AI model like GPT-3, which then generates a storyline in real time based on the user's actions.
[0808] Real-time response generation module
[0809] Based on the generated storyline, the server generates responses from characters and the environment in real time. For example, a scene is generated in which a dragon responds to a user's speech by saying, "Hello, adventurer. What do you want?" Changes to the dragon's movements and the surrounding environment (for example, the dragon moving forward or the lights in a cave turning on) are also generated in real time.
[0810] Output Display Module
[0811] The server sends the generated voice, text, image, and animation data to the device. The device then displays the data on the user's head-mounted display, and the audio is played through the speakers. For example, a scene in which a dragon moves and talks to you in a realistic way, or a scene in which lights in a cave light up, may be displayed on the user's head-mounted display.
[0812] Specific examples
[0813] When a user says "talk to dragon" and waves their hand, the following happens:
[0814] 1. The user says "talk to the dragon" and waves their right hand.
[0815] 2. The device captures these voices and gestures and sends them to the server.
[0816] 3. The server converts the voice data into text and analyzes the gesture data.
[0817] 4. The server inputs the integrated dataset into the generative model and generates a storyline.
[0818] 5. The server generates the dragon's responses based on the generated storyline and generates the dragon's movements in real time.
[0819] 6. The server sends the generated response and action to the terminal, which presents it to the user.
[0820] The system allows users to have a highly personalized, immersive, and interactive experience, and an example prompt using the generative AI model is "Generate a scene where the user talks to the dragon and asks what to do next."
[0821] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0822] Step 1: User Input Capture
[0823] The user wears a head-mounted display and interacts with it using voice and gestures. Specifically, the user says "talk to the dragon" and waves their right hand at the same time. The device captures voice with a microphone and gestures with a camera and gesture sensor. Voice data and gesture data are acquired as input. The output is the captured voice data and gesture data.
[0824] Step 2: Sending data
[0825] The device sends the captured voice and gesture data to the server. Specifically, the data is transferred to the server via the Internet in real time. The input is the captured voice and gesture data. The output is the data sent to the server.
[0826] Step 3: Speech and gesture analysis
[0827] The server converts the received voice data into text data using a voice recognition module. For example, the text "talk to the dragon" is acquired. At the same time, the server analyzes the gesture data using a motion analysis algorithm to identify the user's specific actions. The input is voice data and gesture data, and the output is text data and analyzed motion data.
[0828] Step 4: Data integration and standardization
[0829] The server integrates the text data from speech recognition and the gesture analysis data to create a standardized dataset. This is to unify data from different modes. Specifically, the server standardizes the data format and generates the integrated dataset. The input is text data and analyzed gesture data, and the output is the integrated dataset.
[0830] Step 5: Story Generation
[0831] The server inputs the standardized dataset into a generative model. GPT-3 or other models are used as generative models, generating a storyline based on user behavior. Specifically, the server runs the generative model and generates the next development in the story. The input is the integrated dataset, and the output is the generated storyline.
[0832] Step 6: Real-time response generation
[0833] The server determines the response of the character (for example, a dragon) based on the generated storyline. Environmental changes (for example, turning on the lights in a cave) are also generated. As specific actions, the server generates the dragon's lines and actions. For example, a response such as "Hello, adventurer. What do you want?" is generated. The input is the generated storyline, and the output is the character's response and environmental change data.
[0834] Step 7: Sending Responses and Actions
[0835] The server sends the generated response and action data to the terminal. Specifically, the server sends data to the terminal in real time via the Internet. The input is the character's response and environmental change data, and the output is the data sent to the terminal.
[0836] Step 8: Presenting the output
[0837] The device receives audio, text, images, and animation data from the server and presents it to the user. Specific operations include displaying a scene of a dragon speaking to the user on the head-mounted display and lighting up the cave, and playing audio through the speaker. The input is data sent from the server, and the output is the visuals and audio presented to the user.
[0838] (Application example 1)
[0839] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0840] In conventional storytelling systems, user interactions are fixed, making it difficult to generate stories in real time based on individual actions and utterances. Furthermore, the limited means of interaction mean that users lack a sense of immersion. Furthermore, the inability to integrate and utilize multi-modal data, such as gestures and voice, limits the user experience.
[0841] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0842] In this invention, the server includes: means for acquiring multimodal input from a user; means for integrating and standardizing the multimodal input; means for inputting the standardized data into a generative model to generate a story in real time; means for generating character and environmental responses in real time based on the generated story; means for presenting the responses to the user through a smartphone or tablet display; means for capturing and recognizing user gestures with a camera; means for combining the gestures and voice input to form story generation prompts; and means for playing the generated story aloud. This enables rich interactions that combine the user's voice and gestures, and makes it possible to provide a personalized story experience according to individual actions and utterances.
[0843] "Means for obtaining multimodal input from the user" refers to a function for capturing data in multiple forms, such as text, audio, images, and video, generated by the user.
[0844] "Means to integrate and standardize multimodal input" refers to a function that converts acquired data in multiple formats into a single unified dataset for easier processing.
[0845] "Means of inputting data into a generative model and generating a story in real time" refers to a function that allows standardized data to be input into a generative model (e.g., a generative AI model) to create a dynamic story on the spot.
[0846] "Means for generating character and environmental responses in real time" is a function that allows the environment, such as characters and backgrounds, to react in real time based on the generated story, creating appropriate actions and situations.
[0847] "Means of presenting to the user" refers to the function for showing or listening to the generated story or response to the user through a smartphone or tablet display, speaker, etc.
[0848] "Means for capturing and recognizing user gestures with a camera" refers to a function that uses a camera to capture the user's physical movements (e.g., hand gestures) and analyzes and understands them.
[0849] "Means for combining gestures and voice input to form story generation prompts" is a function for generating instructions (prompts) to be given to a generative AI model based on data that combines the user's gestures and voice input.
[0850] "Means for playing the generated story aloud" is a function for converting the story created by the generative model into audio and playing it through a speaker.
[0851] To implement this invention, it is necessary to acquire user voice and gesture input, integrate the data, and input it into a generative AI model to generate a story in real time. A specific embodiment of this is shown below.
[0852] Hardware and software used
[0853] 1. Hardware
[0854] Smartphone or tablet: Used to provide the user interface.
[0855] Camera: Used to capture user gestures.
[0856] Microphone: Used to collect the user's voice.
[0857] Speaker: Used to play the generated audio back to the user.
[0858] 2. Software
[0859] Google Cloud Speech-to-Text API: Used to convert user speech into text.
[0860] OpenCV: Used to recognize user gestures.
[0861] OpenAI GPT (Generative AI Model): Used to generate stories in real time based on user input.
[0862] gTTS (Google Text-to-Speech) and playsound: Used to play the audio of the generated story.
[0863] System Operation Overview
[0864] 1. Acquiring Multimodal Input
[0865] The user speaks into the smartphone or tablet and simultaneously gestures in front of the camera.
[0866] Smartphones capture audio through their microphones and gestures through their cameras.
[0867] 2. Data integration and standardization
[0868] The captured audio data is converted to text using the Google Cloud Speech-to-Text API.
[0869] Similarly, the captured gesture data is analyzed using OpenCV to create a standardized dataset.
[0870] 3. Story Generation
[0871] The combined dataset is fed into a generative AI model (OpenAI GPT) to generate a story in real time.
[0872] Generative AI models create appropriate character and environmental responses based on user actions and speech.
[0873] 4. Real-time response generation and presentation
[0874] The generated story is converted into audio data using gTTS and played back to the user through the smartphone's speakers.
[0875] At the same time, the display shows related images and text.
[0876] Specific examples
[0877] For example, if a user issues the voice command "talk to the wizard" and makes a hand waving gesture, the system provides the following prompt sentence to the generative AI model:
[0878] text
[0879] User Voice: Talk to the Wizard
[0880] User gesture: wave
[0881] Generate the following story:
[0882] Based on this prompt, the generative AI model generates a story like this:
[0883] text
[0884] "Hello, traveller. My name is Eldritch, how may I help you?" the wizard replied in a kind voice.
[0885] This generated story is then converted into audio by gTTS and played back to the user through the smartphone's speakers, with related images and text also appearing on the display, creating a more immersive experience for the user.
[0886] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0887] Step 1: The user speaks a command (e.g., "Talk to the wizard") into a smartphone or tablet. The smartphone captures this voice through its microphone and saves it as audio data. At the same time, the camera captures the user's hand-waving gesture and saves it as video data. This provides audio and video data as input.
[0888] Step 2: The device converts the captured voice data to text data using the Google Cloud Speech-to-Text API. When converting voice data to text data, the input is the voice data, and the output is the converted text data. For example, the utterance "Talk to the wizard" becomes the text "Talk to the wizard."
[0889] Step 3: The device analyzes the captured video data using OpenCV and recognizes the user's gesture (e.g., waving). The input is the video data, and the output is the recognized gesture information. For example, the user's hand wave is recognized as a "wave."
[0890] Step 4: The device combines the results of converting the voice data to text and the recognized gesture data to generate a standardized dataset. The input is text data and gesture data, and the output is the combined dataset. For example, the dataset might have the format "Voice: Talk to the Wizard" and "Gesture: Wave."
[0891] Step 5: The server receives the integrated dataset and inputs it into the generative AI model (OpenAI GPT). The input is the integrated dataset, and the output is a story text based on the generative AI model. For example, based on the prompts "User voice: Talk to the wizard" and "User gesture: Wave", the story "Hello, traveler. My name is Eldritch. How may I help you?" is generated.
[0892] Step 6: The server converts the generated story text into audio data using gTTS (Google Text-to-Speech). The input is the generated story text, and the output is the audio data of the story. This audio data is sent to the smartphone.
[0893] Step 7: The device plays the received audio data to the user through the speaker. The input is audio data, and the output is audio playback. At the same time, related text and images are displayed on the screen, providing the user with an immersive experience.
[0894] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0895] The present invention is an interactive storytelling system that generates a story in real time based on the user's multimodal input and presents it to the user, and further incorporates an emotion engine to recognize the user's emotions and adjust the story and character responses based on those emotions. The system of the present invention mainly consists of the following components: a user input processing module, an emotion engine, a multimodal data integration module, a generative model module, a real-time response generation module, and an output display module.
[0896] User Input Processing Module
[0897] Users wear a VR headset and interact with the game using voice and gestures. The device captures these interactions through a microphone, camera, and gesture sensors. For example, a user can speak to a dragon by issuing a voice command or waving their hand.
[0898] Emotion Engine
[0899] The server operates an emotion engine using voice data and image data acquired from the user. The emotion engine recognizes the user's emotions using voice analysis and detects and identifies the user's facial expressions using image analysis. This allows the engine to determine the user's current emotion (e.g., joy, anger, sadness) from the tone and facial expression of the user when speaking.
[0900] Multimodal Data Integration Module
[0901] The voice and gesture data collected by the device is sent to a server, where it is converted into text using speech recognition technology. Gesture data is also sent to the server, where the user's movements are analyzed using a motion analysis algorithm. These multiple input data are integrated to create a standardized dataset, which also includes emotional information recognized by an emotion engine.
[0902] Generative Model Module
[0903] The standardized dataset created above is input into a generative model. The server runs the generative model (e.g., GPT-3) to generate a storyline based on the user's input in real time. The generative model also uses the user's emotional information to adjust the story and generate a more convincing response.
[0904] Real-time response generation module
[0905] Based on the generated storyline, the server generates responses for characters (e.g., dragons) and environments (e.g., the caves where dragons live). The generation process includes specific actions such as the dragon nodding and talking. Reactions based on the user's emotions are also incorporated. For example, if the user is angry, a scenario will be generated in which the dragon speaks words to calm them down.
[0906] Output Display Module
[0907] The data generated by the server is sent to the device, which then displays it in an appropriate format for the user. Audio data is played through the speaker, and text, images, and animations are displayed on the screen. For example, a dragon may talk to the user while moving realistically, presenting new options to the user (e.g., "Ask the dragon" or "Fight the dragon").
[0908] Specific examples
[0909] Consider a scenario where a user issues the voice command "talk to the dragon" and simultaneously smiles and waves:
[0910] 1. The user says "talk to the dragon" and smiles and waves.
[0911] 2. The device captures voice, gesture, and facial expression data and sends it to the server.
[0912] 3. The server converts the speech into text and analyzes gestures and facial expressions.
[0913] 4. The emotion engine recognizes the user's smile and determines it to be a positive emotion (joy).
[0914] 5. The integrated dataset is fed into a generative model to generate a storyline.
[0915] 6. The server generates a dragon response based on the positive emotion (e.g., "You're in a good mood today, adventurer. How can I help you?").
[0916] 7. The server sends the generated response to the terminal, which presents it to the user in real time.
[0917] This system allows users to have a highly personalized, emotionally immersive gaming experience, while also allowing game developers to eliminate the need for complex scenario design, improving development flexibility and creativity.
[0918] The processing flow will be explained below.
[0919] Step 1:
[0920] The user puts on the VR headset and says "talk to the dragon." The device captures this voice through the microphone. At the same time, the device captures the user's hand gesture through the camera and sensors. The device also captures the user's facial expressions (e.g., smiling).
[0921] Step 2:
[0922] The terminal transmits the captured voice data, gesture data, and facial expression data to a server.
[0923] Step 3:
[0924] The server converts the received voice data into text using voice recognition technology. The voice recognition engine analyzes the voice waveform and generates the corresponding text data.
[0925] Step 4:
[0926] The server analyzes the gesture data and uses a motion analysis algorithm to identify the user's hand gestures and analyze their meaning. For example, it recognizes the "waving" gesture.
[0927] Step 5:
[0928] The server analyzes the facial expression data. It uses an expression analysis algorithm to recognize the user's facial expressions and identify their emotional state (e.g., joy, anger, sadness, or happiness). For example, it recognizes the emotion of "joy" from the user's smile.
[0929] Step 6:
[0930] The server integrates the speech, gesture, and facial expression data. The integration process brings together speech text, gesture data, and emotion information into a single standardized dataset.
[0931] Step 7:
[0932] The server inputs the combined dataset into a generative model (e.g., GPT-3), which then runs the model and generates a storyline based on the user's input in real time.
[0933] Step 8:
[0934] The server analyzes the generated storyline and generates a character (e.g., a dragon)'s reaction based on the user's emotional information. For example, if the user has the emotion "joy," the server generates a scenario in which the dragon responds in a friendly manner.
[0935] Step 9:
[0936] The server sends generated data, including character reactions, to the device. This data includes text, audio, and animation data.
[0937] Step 10:
[0938] The device processes the received data and presents it to the user in real time, playing the dragon's voice through the speaker and showing the dragon's movements on the display.
[0939] Step 11:
[0940] The device presents the user with new options, such as "Ask the dragon" or "Fight the dragon."
[0941] Step 12:
[0942] The user selects their next action. Once a new selection is made, the device again sends this information to the server, and the process begins again at step 1.
[0943] This series of steps enables the system to dynamically generate a story while taking into account the user's emotions, providing an interactive and immersive gaming experience.
[0944] Example 2
[0945] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0946] Previous interactive storytelling systems lacked the ability to recognize users' emotions in real time and adjust the story and character responses based on those emotions. As a result, users often did not receive responses that reflected their own emotions and actions, resulting in a less immersive experience. Furthermore, there was a lack of technical means to provide flexible and advanced interactions.
[0947] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0948] In this invention, the server includes means for acquiring multimodal input from a user, means for integrating and standardizing the multimodal input, means for inputting the standardized data into a generative model to generate a story in real time, means for recognizing the user's emotions, means for reflecting the recognized emotional information in a real-time response, and means for presenting the response to the user. This enables real-time responses that are in line with the user's emotions and behavior, providing a more immersive interactive experience.
[0949] "Multimodal input" refers to input from a user that includes data in multiple formats, such as text, audio, images, and video.
[0950] "Means of integration and standardization" refers to the process of converting data in multiple different formats into a consistent format and compiling it into a single data set.
[0951] A "generative model" refers to an artificial intelligence model that generates new text or storylines based on given input data.
[0952] A "means for generating stories in real time" is a means for instantly responding to input from the user and generating new stories on the spot.
[0953] "Means for generating character and environmental responses in real time" refers to means for instantly generating character actions and environmental changes based on the generated story.
[0954] "Means for recognizing user emotions" refers to means for identifying a user's emotional state in real time using voice analysis or image analysis.
[0955] "Means for reflecting emotional information in real-time responses" refers to means for adjusting the responses and stories generated based on the recognized emotions of the user.
[0956] "Means for presenting responses to users" refers to the means by which the generated responses or stories are conveyed to users in the form of audio, text, animation, etc.
[0957] This invention is an interactive storytelling system that generates a story in real time based on the user's multimodal input and presents it to the user. It also incorporates an emotion engine, which recognizes the user's emotions and adjusts the story and character responses based on those emotions. This system mainly consists of the following components: a user input processing module, an emotion engine, a multimodal data integration module, a generative model module, a real-time response generation module, and an output display module.
[0958] User Input Processing Module
[0959] Users wear a VR headset and interact with the game using voice and gestures. The device captures these interactions through a microphone, camera, and gesture sensor. For example, a user can issue a voice command to "talk to the dragon" and wave their hand.
[0960] Emotion Engine
[0961] The server runs an emotion engine using voice and image data acquired from the user. The emotion engine recognizes the user's emotions using voice analysis and detects and identifies the user's facial expressions using image analysis. This allows the engine to determine the user's current emotion (e.g., joy, anger, sadness) from the tone and facial expression of the user when speaking.
[0962] Multimodal Data Integration Module
[0963] The device sends collected voice and gesture data to a server, which converts the data into text using speech recognition technology. Gesture data is also sent to the server, where the user's movements are analyzed using a motion analysis algorithm. These multiple inputs are combined to create a standardized dataset, which also includes emotional information recognized by an emotion engine.
[0964] Generative Model Module
[0965] The standardized dataset is then fed into a generative model. The server runs the generative model (e.g., a large-scale language model) to generate a storyline based on the user's input in real time. The generative model also uses the user's emotional information to adjust the story and generate a more convincing response.
[0966] Real-time response generation module
[0967] Based on the generated storyline, the server generates responses for characters (e.g., dragons) and environments (e.g., the caves where the dragons live). The generation process includes specific actions such as the dragon nodding and talking. Reactions based on the user's emotions can also be incorporated. For example, if the user is angry, a scenario will be generated in which the dragon speaks words to calm them down.
[0968] Output Display Module
[0969] The data generated by the server is sent to the device, which then displays it in an appropriate format for the user. Audio data is played through the speaker, and text, images, and animations are displayed on the screen. For example, a dragon moves realistically and speaks to the user, presenting new options to the user.
[0970] Specific examples
[0971] Consider a scenario where the user gives the voice command "talk to the dragon" and smiles and waves:
[0972] 1. The user says "talk to the dragon" and smiles and waves.
[0973] 2. The device captures voice, gesture, and facial expression data and sends it to the server.
[0974] 3. The server converts the speech into text and analyzes gestures and facial expressions.
[0975] 4. The emotion engine recognizes the user's smile and determines it to be a positive emotion (joy).
[0976] 5. The integrated dataset is fed into a generative model to generate a storyline.
[0977] 6. The server generates a dragon response based on the positive emotion (e.g., "You're in a good mood today, adventurer. How can I help you?").
[0978] 7. The server sends the generated response to the terminal, which presents it to the user in real time.
[0979] Prompt Sentence Examples
[0980] "The user smiles and says the voice command 'talk to the dragon.' Generate a response from the dragon."
[0981] "If the user is angry, write a line for the dragon that will calm them down."
[0982] This system allows users to have a highly personalized, emotionally immersive experience, while also allowing game developers to eliminate the need for complex scenario design, improving development flexibility and creativity.
[0983] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0984] Step 1:
[0985] The user puts on the VR headset and begins interacting with it using voice and gestures. For example, they might say, "Talk to the dragon," and then wave their hand. The device's microphone, camera, and gesture sensors capture this voice, image, and movement data.
[0986] Input: Audio, video, gesture data
[0987] Output: Raw captured data
[0988] Step 2:
[0989] The voice data and gesture data captured by the device are sent to the server via the network, along with the voice data captured by the microphone and the gesture data captured by the camera and gesture sensor.
[0990] Input: Raw captured data
[0991] Output: Voice and gesture data sent to the server
[0992] Step 3:
[0993] The server uses voice recognition technology (e.g., a voice-to-text conversion system) to convert the captured voice data into text, and uses a motion analysis algorithm (e.g., a motion detection algorithm) to analyze the gesture data and recognize the user's movements.
[0994] Input: Voice data, gesture data
[0995] Data processing: Converts voice data into text, and gesture data into motion recognition information
[0996] Output: Text data, motion recognition data
[0997] Step 4:
[0998] The server runs an emotion engine (e.g., a voice emotion analysis system, an expression analysis system) to analyze the voice tone and facial expressions, thereby determining whether the user is expressing an emotion such as joy or anger.
[0999] Input: Audio data, video data
[1000] Data processing: voice tone analysis, facial expression analysis
[1001] Output: Emotion recognition data
[1002] Step 5:
[1003] The server combines the text data, action recognition data, and emotion recognition data to create a standardized dataset, which is then input into the subsequent generative model.
[1004] Input: Text data, action recognition data, emotion recognition data
[1005] Data processing: data integration and standardization
[1006] Output: Standardized dataset
[1007] Step 6:
[1008] The server runs a generative AI model (e.g., a large-scale language model) to generate storylines in real time based on the integrated dataset, for example, generating positive stories based on what makes the user happy.
[1009] Input: Standardized dataset
[1010] Data processing: Story generation using generative AI models
[1011] Output: Generated storyline
[1012] Step 7:
[1013] The server generates character actions and conversations based on the generated storyline. For example, it generates lines and actions for a dragon to speak to the user. The response changes depending on the user's emotions.
[1014] Input: Storyline, emotion recognition data
[1015] Data processing: Character response generation
[1016] Output: Character movement data, character conversation data
[1017] Step 8:
[1018] The server sends the generated character movement data and conversation data to the device, which then presents this to the user in real time. Audio is played through the speaker and animation is shown on the display. For example, a dragon could move realistically and say, "You're in a good mood today, adventurer. Is there anything I can help you with?"
[1019] Input: Character movement data, character conversation data
[1020] Output: Real-time display to the user
[1021] (Application example 2)
[1022] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1023] Conventional interactive storytelling systems do not take into account the user's emotional state, resulting in a lack of individualized optimization of the user's experience and difficulty in achieving emotional empathy. Furthermore, it is difficult to effectively integrate the user's diverse inputs (voice, facial expressions, gestures), making it difficult to generate appropriate responses in real time. Furthermore, it has been challenging to realize a system that is not dependent on a specific device and uses general-purpose hardware.
[1024] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1025] In this invention, the server includes: means for acquiring multimodal input from a user; means for integrating and standardizing the multimodal input; means for inputting the standardized data into a generative model to generate a story in real time; means for generating responses of characters and the environment in real time based on the generated story; means for presenting the responses to the user; means for analyzing the user's emotions and adjusting the generated story and character responses based on the emotions; means for acquiring the user's voice, facial expressions, and gestures using a smartphone camera and microphone; and means for displaying the story generated in real time on a display device according to the user's emotional state, thereby enabling interactive storytelling that is individually optimized according to the user's emotional state.
[1026] "Means for acquiring multimodal input from a user" refers to a device or system for simultaneously acquiring multiple forms of input, such as a user's voice, facial expressions, and gestures.
[1027] The "means for integrating and standardizing the multimodal input" refers to a device or system that processes multiple acquired input formats in a unified manner and converts them into a format that is easy to analyze.
[1028] "Means for inputting the standardized data into a generative model and generating a story in real time" refers to a device or system that generates a story in real time using an artificial intelligence model based on the integrated and standardized data.
[1029] "Means for generating character and environmental responses in real time based on the generated story" refers to a device or system that generates character and environmental responses in real time according to a story created by a generative model.
[1030] "Means for presenting the response to the user" refers to a device or system that displays or outputs the response of the generated character or environment to the user.
[1031] "Means for analyzing the user's emotions and adjusting the generated story and character responses based on said emotions" refers to a device or system that analyzes the user's emotional state and adjusts the generated story and character responses based on that data.
[1032] "Means for capturing a user's voice, facial expressions, and gestures using a smartphone's camera and microphone" refers to a device or system that captures a user's voice, facial expressions, and gestures using a smartphone's built-in camera and microphone.
[1033] "Means for displaying a story generated in real time on a display device according to the user's emotional state" refers to a device or system for appropriately displaying to a user a story generated in real time based on the user's emotional state.
[1034] This invention uses a system that analyzes various user inputs (voice, facial expressions, gestures) in real time and generates an interactive story based on the analysis. This system mainly uses a smartphone and its built-in camera and microphone.
[1035] Hardware and software used
[1036] 1. Smartphone: Uses a camera and microphone to capture the user's voice, facial expressions, and gestures.
[1037] 2. Server: Data processing, sentiment analysis, story generation, character response generation. The main software used here is as follows:
[1038] OpenCV: A library used for camera capture and image analysis.
[1039] speech_recognition: The library to use for speech recognition.
[1040] transformers (GPT-2): A library used as a generative model.
[1041] TextBlob: A library used for sentiment analysis.
[1042] Processing flow
[1043] 1. User Input Processing:
[1044] The user speaks and gestures into the smartphone's camera, which then captures the audio and image data.
[1045] 2. Multimodal Data Integration:
[1046] The smartphone captures audio and image data and sends it to the server. The audio data is converted to text using the speech_recognition library. The image data is analyzed for gestures and facial expressions using OpenCV.
[1047] 3. Emotion analysis:
[1048] The converted voice data and image analysis results are sent to the server, where the TextBlob library analyzes the user's emotions, classifying them as positive, negative, or neutral.
[1049] 4. Story generation:
[1050] The data combined with the sentiment analysis results is fed into a GPT-2 model from the transformers library to generate an interactive story based on the user's emotional state.
[1051] 5. Character response generation:
[1052] Based on the story created by the generative model, specific responses from the character are generated. For example, when a dragon speaks to the user, it responds appropriately based on the user's emotions.
[1053] 6. Real-time display:
[1054] The generated story and character responses are displayed on the smartphone screen and played back aloud.
[1055] Specific examples
[1056] Consider a scenario where the user says "talk to the dragon" and smiles and waves:
[1057] 1. The smartphone's microphone and camera capture the user's voice, facial expressions, and gestures.
[1058] 2. The captured data is sent to a server, where the voice data is converted into text and the image data is analyzed for gestures and facial expressions.
[1059] 3. The TextBlob library parses the user's smile as a positive emotion.
[1060] 4. The combined data is fed into the GPT-2 model to generate prompts for the user to talk to the dragon in a positive manner.
[1061] 5. An interactive story is generated based on the generated prompts.
[1062] 6. As part of the story, the dragon responds to the user by saying, "You're in a good mood today, adventurer. How can I help you?"
[1063] Prompt Sentence Examples
[1064] An example prompt that describes a situation in which the user says "talk to the dragon" and smiles and waves is:
[1065] The user says "talk to the dragon," smiles, and waves. The user is pleased. How does the dragon respond?
[1066] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1067] Step 1:
[1068] The user speaks and gestures into the smartphone camera. The smartphone camera and microphone capture audio and image data. In this step, the user's voice input and gesture input are collected. The input data are audio files and image files.
[1069] Step 2:
[1070] The device sends captured voice and image data to the server. The voice data is converted to text using the speech_recognition library. The image data is analyzed for gestures and facial expressions using OpenCV. The input is the captured voice data and image data, and the output is text data and analyzed gesture and facial expression data.
[1071] Step 3:
[1072] The server inputs the converted voice data and image analysis results into the TextBlob library to analyze the user's emotions. The emotion analysis results are classified as positive, negative, or neutral. The input is text data and image analysis data, and the output is the user's emotional state.
[1073] Step 4:
[1074] The integrated data (text data, gesture analysis data, and emotion analysis data) is input to the GPT-2 model in the transformers library on the server to generate an interactive story based on the user's emotional state. The input is the integrated data, and the output is the generated story.
[1075] Step 5:
[1076] The server generates specific responses for the characters based on the generated story. For example, when a dragon speaks to a user, it responds appropriately based on the user's emotions. The input is the generated story, and the output is the character's response.
[1077] Step 6:
[1078] The server sends the generated story and character responses to the device, which displays them on the smartphone screen and plays them audibly. The input is the character responses and the generated story, and the output is an interactive story experience that is displayed and played back to the user.
[1079] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1080] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1081] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1082] [Fourth embodiment]
[1083] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1084] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1085] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1086] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1087] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1088] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1089] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1090] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1091] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1092] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1093] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1094] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1095] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1096] This invention is an interactive storytelling system that generates stories in real time based on the user's multimodal input and presents them to the user. This system mainly consists of the following components: a user input processing module, a multimodal data integration module, a generative model module, a real-time response generation module, and an output display module.
[1097] User Input Processing Module
[1098] Users wear a VR headset and interact with the game using voice and gestures, which are then captured by the device through a microphone, camera, and gesture sensors. For example, a user can speak to a dragon by waving their hand or issuing a voice command.
[1099] Multimodal Data Integration Module
[1100] The voice data collected by the device is sent to a server, where it is converted to text using speech recognition technology. Gesture data is also sent to the server, where the user's movements are analyzed using a motion analysis algorithm. These multiple inputs are then combined to create a standardized data set.
[1101] Generative Model Module
[1102] The standardized dataset created above is input into a generative model, which is then run by a server to generate a storyline based on the user's choices and actions. This generative model can be, for example, GPT or another generative AI model, and generates a story in real time in response to the user's actions.
[1103] Real-time response generation module
[1104] Based on the generated storyline, the server generates responses for characters (e.g., dragons) and environments (e.g., the caves in which the dragons live), including the dragon's dialogue, actions, and changes to the surrounding environment. For example, the dragon may start speaking to answer the user's questions.
[1105] Output Display Module
[1106] The data generated by the server is sent to the device, which then displays it in the appropriate format for the user: audio data is played through the speakers, and text, images, and animations are displayed on the screen. For example, a scene in which a dragon moves and talks to the user in a realistic manner is projected on the user's VR headset.
[1107] Specific examples
[1108] When a user issues the voice command "talk to dragon", the following occurs:
[1109] 1. The user says "talk to dragon" aloud.
[1110] 2. The device captures the user's hand-waving gesture along with the audio.
[1111] 3. The device sends the captured data to the server.
[1112] 4. The server converts the speech to text and analyzes the gestures.
[1113] 5. The integrated dataset is fed into a generative model to generate a storyline.
[1114] 6. The server generates the dragon's response (e.g., "Hello, adventurer. What do you want?").
[1115] 7. The server sends the generated response to the terminal, which presents it to the user in real time.
[1116] This system allows users to have a highly personalized and immersive gaming experience, while also allowing game developers to omit complex scenario design, improving development flexibility and creativity.
[1117] The processing flow will be explained below.
[1118] Step 1:
[1119] The user puts on the VR headset and says, "Talk to the dragon." The device captures this voice through the microphone. At the same time, the device captures the user's hand gestures through the camera and sensors.
[1120] Step 2:
[1121] The device sends the captured voice data to the server, and also sends the captured gesture data to the server along with the voice data.
[1122] Step 3:
[1123] The server converts the received voice data into text using voice recognition technology. The voice recognition engine analyzes the voice waveform and generates the corresponding text data.
[1124] Step 4:
[1125] The server analyzes the gesture data and uses a motion analysis algorithm to identify the user's hand gestures and analyze their meaning. For example, it recognizes the "waving" gesture.
[1126] Step 5:
[1127] The server merges the speech and gesture data. The merger process combines the speech text and gesture data into a single standardized data set.
[1128] Step 6:
[1129] The server inputs the combined dataset into a generative model (e.g., GPT-3), which then runs the model and generates a storyline based on the user's input in real time.
[1130] Step 7:
[1131] The server analyzes the generated storyline and generates the character (dragon)'s reactions. The generation process includes specific actions such as the dragon nodding and talking.
[1132] Step 8:
[1133] The server sends generated data, including the dragon's reactions, to the device. This data includes text, audio, and animation data.
[1134] Step 9:
[1135] The device processes the received data and presents it to the user in real time, playing the dragon's voice through the speaker and showing the dragon's movements on the display.
[1136] Step 10:
[1137] The device presents the user with new options, such as "Ask the dragon" or "Fight the dragon."
[1138] Step 11:
[1139] The user selects their next action. Once a new selection is made, the device again sends this information to the server, and the process begins again at step 1.
[1140] Through this series of steps, the system dynamically generates a story based on user input, providing an interactive and immersive gaming experience.
[1141] Example 1
[1142] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1143] Conventional interactive entertainment systems have struggled to respond to user input in real time or generate complex storylines. Furthermore, they lacked technology to integrate and standardize multiple input modes (voice, gestures, etc.), making it difficult to provide an immersive experience that takes advantage of diverse user interactions. Therefore, there is a need for systems that allow users to enjoy more advanced interactive experiences and increase the flexibility and creativity of game developers.
[1144] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1145] In this invention, the server includes means for acquiring multimodal input from a user, means for integrating and standardizing the multimodal input, means for inputting the standardized data into a generative model and generating a story in real time, means for generating character and environmental responses in real time based on the generated story, means for presenting the responses to the user, means for the user to wear a head-mounted display and interact with the user through voice and gestures, means for converting voice data into text and analyzing gesture data, means for generating a storyline based on the user's actions using an artificial intelligence model as a generative model, and means for outputting the generated character's actions and environmental changes in real time, thereby enabling the user to have a highly personalized and immersive interactive experience.
[1146] "Multimodal input" refers to data input by users in multiple different modes, such as voice, gestures, images, and video.
[1147] A "head-mounted display" is a device worn by the user that displays visual information directly in front of the user's eyes.
[1148] A "generative model" refers to an artificial intelligence model that generates new content or stories based on user actions and input.
[1149] A "voice recognition module" is a piece of software that converts voice data into text data.
[1150] A "gesture sensor" is a device that detects a user's movements and captures their action data.
[1151] "Motion analysis algorithm" refers to an algorithm that analyzes captured gesture data to identify specific user movements.
[1152] An "integrated dataset" is a set of data that has been integrated and standardized from data input from multiple different modes.
[1153] "Real-time response generation" refers to the process of generating instant character and environmental reactions based on user input.
[1154] The "output display module" is a module for presenting data generated by the server to the user in an appropriate format.
[1155] The present invention relates to an interactive entertainment system that generates a story in real time based on a user's multimodal input and presents it to the user. The system consists of a user input processing module, a multimodal data integration module, a generative model module, a real-time response generation module, and an output display module.
[1156] User Input Processing Module
[1157] The user wears a head-mounted display and interacts with it using voice and gestures. The device captures these voices and gestures using a microphone, camera, and gesture sensor. For example, if a user says "talk to the dragon" and waves their hand, this data is captured.
[1158] Multimodal Data Integration Module
[1159] The voice data captured by the device is sent to a server where it is converted into text using speech recognition technology. Similarly, gesture data is sent to a server where a motion analysis algorithm analyzes the user's specific movements. This process standardizes the voice and gesture data and compiles them into a unified data set.
[1160] Generative Model Module
[1161] The server inputs the combined dataset into a generative model, such as an advanced generative AI model like GPT-3, which then generates a storyline in real time based on the user's actions.
[1162] Real-time response generation module
[1163] Based on the generated storyline, the server generates responses from characters and the environment in real time. For example, a scene is generated in which a dragon responds to a user's speech by saying, "Hello, adventurer. What do you want?" Changes to the dragon's movements and the surrounding environment (for example, the dragon moving forward or the lights in a cave turning on) are also generated in real time.
[1164] Output Display Module
[1165] The server sends the generated voice, text, image, and animation data to the device. The device then displays the data on the user's head-mounted display, and the audio is played through the speakers. For example, a scene in which a dragon moves and talks to you in a realistic way, or a scene in which lights in a cave light up, may be displayed on the user's head-mounted display.
[1166] Specific examples
[1167] When a user says "talk to dragon" and waves their hand, the following happens:
[1168] 1. The user says "talk to the dragon" and waves their right hand.
[1169] 2. The device captures these voices and gestures and sends them to the server.
[1170] 3. The server converts the voice data into text and analyzes the gesture data.
[1171] 4. The server inputs the integrated dataset into the generative model and generates a storyline.
[1172] 5. The server generates the dragon's responses based on the generated storyline and generates the dragon's movements in real time.
[1173] 6. The server sends the generated response and action to the terminal, which presents it to the user.
[1174] The system allows users to have a highly personalized, immersive, and interactive experience, and an example prompt using the generative AI model is "Generate a scene where the user talks to the dragon and asks what to do next."
[1175] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1176] Step 1: User Input Capture
[1177] The user wears a head-mounted display and interacts with it using voice and gestures. Specifically, the user says "talk to the dragon" and waves their right hand at the same time. The device captures voice with a microphone and gestures with a camera and gesture sensor. Voice data and gesture data are acquired as input. The output is the captured voice data and gesture data.
[1178] Step 2: Sending data
[1179] The device sends the captured voice and gesture data to the server. Specifically, the data is transferred to the server via the Internet in real time. The input is the captured voice and gesture data. The output is the data sent to the server.
[1180] Step 3: Speech and gesture analysis
[1181] The server converts the received voice data into text data using a voice recognition module. For example, the text "talk to the dragon" is acquired. At the same time, the server analyzes the gesture data using a motion analysis algorithm to identify the user's specific actions. The input is voice data and gesture data, and the output is text data and analyzed motion data.
[1182] Step 4: Data integration and standardization
[1183] The server integrates the text data from speech recognition and the gesture analysis data to create a standardized dataset. This is to unify data from different modes. Specifically, the server standardizes the data format and generates the integrated dataset. The input is text data and analyzed gesture data, and the output is the integrated dataset.
[1184] Step 5: Story Generation
[1185] The server inputs the standardized dataset into a generative model. GPT-3 or other models are used as generative models, generating a storyline based on user behavior. Specifically, the server runs the generative model and generates the next development in the story. The input is the integrated dataset, and the output is the generated storyline.
[1186] Step 6: Real-time response generation
[1187] The server determines the response of the character (for example, a dragon) based on the generated storyline. Environmental changes (for example, turning on the lights in a cave) are also generated. As specific actions, the server generates the dragon's lines and actions. For example, a response such as "Hello, adventurer. What do you want?" is generated. The input is the generated storyline, and the output is the character's response and environmental change data.
[1188] Step 7: Sending Responses and Actions
[1189] The server sends the generated response and action data to the terminal. Specifically, the server sends data to the terminal in real time via the Internet. The input is the character's response and environmental change data, and the output is the data sent to the terminal.
[1190] Step 8: Presenting the output
[1191] The device receives audio, text, images, and animation data from the server and presents it to the user. Specific operations include displaying a scene of a dragon speaking to the user on the head-mounted display and lighting up the cave, and playing audio through the speaker. The input is data sent from the server, and the output is the visuals and audio presented to the user.
[1192] (Application example 1)
[1193] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1194] In conventional storytelling systems, user interactions are fixed, making it difficult to generate stories in real time based on individual actions and utterances. Furthermore, the limited means of interaction mean that users lack a sense of immersion. Furthermore, the inability to integrate and utilize multi-modal data, such as gestures and voice, limits the user experience.
[1195] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1196] In this invention, the server includes: means for acquiring multimodal input from a user; means for integrating and standardizing the multimodal input; means for inputting the standardized data into a generative model to generate a story in real time; means for generating character and environmental responses in real time based on the generated story; means for presenting the responses to the user through a smartphone or tablet display; means for capturing and recognizing user gestures with a camera; means for combining the gestures and voice input to form story generation prompts; and means for playing the generated story aloud. This enables rich interactions that combine the user's voice and gestures, and makes it possible to provide a personalized story experience according to individual actions and utterances.
[1197] "Means for obtaining multimodal input from the user" refers to a function for capturing data in multiple forms, such as text, audio, images, and video, generated by the user.
[1198] "Means to integrate and standardize multimodal input" refers to a function that converts acquired data in multiple formats into a single unified dataset for easier processing.
[1199] "Means of inputting data into a generative model and generating a story in real time" refers to a function that allows standardized data to be input into a generative model (e.g., a generative AI model) to create a dynamic story on the spot.
[1200] "Means for generating character and environmental responses in real time" is a function that allows the environment, such as characters and backgrounds, to react in real time based on the generated story, creating appropriate actions and situations.
[1201] "Means of presenting to the user" refers to the function for showing or listening to the generated story or response to the user through a smartphone or tablet display, speaker, etc.
[1202] "Means for capturing and recognizing user gestures with a camera" refers to a function that uses a camera to capture the user's physical movements (e.g., hand gestures) and analyzes and understands them.
[1203] "Means for combining gestures and voice input to form story generation prompts" is a function for generating instructions (prompts) to be given to a generative AI model based on data that combines the user's gestures and voice input.
[1204] "Means for playing the generated story aloud" is a function for converting the story created by the generative model into audio and playing it through a speaker.
[1205] To implement this invention, it is necessary to acquire user voice and gesture input, integrate the data, and input it into a generative AI model to generate a story in real time. A specific embodiment of this is shown below.
[1206] Hardware and software used
[1207] 1. Hardware
[1208] Smartphone or tablet: Used to provide the user interface.
[1209] Camera: Used to capture user gestures.
[1210] Microphone: Used to collect the user's voice.
[1211] Speaker: Used to play the generated audio back to the user.
[1212] 2. Software
[1213] Google Cloud Speech-to-Text API: Used to convert user speech into text.
[1214] OpenCV: Used to recognize user gestures.
[1215] OpenAI GPT (Generative AI Model): Used to generate stories in real time based on user input.
[1216] gTTS (Google Text-to-Speech) and playsound: Used to play the audio of the generated story.
[1217] System Operation Overview
[1218] 1. Acquiring Multimodal Input
[1219] The user speaks into the smartphone or tablet and simultaneously gestures in front of the camera.
[1220] Smartphones capture audio through their microphones and gestures through their cameras.
[1221] 2. Data integration and standardization
[1222] The captured audio data is converted to text using the Google Cloud Speech-to-Text API.
[1223] Similarly, the captured gesture data is analyzed using OpenCV to create a standardized dataset.
[1224] 3. Story Generation
[1225] The combined dataset is fed into a generative AI model (OpenAI GPT) to generate a story in real time.
[1226] Generative AI models create appropriate character and environmental responses based on user actions and speech.
[1227] 4. Real-time response generation and presentation
[1228] The generated story is converted into audio data using gTTS and played back to the user through the smartphone's speakers.
[1229] At the same time, the display shows related images and text.
[1230] Specific examples
[1231] For example, if a user issues the voice command "talk to the wizard" and makes a hand waving gesture, the system provides the following prompt sentence to the generative AI model:
[1232] text
[1233] User Voice: Talk to the Wizard
[1234] User gesture: wave
[1235] Generate the following story:
[1236] Based on this prompt, the generative AI model generates a story like this:
[1237] text
[1238] "Hello, traveller. My name is Eldritch, how may I help you?" the wizard replied in a kind voice.
[1239] This generated story is then converted into audio by gTTS and played back to the user through the smartphone's speakers, with related images and text also appearing on the display, creating a more immersive experience for the user.
[1240] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1241] Step 1: The user speaks a command (e.g., "Talk to the wizard") into a smartphone or tablet. The smartphone captures this voice through its microphone and saves it as audio data. At the same time, the camera captures the user's hand-waving gesture and saves it as video data. This provides audio and video data as input.
[1242] Step 2: The device converts the captured voice data to text data using the Google Cloud Speech-to-Text API. When converting voice data to text data, the input is the voice data, and the output is the converted text data. For example, the utterance "Talk to the wizard" becomes the text "Talk to the wizard."
[1243] Step 3: The device analyzes the captured video data using OpenCV and recognizes the user's gesture (e.g., waving). The input is the video data, and the output is the recognized gesture information. For example, the user's hand wave is recognized as a "wave."
[1244] Step 4: The device combines the results of converting the voice data to text and the recognized gesture data to generate a standardized dataset. The input is text data and gesture data, and the output is the combined dataset. For example, the dataset might have the format "Voice: Talk to the Wizard" and "Gesture: Wave."
[1245] Step 5: The server receives the integrated dataset and inputs it into the generative AI model (OpenAI GPT). The input is the integrated dataset, and the output is a story text based on the generative AI model. For example, based on the prompts "User voice: Talk to the wizard" and "User gesture: Wave", the story "Hello, traveler. My name is Eldritch. How may I help you?" is generated.
[1246] Step 6: The server converts the generated story text into audio data using gTTS (Google Text-to-Speech). The input is the generated story text, and the output is the audio data of the story. This audio data is sent to the smartphone.
[1247] Step 7: The device plays the received audio data to the user through the speaker. The input is audio data, and the output is audio playback. At the same time, related text and images are displayed on the screen, providing the user with an immersive experience.
[1248] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1249] The present invention is an interactive storytelling system that generates a story in real time based on the user's multimodal input and presents it to the user, and further incorporates an emotion engine to recognize the user's emotions and adjust the story and character responses based on those emotions. The system of the present invention mainly consists of the following components: a user input processing module, an emotion engine, a multimodal data integration module, a generative model module, a real-time response generation module, and an output display module.
[1250] User Input Processing Module
[1251] Users wear a VR headset and interact with the game using voice and gestures. The device captures these interactions through a microphone, camera, and gesture sensors. For example, a user can speak to a dragon by issuing a voice command or waving their hand.
[1252] Emotion Engine
[1253] The server operates an emotion engine using voice data and image data acquired from the user. The emotion engine recognizes the user's emotions using voice analysis and detects and identifies the user's facial expressions using image analysis. This allows the engine to determine the user's current emotion (e.g., joy, anger, sadness) from the tone and facial expression of the user when speaking.
[1254] Multimodal Data Integration Module
[1255] The voice and gesture data collected by the device is sent to a server, where it is converted into text using speech recognition technology. Gesture data is also sent to the server, where the user's movements are analyzed using a motion analysis algorithm. These multiple input data are integrated to create a standardized dataset, which also includes emotional information recognized by an emotion engine.
[1256] Generative Model Module
[1257] The standardized dataset created above is input into a generative model. The server runs the generative model (e.g., GPT-3) to generate a storyline based on the user's input in real time. The generative model also uses the user's emotional information to adjust the story and generate a more convincing response.
[1258] Real-time response generation module
[1259] Based on the generated storyline, the server generates responses for characters (e.g., dragons) and environments (e.g., the caves where dragons live). The generation process includes specific actions such as the dragon nodding and talking. Reactions based on the user's emotions are also incorporated. For example, if the user is angry, a scenario will be generated in which the dragon speaks words to calm them down.
[1260] Output Display Module
[1261] The data generated by the server is sent to the device, which then displays it in an appropriate format for the user. Audio data is played through the speaker, and text, images, and animations are displayed on the screen. For example, a dragon may talk to the user while moving realistically, presenting new options to the user (e.g., "Ask the dragon" or "Fight the dragon").
[1262] Specific examples
[1263] Consider a scenario where a user issues the voice command "talk to the dragon" and simultaneously smiles and waves:
[1264] 1. The user says "talk to the dragon" and smiles and waves.
[1265] 2. The device captures voice, gesture, and facial expression data and sends it to the server.
[1266] 3. The server converts the speech into text and analyzes gestures and facial expressions.
[1267] 4. The emotion engine recognizes the user's smile and determines it to be a positive emotion (joy).
[1268] 5. The integrated dataset is fed into a generative model to generate a storyline.
[1269] 6. The server generates a dragon response based on the positive emotion (e.g., "You're in a good mood today, adventurer. How can I help you?").
[1270] 7. The server sends the generated response to the terminal, which presents it to the user in real time.
[1271] This system allows users to have a highly personalized, emotionally immersive gaming experience, while also allowing game developers to eliminate the need for complex scenario design, improving development flexibility and creativity.
[1272] The processing flow will be explained below.
[1273] Step 1:
[1274] The user puts on the VR headset and says "talk to the dragon." The device captures this voice through the microphone. At the same time, the device captures the user's hand gesture through the camera and sensors. The device also captures the user's facial expressions (e.g., smiling).
[1275] Step 2:
[1276] The terminal transmits the captured voice data, gesture data, and facial expression data to a server.
[1277] Step 3:
[1278] The server converts the received voice data into text using voice recognition technology. The voice recognition engine analyzes the voice waveform and generates the corresponding text data.
[1279] Step 4:
[1280] The server analyzes the gesture data and uses a motion analysis algorithm to identify the user's hand gestures and analyze their meaning. For example, it recognizes the "waving" gesture.
[1281] Step 5:
[1282] The server analyzes the facial expression data. It uses an expression analysis algorithm to recognize the user's facial expressions and identify their emotional state (e.g., joy, anger, sadness, or happiness). For example, it recognizes the emotion of "joy" from the user's smile.
[1283] Step 6:
[1284] The server integrates the speech, gesture, and facial expression data. The integration process brings together speech text, gesture data, and emotion information into a single standardized dataset.
[1285] Step 7:
[1286] The server inputs the combined dataset into a generative model (e.g., GPT-3), which then runs the model and generates a storyline based on the user's input in real time.
[1287] Step 8:
[1288] The server analyzes the generated storyline and generates a character (e.g., a dragon)'s reaction based on the user's emotional information. For example, if the user has the emotion "joy," the server generates a scenario in which the dragon responds in a friendly manner.
[1289] Step 9:
[1290] The server sends generated data, including character reactions, to the device. This data includes text, audio, and animation data.
[1291] Step 10:
[1292] The device processes the received data and presents it to the user in real time, playing the dragon's voice through the speaker and showing the dragon's movements on the display.
[1293] Step 11:
[1294] The device presents the user with new options, such as "Ask the dragon" or "Fight the dragon."
[1295] Step 12:
[1296] The user selects their next action. Once a new selection is made, the device again sends this information to the server, and the process begins again at step 1.
[1297] This series of steps enables the system to dynamically generate a story while taking into account the user's emotions, providing an interactive and immersive gaming experience.
[1298] Example 2
[1299] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1300] Previous interactive storytelling systems lacked the ability to recognize users' emotions in real time and adjust the story and character responses based on those emotions. As a result, users often did not receive responses that reflected their own emotions and actions, resulting in a less immersive experience. Furthermore, there was a lack of technical means to provide flexible and advanced interactions.
[1301] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1302] In this invention, the server includes means for acquiring multimodal input from a user, means for integrating and standardizing the multimodal input, means for inputting the standardized data into a generative model to generate a story in real time, means for recognizing the user's emotions, means for reflecting the recognized emotional information in a real-time response, and means for presenting the response to the user. This enables real-time responses that are in line with the user's emotions and behavior, providing a more immersive interactive experience.
[1303] "Multimodal input" refers to input from a user that includes data in multiple formats, such as text, audio, images, and video.
[1304] "Means of integration and standardization" refers to the process of converting data in multiple different formats into a consistent format and compiling it into a single data set.
[1305] A "generative model" refers to an artificial intelligence model that generates new text or storylines based on given input data.
[1306] A "means for generating stories in real time" is a means for instantly responding to input from the user and generating new stories on the spot.
[1307] "Means for generating character and environmental responses in real time" refers to means for instantly generating character actions and environmental changes based on the generated story.
[1308] "Means for recognizing user emotions" refers to means for identifying a user's emotional state in real time using voice analysis or image analysis.
[1309] "Means for reflecting emotional information in real-time responses" refers to means for adjusting the responses and stories generated based on the recognized emotions of the user.
[1310] "Means for presenting responses to users" refers to the means by which the generated responses or stories are conveyed to users in the form of audio, text, animation, etc.
[1311] This invention is an interactive storytelling system that generates a story in real time based on the user's multimodal input and presents it to the user. It also incorporates an emotion engine, which recognizes the user's emotions and adjusts the story and character responses based on those emotions. This system mainly consists of the following components: a user input processing module, an emotion engine, a multimodal data integration module, a generative model module, a real-time response generation module, and an output display module.
[1312] User Input Processing Module
[1313] Users wear a VR headset and interact with the game using voice and gestures. The device captures these interactions through a microphone, camera, and gesture sensor. For example, a user can issue a voice command to "talk to the dragon" and wave their hand.
[1314] Emotion Engine
[1315] The server runs an emotion engine using voice and image data acquired from the user. The emotion engine recognizes the user's emotions using voice analysis and detects and identifies the user's facial expressions using image analysis. This allows the engine to determine the user's current emotion (e.g., joy, anger, sadness) from the tone and facial expression of the user when speaking.
[1316] Multimodal Data Integration Module
[1317] The device sends collected voice and gesture data to a server, which converts the data into text using speech recognition technology. Gesture data is also sent to the server, where the user's movements are analyzed using a motion analysis algorithm. These multiple inputs are combined to create a standardized dataset, which also includes emotional information recognized by an emotion engine.
[1318] Generative Model Module
[1319] The standardized dataset is then fed into a generative model. The server runs the generative model (e.g., a large-scale language model) to generate a storyline based on the user's input in real time. The generative model also uses the user's emotional information to adjust the story and generate a more convincing response.
[1320] Real-time response generation module
[1321] Based on the generated storyline, the server generates responses for characters (e.g., dragons) and environments (e.g., the caves where the dragons live). The generation process includes specific actions such as the dragon nodding and talking. Reactions based on the user's emotions can also be incorporated. For example, if the user is angry, a scenario will be generated in which the dragon speaks words to calm them down.
[1322] Output Display Module
[1323] The data generated by the server is sent to the device, which then displays it in an appropriate format for the user. Audio data is played through the speaker, and text, images, and animations are displayed on the screen. For example, a dragon moves realistically and speaks to the user, presenting new options to the user.
[1324] Specific examples
[1325] Consider a scenario where the user gives the voice command "talk to the dragon" and smiles and waves:
[1326] 1. The user says "talk to the dragon" and smiles and waves.
[1327] 2. The device captures voice, gesture, and facial expression data and sends it to the server.
[1328] 3. The server converts the speech into text and analyzes gestures and facial expressions.
[1329] 4. The emotion engine recognizes the user's smile and determines it to be a positive emotion (joy).
[1330] 5. The integrated dataset is fed into a generative model to generate a storyline.
[1331] 6. The server generates a dragon response based on the positive emotion (e.g., "You're in a good mood today, adventurer. How can I help you?").
[1332] 7. The server sends the generated response to the terminal, which presents it to the user in real time.
[1333] Prompt Sentence Examples
[1334] "The user smiles and says the voice command 'talk to the dragon.' Generate a response from the dragon."
[1335] "If the user is angry, write a line for the dragon that will calm them down."
[1336] This system allows users to have a highly personalized, emotionally immersive experience, while also allowing game developers to eliminate the need for complex scenario design, improving development flexibility and creativity.
[1337] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1338] Step 1:
[1339] The user puts on the VR headset and begins interacting with it using voice and gestures. For example, they might say, "Talk to the dragon," and then wave their hand. The device's microphone, camera, and gesture sensors capture this voice, image, and movement data.
[1340] Input: Audio, video, gesture data
[1341] Output: Raw captured data
[1342] Step 2:
[1343] The voice data and gesture data captured by the device are sent to the server via the network, along with the voice data captured by the microphone and the gesture data captured by the camera and gesture sensor.
[1344] Input: Raw captured data
[1345] Output: Voice and gesture data sent to the server
[1346] Step 3:
[1347] The server uses voice recognition technology (e.g., a voice-to-text conversion system) to convert the captured voice data into text, and uses a motion analysis algorithm (e.g., a motion detection algorithm) to analyze the gesture data and recognize the user's movements.
[1348] Input: Voice data, gesture data
[1349] Data processing: Converts voice data into text, and gesture data into motion recognition information
[1350] Output: Text data, motion recognition data
[1351] Step 4:
[1352] The server runs an emotion engine (e.g., a voice emotion analysis system, an expression analysis system) to analyze the voice tone and facial expressions, thereby determining whether the user is expressing an emotion such as joy or anger.
[1353] Input: Audio data, video data
[1354] Data processing: voice tone analysis, facial expression analysis
[1355] Output: Emotion recognition data
[1356] Step 5:
[1357] The server combines the text data, action recognition data, and emotion recognition data to create a standardized dataset, which is then input into the subsequent generative model.
[1358] Input: Text data, action recognition data, emotion recognition data
[1359] Data processing: data integration and standardization
[1360] Output: Standardized dataset
[1361] Step 6:
[1362] The server runs a generative AI model (e.g., a large-scale language model) to generate storylines in real time based on the integrated dataset, for example, generating positive stories based on what makes the user happy.
[1363] Input: Standardized dataset
[1364] Data processing: Story generation using generative AI models
[1365] Output: Generated storyline
[1366] Step 7:
[1367] The server generates character actions and conversations based on the generated storyline. For example, it generates lines and actions for a dragon to speak to the user. The response changes depending on the user's emotions.
[1368] Input: Storyline, emotion recognition data
[1369] Data processing: Character response generation
[1370] Output: Character movement data, character conversation data
[1371] Step 8:
[1372] The server sends the generated character movement data and conversation data to the device, which then presents this to the user in real time. Audio is played through the speaker and animation is shown on the display. For example, a dragon could move realistically and say, "You're in a good mood today, adventurer. Is there anything I can help you with?"
[1373] Input: Character movement data, character conversation data
[1374] Output: Real-time display to the user
[1375] (Application example 2)
[1376] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1377] Conventional interactive storytelling systems do not take into account the user's emotional state, resulting in a lack of individualized optimization of the user's experience and difficulty in achieving emotional empathy. Furthermore, it is difficult to effectively integrate the user's diverse inputs (voice, facial expressions, gestures), making it difficult to generate appropriate responses in real time. Furthermore, it has been challenging to realize a system that is not dependent on a specific device and uses general-purpose hardware.
[1378] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1379] In this invention, the server includes: means for acquiring multimodal input from a user; means for integrating and standardizing the multimodal input; means for inputting the standardized data into a generative model to generate a story in real time; means for generating responses of characters and the environment in real time based on the generated story; means for presenting the responses to the user; means for analyzing the user's emotions and adjusting the generated story and character responses based on the emotions; means for acquiring the user's voice, facial expressions, and gestures using a smartphone camera and microphone; and means for displaying the story generated in real time on a display device according to the user's emotional state, thereby enabling interactive storytelling that is individually optimized according to the user's emotional state.
[1380] "Means for acquiring multimodal input from a user" refers to a device or system for simultaneously acquiring multiple forms of input, such as a user's voice, facial expressions, and gestures.
[1381] The "means for integrating and standardizing the multimodal input" refers to a device or system that processes multiple acquired input formats in a unified manner and converts them into a format that is easy to analyze.
[1382] "Means for inputting the standardized data into a generative model and generating a story in real time" refers to a device or system that generates a story in real time using an artificial intelligence model based on the integrated and standardized data.
[1383] "Means for generating character and environmental responses in real time based on the generated story" refers to a device or system that generates character and environmental responses in real time according to a story created by a generative model.
[1384] "Means for presenting the response to the user" refers to a device or system that displays or outputs the response of the generated character or environment to the user.
[1385] "Means for analyzing the user's emotions and adjusting the generated story and character responses based on said emotions" refers to a device or system that analyzes the user's emotional state and adjusts the generated story and character responses based on that data.
[1386] "Means for capturing a user's voice, facial expressions, and gestures using a smartphone's camera and microphone" refers to a device or system that captures a user's voice, facial expressions, and gestures using a smartphone's built-in camera and microphone.
[1387] "Means for displaying a story generated in real time on a display device according to the user's emotional state" refers to a device or system for appropriately displaying to a user a story generated in real time based on the user's emotional state.
[1388] This invention uses a system that analyzes various user inputs (voice, facial expressions, gestures) in real time and generates an interactive story based on the analysis. This system mainly uses a smartphone and its built-in camera and microphone.
[1389] Hardware and software used
[1390] 1. Smartphone: Uses a camera and microphone to capture the user's voice, facial expressions, and gestures.
[1391] 2. Server: Data processing, sentiment analysis, story generation, character response generation. The main software used here is as follows:
[1392] OpenCV: A library used for camera capture and image analysis.
[1393] speech_recognition: The library to use for speech recognition.
[1394] transformers (GPT-2): A library used as a generative model.
[1395] TextBlob: A library used for sentiment analysis.
[1396] Processing flow
[1397] 1. User Input Processing:
[1398] The user speaks and gestures into the smartphone's camera, which then captures the audio and image data.
[1399] 2. Multimodal Data Integration:
[1400] The smartphone captures audio and image data and sends it to the server. The audio data is converted to text using the speech_recognition library. The image data is analyzed for gestures and facial expressions using OpenCV.
[1401] 3. Emotion analysis:
[1402] The converted voice data and image analysis results are sent to the server, where the TextBlob library analyzes the user's emotions, classifying them as positive, negative, or neutral.
[1403] 4. Story generation:
[1404] The data combined with the sentiment analysis results is fed into a GPT-2 model from the transformers library to generate an interactive story based on the user's emotional state.
[1405] 5. Character response generation:
[1406] Based on the story created by the generative model, specific responses from the character are generated. For example, when a dragon speaks to the user, it responds appropriately based on the user's emotions.
[1407] 6. Real-time display:
[1408] The generated story and character responses are displayed on the smartphone screen and played back aloud.
[1409] Specific examples
[1410] Consider a scenario where the user says "talk to the dragon" and smiles and waves:
[1411] 1. The smartphone's microphone and camera capture the user's voice, facial expressions, and gestures.
[1412] 2. The captured data is sent to a server, where the voice data is converted into text and the image data is analyzed for gestures and facial expressions.
[1413] 3. The TextBlob library parses the user's smile as a positive emotion.
[1414] 4. The combined data is fed into the GPT-2 model to generate prompts for the user to talk to the dragon in a positive manner.
[1415] 5. An interactive story is generated based on the generated prompts.
[1416] 6. As part of the story, the dragon responds to the user by saying, "You're in a good mood today, adventurer. How can I help you?"
[1417] Prompt Sentence Examples
[1418] An example prompt that describes a situation in which the user says "talk to the dragon" and smiles and waves is:
[1419] The user says "talk to the dragon," smiles, and waves. The user is pleased. How does the dragon respond?
[1420] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1421] Step 1:
[1422] The user speaks and gestures into the smartphone camera. The smartphone camera and microphone capture audio and image data. In this step, the user's voice input and gesture input are collected. The input data are audio files and image files.
[1423] Step 2:
[1424] The device sends captured voice and image data to the server. The voice data is converted to text using the speech_recognition library. The image data is analyzed for gestures and facial expressions using OpenCV. The input is the captured voice data and image data, and the output is text data and analyzed gesture and facial expression data.
[1425] Step 3:
[1426] The server inputs the converted voice data and image analysis results into the TextBlob library to analyze the user's emotions. The emotion analysis results are classified as positive, negative, or neutral. The input is text data and image analysis data, and the output is the user's emotional state.
[1427] Step 4:
[1428] The integrated data (text data, gesture analysis data, and emotion analysis data) is input to the GPT-2 model in the transformers library on the server to generate an interactive story based on the user's emotional state. The input is the integrated data, and the output is the generated story.
[1429] Step 5:
[1430] The server generates specific responses for the characters based on the generated story. For example, when a dragon speaks to a user, it responds appropriately based on the user's emotions. The input is the generated story, and the output is the character's response.
[1431] Step 6:
[1432] The server sends the generated story and character responses to the device, which displays them on the smartphone screen and plays them audibly. The input is the character responses and the generated story, and the output is an interactive story experience that is displayed and played back to the user.
[1433] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1434] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1435] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1436] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1437] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1438] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1439] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1440] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1441] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1442] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1443] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1444] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1445] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1446] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1447] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1448] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1449] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1450] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1451] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1452] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1453] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1454] The following is further disclosed regarding the above embodiment.
[1455] (Claim 1)
[1456] a means for obtaining multimodal input from a user;
[1457] means for integrating and standardizing said multimodal input;
[1458] means for inputting the standardized data into a generative model to generate a story in real time;
[1459] means for generating character and environmental responses in real time based on the generated story;
[1460] means for presenting said response to a user.
[1461] (Claim 2)
[1462] 2. The system of claim 1, wherein the multimodal input is at least one of text, audio, image, and video.
[1463] (Claim 3)
[1464] The system of claim 1 , wherein the generative model is an artificial intelligence model.
[1465] "Example 1"
[1466] (Claim 1)
[1467] a means for obtaining multimodal input from a user;
[1468] means for integrating and standardizing said multimodal input;
[1469] means for inputting the standardized data into a generative model to generate a story in real time;
[1470] means for generating character and environmental responses in real time based on the generated story;
[1471] means for presenting said response to a user;
[1472] A means for users to wear a head-mounted display and interact with it using voice and gestures;
[1473] means for converting voice data to text and analyzing gesture data;
[1474] a means for generating a storyline based on user behavior using an artificial intelligence model as a generative model;
[1475] A system that includes a means for outputting the generated character's movements and environmental changes in real time.
[1476] (Claim 2)
[1477] 2. The system of claim 1, wherein the multimodal input is at least one of text, audio, image, and video.
[1478] (Claim 3)
[1479] The system of claim 1, wherein the generative model is a generative AI model that generates a story using prompts based on user behavior.
[1480] "Application Example 1"
[1481] (Claim 1)
[1482] a means for obtaining multimodal input from a user;
[1483] means for integrating and standardizing said multimodal input;
[1484] means for inputting the standardized data into a generative model to generate a story in real time;
[1485] means for generating character and environmental responses in real time based on the generated story;
[1486] means for presenting said response to the user via a display of a smartphone or tablet;
[1487] A means for capturing and recognizing user gestures with a camera;
[1488] means for combining the gestures and speech input to form story generation prompts;
[1489] The system includes means for audibly playing the generated story.
[1490] (Claim 2)
[1491] 2. The system of claim 1, wherein the multimodal input is at least one of text, audio, image, and video.
[1492] (Claim 3)
[1493] The system of claim 1, wherein the generative model is an artificial intelligence model that generates a story based on a generative prompt sentence.
[1494] "Example 2: Combining Emotion Engines"
[1495] (Claim 1)
[1496] a means for obtaining multimodal input from a user;
[1497] means for integrating and standardizing said multimodal input;
[1498] means for inputting the standardized data into a generative model to generate a story in real time;
[1499] means for generating character and environmental responses in real time based on the generated story;
[1500] a means of recognizing a user's emotions;
[1501] means for reflecting the recognized emotion information in a real-time response;
[1502] means for presenting said response to a user.
[1503] (Claim 2)
[1504] 2. The system of claim 1, wherein the multimodal input is at least one of text, audio, image, and video.
[1505] (Claim 3)
[1506] The system of claim 1 , wherein the generative model is an artificial intelligence model.
[1507] "Application example 2 when combining emotion engines"
[1508] (Claim 1)
[1509] a means for obtaining multimodal input from a user;
[1510] means for integrating and standardizing said multimodal input;
[1511] means for inputting the standardized data into a generative model to generate a story in real time;
[1512] means for generating character and environmental responses in real time based on the generated story;
[1513] means for presenting said response to a user;
[1514] means for analyzing a user's emotions and adjusting the generated story and character responses based on said emotions;
[1515] A means for capturing a user's voice, facial expressions, and gestures using a camera and microphone of the smartphone;
[1516] A system including means for displaying a real-time generated story on a display device in response to a user's emotional state.
[1517] (Claim 2)
[1518] 10. The system of claim 1, wherein the multimodal input is at least one of text, audio, image, and video.
[1519] (Claim 3)
[1520] The system of claim 1 , wherein the generative model is an artificial intelligence model. [Explanation of symbols]
[1521] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for obtaining multimodal input from a user; means for integrating and standardizing said multimodal input; means for inputting the standardized data into a generative model to generate a story in real time; means for generating character and environmental responses in real time based on the generated story; means for presenting said response to a user.
2. The system of claim 1 , wherein the multimodal input is at least one of text, audio, images, and video.
3. The system of claim 1 , wherein the generative model is an artificial intelligence model.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A