System

The system addresses the challenge of providing a creative and engaging storytelling experience for children by using generation AI for natural language processing, enabling flexible responses and real-time interaction with robots, thus enhancing creativity and involvement.

JP2025072326APending Publication Date: 2025-05-09SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024184315
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-24
Filing Date
2024-10-18
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Existing systems struggle to provide a fun and creative experience for children in story creation and dialogue, as they require flexible and diverse responses to user suggestions and questions, and fail to instantly reflect story development and character responses.

Method used

A system that uses natural language processing with generation AI to receive story suggestions and questions from users, generate flexible and diverse responses, and display robot actions and reactions in real time, enhancing user involvement and creativity.

Benefits of technology

The system effectively increases children's creativity and sense of involvement by providing dynamic and natural narrative development and character responses, creating an immersive and interactive storytelling experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025072326000001_ABST
    Figure 2025072326000001_ABST
Patent Text Reader

Abstract

To provide a system.SOLUTION: A system comprises: means for using an emotion engine which recognizes a user's emotion on the basis of the user's voice and image, to recognize the user's emotion; means for receiving a suggestion of or a question about a story from the user; means for generating an evolution of a story and a character's response through generative AI, as well as a prompt for ordering generation of the evolution of a story and a character's response on the basis of the suggestion of and the question about a story that are received, and the user's emotion; and means for controlling a robot's behavior on the basis of the generated evolution of the story or the character's response.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including a description and related instruction sentence regarding the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2022-180282 A Summary of the Invention [Problem to be solved by the invention]

[0004] In order to provide a fun and creative experience with children through story creation and dialogue, there are several challenges:

[0005] In order for children to freely create their own stories, they need flexible and diverse responses to suggestions and questions about characters and plot.

[0006] · The story development and characters’ reactions should be immediate to children’s suggestions and questions, enhancing children’s creativity and sense of involvement. [Means for solving the problem]

[0007] In order to solve the above problems, the following measures are proposed.

[0008] It receives story suggestions and questions from users and generates story developments and character responses based on them, achieving flexible and diverse responses.

[0009] Generative AI is used to perform natural language processing in response to suggestions and questions, and the robot then expresses the generated developments and reactions as actions, thereby increasing children's sense of engagement.

[0010] By displaying the robot's actions and reactions on the device, children can react to the story's development in real time, encouraging creative dialogue.

[0011] These tools allow you to solve story-making challenges together with your child, enhancing their creativity and sense of engagement.

[0012] "Robot application system" refers to programs and applications developed for story creation and dialogue.

[0013] "Narrative development" refers to the progression of a story or the development of a storyline. It means that the story progresses and develops through changes and developments in the characters and plot.

[0014] "Character reaction" refers to the emotions and behaviors that a character in a story shows in response to the development of the story or a user's suggestions or questions. It includes how the character reacts to events that occur in the story.

[0015] "Generative AI" refers to models or systems that use natural language processing techniques to conduct dialogue. Generative AI is an AI model that can generate natural responses based on user input. [Brief description of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Diagram 2]1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. FIG. [Diagram 3] FIG. 11 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Diagram 5] FIG. 13 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 13 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 13 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] 4 is a sequence diagram showing a process flow of the data processing system according to the first embodiment. FIG. [Figure 12] 11 is a sequence diagram showing a process flow of the data processing system in application example 1. FIG. [Figure 13] FIG. 11 is a sequence diagram showing the flow of processing of the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 11 is a sequence diagram showing the flow of processing in the data processing system in application example 2 when combined with an emotion engine. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a signed processor (hereinafter simply referred to as a "processor") may be one arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be one type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), etc.

[0020] In the following embodiments, a signed RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by the processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a code is an interface including a communication processor and an antenna. The communication I / F controls communication between multiple computers. An example of a communication standard applied to the communication I / F is a wireless communication standard including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. In addition, in this specification, the same idea as "A and / or B" is also applied when three or more things are expressed by connecting them with "and / or."

[0024] [First embodiment]

[0025] Fig. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment. As shown in Fig. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).

[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0028] The reception device 38 includes a touch panel 38A and a microphone 38B, and receives user input. The touch panel 38A detects contact with a pointer (e.g., a pen or a finger) to receive user input by the touch of the pointer. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (e.g., voice and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs voice according to instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a Complementary Metal-Oxide-Semiconductor (CMOS) image sensor or a Charge Coupled Device (CCD) image sensor.

[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54.

[0031] FIG. 2 shows an example of main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Fig. 2, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32. The specific process program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific process program 56 from the storage 32, and executes the read specific process program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific process program 56 executed on the RAM 30.

[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores a reception output program 60. The reception output program 60 is used together with the specific processing program 56 by the data processing system 10. The processor 46 reads out the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0035] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0036] An embodiment for implementing the present invention includes the following elements.

[0037] 1. Server

[0038] A server is provided that executes a program for generating story development and character reactions.

[0039] It receives story suggestions and questions from users, and uses generative AI to generate developments and responses based on them.

[0040] The generated developments and reactions are sent to the robot.

[0041] 2. Terminal

[0042] It provides an interface for users to make story suggestions and questions.

[0043] The user's input information is sent to the server.

[0044] 3. Robot

[0045] Act on developments and reactions sent from the server.

[0046] As a character in the story, it will act and express itself according to the generated developments and reactions.

[0047] The response is displayed on the terminal, enabling a dialogue with the user.

[0048] (Specific examples)

[0049] For example, if a user suggests through a device, "Make a story where the main character goes on a journey into space," the device will send that information to the server. The server will use generative AI based on the suggestion to generate a plot related to space travel. The generated plot will be sent to the robot, which will then act out the journey into space.

[0050] Also, if the user asks, "What would happen if the main character met an alien?", the device sends that information to the server. The server uses generative AI to generate the character's reaction to meeting an alien. The robot acts based on that reaction and expresses its reaction to meeting an alien.

[0051] The above is an example of an embodiment of the present invention. A system for creating a story and realizing a dialogue is specifically embodied by the roles and cooperation of a server, a terminal, and a robot.

[0052] The process flow will be explained below.

[0053] Step 1: Receiving User Input

[0054] The terminal provides an interface for the user to make story suggestions and questions, and receives information entered by the user.

[0055] Step 2: Submit your input

[0056] The device sends the received user input information, including suggestions and questions, to the server.

[0057] Step 3: Receiving and processing information

[0058] The server receives the information sent from the device and uses generative AI to generate the story development and character reactions based on the received information.

[0059] Step 4: Development and reaction generation

[0060] The server uses generative AI to generate story developments and character reactions based on the information it receives, which are then sent to the robot in the next step.

[0061] Step 5: Robot behavior

[0062] The robot receives the developments and reactions sent from the server. Based on the information it receives, it acts as a character in the story. The specific actions it takes vary depending on the development of the story.

[0063] Step 6: Viewing the Robot's Response

[0064] The robot displays responses on the device based on its own actions and reactions, including story developments and character reactions.

[0065] The above is the specific flow of the program's processing. The terminal receives the user's input, the server processes the information and generates developments and reactions, the robot acts based on that, and finally the robot's response is displayed on the terminal. This is how the story is created and the dialogue is carried out.

[0066] Example 1

[0067] Next, a description will be given of Example 1. In the following description, the data processing device 12 is referred to as a "server" and the smart device 14 is referred to as a "terminal."

[0068] Conventional story generation systems have difficulty generating natural story developments and character reactions to user input in real time, and there is a lack of systems that can express them as real objects and characters. This limits interactive dialogue with the user, resulting in a poor quality experience.

[0069] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0070] In this invention, the server includes a means for receiving story suggestions and questions from a user, a means for generating a story development and a character's reaction using a generative AI model based on the received suggestions and questions, and a means for the robot to act based on the generated development and reaction. This enables the generation and expression of a dynamic and natural story development and character reaction in response to user input.

[0071] "User" refers to the person or user who inputs story suggestions and questions.

[0072] "Terminal" refers to an electronic device that provides an interface through which a user can enter story suggestions and questions and transmit them to a server.

[0073] The "server" refers to a computer system that receives suggestions and questions sent by users, uses a generative AI model to generate story developments and character responses, and transmits them to the robot.

[0074] A "generative AI model" is a model that uses artificial intelligence technology and is an algorithm that performs natural language processing based on user input to generate story development and character reactions.

[0075] A "prompt sentence" refers to a text sentence entered into a generative AI model to give it a specific instruction or request.

[0076] A "robot" refers to a physical device that acts as a character in a story based on the generation results sent from a server, performs actions and expressions, and displays its responses on a terminal.

[0077] "Story suggestions and questions" refer to instructions or questions that a user inputs via a terminal regarding the content of the story or the actions of a character and that are then transmitted to the server.

[0078] "Development and response" refers to the story progression and character responses that the generative AI model generates based on user suggestions and questions.

[0079] To implement this invention, hardware such as a server, terminal, and robot, as well as software such as a generative AI model, are required. This system generates story development and character reactions based on user input, and expresses them as real objects, thereby realizing interactive dialogue with the user.

[0080] System Configuration

[0081] 1. Server

[0082] Hardware: High-performance computer systems (e.g., a server rack with multiple processors, memory, storage, etc.)

[0083] Software: Generative AI models (e.g. ChatGPT(R) by OpenAI(R)), data processing algorithms, HTTP servers, etc.

[0084] Function: Receives story suggestions and questions sent by the user via the device, and uses a generative AI model to generate story development and character responses based on those suggestions. The generated data is then sent to the robot.

[0085] 2. Terminal

[0086] Hardware: User interface devices such as smartphones, tablets, and computers

[0087] Software: Web applications, mobile applications

[0088] Function: Provides an interface for users to enter story suggestions and questions, and sends the information to a server.

[0089] 3. Robots

[0090] Hardware: Physically shaped robots (e.g. humanoid robots, animal-like robots, etc.), sensors and actuators for movement and voice output

[0091] Software: Robot control program

[0092] Function: Acts and expresses itself based on the story development and character reactions sent from the server, and displays the results on the device.

[0093] How it works

[0094] When a user uses a device to input story suggestions or questions, the input data is sent from the device to a server. The server analyzes the received input data and uses a generative AI model to generate the story development and character reactions. The generated data is sent to the robot, which then operates based on that data. As a result, the robot behaves and speaks as a character in the story, and displays this on the device, realizing an interactive dialogue with the user.

[0095] Examples

[0096] For example, if a user uses a terminal to suggest, "Create a story where the main character goes on a journey into space," the process would go like this:

[0097] User: Type "Create a story where the main character goes on a space journey" into the terminal.

[0098] Terminal: Sends the entered information to the server.

[0099] Server: Sends the prompt "Create a story in which the protagonist goes on a journey into space" to the generative AI model and generates the story development.

[0100] Server: Sends the generated story development to the robot.

[0101] Robot: Performs movements that represent "the protagonist setting off into space."

[0102] Robot: Display the results on the terminal.

[0103] Also, if a user asks, "What happens if the main character meets an alien?"

[0104] User: Type into terminal, "What happens when the main character meets an alien?"

[0105] Terminal: Sends the entered information to the server.

[0106] Server: Sends this prompt to the generative AI model and generates the character's response.

[0107] Server: Sends the generated responses to the robot.

[0108] Robot: Performs actions that express the reaction of the protagonist when he meets the alien.

[0109] Robot: Display the results on the terminal.

[0110] In this manner, the present invention enables the generation and expression of dynamic and natural story development and character reactions to user input.

[0111] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0112] Step 1:

[0113] A user uses the terminal to input story suggestions and questions.

[0114] Specific operations: The user enters "Create a story in which the main character goes on a journey into space" into the input form displayed on the device screen.

[0115] Input: User suggestions or questions ("Create a story in which the main character travels into space")

[0116] Output: Input data that the device sends to the server

[0117] Step 2:

[0118] The terminal transmits the user's input information to the server.

[0119] Specific operation: The terminal sends the entered string data to the server as an HTTP request.

[0120] Input: String data entered by the user

[0121] Output: HTTP request to the server

[0122] Step 3:

[0123] The server receives and analyzes the input data sent from the terminal.

[0124] Specific operation: The server receives an HTTP request and extracts the user input contained in the request body.

[0125] Input: HTTP request from the terminal

[0126] Output: Extraction result of user input data

[0127] Step 4:

[0128] The server sends a prompt to the generative AI model based on the received input data.

[0129] Specific operation: The server generates a prompt such as "Create a story in which the protagonist goes on a journey into space" and sends it to the generative AI model as an API request.

[0130] Input: Data entered by the user

[0131] Output: API request to the generative AI model

[0132] Step 5:

[0133] The generative AI model generates a story development in response to the prompt sentence and returns the result to the server.

[0134] Specific operation: The generative AI model analyzes the prompt sentence, generates a story development, and returns it to the server as an API response.

[0135] Input: API request to the generative AI model (prompt text)

[0136] Output: Generated story development

[0137] Step 6:

[0138] The server receives and analyzes the story development returned by the generative AI model.

[0139] Specific operation: The server receives the API response returned from the generative AI model and parses the contents in JSON format.

[0140] Input: API response from the generative AI model

[0141] Output: Parsed story development

[0142] Step 7:

[0143] The server transmits the parsed story development to the robot.

[0144] Specific operation: The server sends the analysis results to the robot via a dedicated API endpoint.

[0145] Input: Parsed story arc

[0146] Output: API request to the robot

[0147] Step 8:

[0148] The robot acts based on the received story development and displays it to the user.

[0149] Specific behavior: The robot reproduces the movements of the "protagonist who sets off into space" and utters the relevant lines. In addition, it sends the generated story text to the terminal, which displays it on the screen.

[0150] Input: API request from the server (the story unfolds)

[0151] Output: Robot movement, display on terminal

[0152] (Application example 1)

[0153] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0154] In recent years, many story generation and interactive robot systems have appeared, but these systems can only provide static stories and have difficulty responding flexibly to users' interests and requests. In addition, conventional systems have limited story visualization, and more immersive experiences are required. Furthermore, the functionality that allows users to enjoy stories in 3D space using smart glasses or head-mounted displays has not been fully achieved. To address these challenges, there is a demand for more engaging and immersive story experiences that utilize user interaction.

[0155] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0156] In this invention, the server includes a means for receiving story suggestions and questions from a user, a means for generating a story development and a character's reaction based on the received suggestions and questions, a means for a robot to act based on the generated development and reaction, a means for displaying a response based on the robot's action and reaction, a means for transmitting the generated story to a user's terminal via a network and visualizing it in a 3D space, and a means for displaying the generated story using smart glasses or a head-mounted display. This allows the user to enjoy a dynamic and personalized story based on their suggestions and questions in a 3D space. Furthermore, by using smart glasses or a head-mounted display, a more realistic story experience can be provided, improving the user's sense of immersion.

[0157] The term "user" refers to a human subject who performs operations and inputs, and who makes story suggestions and asks questions to the system.

[0158] A "story suggestion or question" is a question about the story's plot or characters that a user provides to the system.

[0159] "Generative AI" refers to artificial intelligence that generates new story developments and character reactions based on previously learned data.

[0160] The "server" is a computer system that receives input information from users, uses generative AI to generate story development and character reactions, and sends the results to each terminal and robot.

[0161] A "robot" is a mechanical device that can act based on the story development and character reactions transmitted from the server.

[0162] A "terminal" is an electronic device that acts as an interface for users to input story suggestions and questions and that communicates with the server.

[0163] A "network" is a communications infrastructure for data communication between a server, a terminal, and a robot.

[0164] "Means of visualization in 3D space" refers to technology for presenting the generated story to the user as a three-dimensional stereoscopic image.

[0165] "Smart glasses" are glasses-type devices that are worn by a user to visually display the generated story in three-dimensional space.

[0166] A "head-mounted display" is a device that a user wears and can directly view three-dimensional images projected on a display in front of the user's eyes.

[0167] The "means for displaying responses" is a function for visually presenting the robot's actions and reactions to the user on the terminal.

[0168] "Story development" refers to the progression of events or scenes that occur within a story.

[0169] "Character response" refers to the actions or lines that characters in the story show in response to the user's suggestions or questions.

[0170] An embodiment of the present invention is a system including the following elements.

[0171] 1. Hardware and Software Configuration

[0172] The servers use Amazon Web Services (AWS (registered trademark)) Lambda (registered trademark) and EC2 servers.

[0173] The user's terminal is a smartphone or tablet, which provides an interface for suggesting stories and inputting questions.

[0174] The smart glasses use Google® Glass® or Microsoft® HoloLens® to visualize the story in 3D space.

[0175] The robot uses programmable robots such as Pepper(R) to perform story-based actions.

[0176] 2. Data processing and data calculation processing

[0177] User input:

[0178] Users enter questions about the story's plot and characters on their smartphone or tablet, and this data is sent to the server via an HTTP POST request.

[0179] Example: User: "Make a story where the protagonist saves a medieval kingdom."

[0180] On the server:

[0181] The server receives the user's input data and inputs it as a prompt sentence to a generative AI model such as ChatGPT (registered trademark). The generative AI model generates a story development based on the received prompt sentence.

[0182] Example prompt: "User input: A story about saving a medieval kingdom. Generated story:"

[0183] Delivery and visualization of generative narratives:

[0184] The generated story is transmitted to the user's terminal via the network, where it is visualized using a 3D rendering engine (e.g., Unity3D (registered trademark)).

[0185] At the same time, operating instructions are sent to the robot, and the robot acts based on the story.

[0186] Smart Glasses and Head Mounted Displays Applications:

[0187] The generated story is also sent to smart glasses or a head-mounted display, through which the story is displayed to the user in 3D space.

[0188] 3. Specific application examples

[0189] For example, if a user suggests "create a story in which the protagonist saves a medieval kingdom," the server will generate a storyline based on the suggestion, send the generated story to the user's device, and visualize it in 3D space. The robot will also act out the story.

[0190] If the user then asks, "What happens if the main character meets a dragon?", the server uses generative AI to generate a dragon encounter scene in response to the question. The robot acts out the scene, which is then visually displayed on the smart glasses or head-mounted display.

[0191] The above is a specific embodiment for carrying out the present invention. With this configuration, the user can enjoy a dynamic and realistic story that responds to the user's input.

[0192] The flow of the specific process in the application example 1 will be described with reference to FIG.

[0193] Step 1:

[0194] Users enter suggestions or questions about the story's plot or characters on their smartphone or tablet, and that input (e.g., "Create a story where the hero saves a medieval kingdom") is sent to the server through the application's interface as an HTTP POST request.

[0195] Step 2:

[0196] The server receives a request sent by a user. Based on the received input data, the server generates a prompt to be given to a generative AI model (e.g. ChatGPT).

[0197] Example: Prompt: "User input: A story about saving a medieval kingdom. Generated story:"

[0198] Step 3:

[0199] The server uses a generative AI model to generate story development and character responses based on the prompt text. The result of this process is a new story text. The generated story text is stored on the server.

[0200] Input: prompt statement

[0201] Output: Generated narrative text

[0202] Step 4:

[0203] The server transmits the generated story text to the robot, and at the same time, the story text is also transmitted to the user's terminal via the network.

[0204] Input: Generated narrative text

[0205] Output: Data sent to the robot and the user's device

[0206] Step 5:

[0207] The user's device passes the received story text to a 3D rendering engine (e.g. Unity3D), which visualizes the story as a 3D space on the device and displays it to the user.

[0208] Input: Generated narrative text

[0209] Output: A narrative visualized in 3D space.

[0210] Step 6:

[0211] The robot acts based on the story text sent from the server. The robot executes the specified scenario and acts out the story with actions and sounds.

[0212] Input: Generated narrative text

[0213] Output: Robot behavior and audio output

[0214] Step 7:

[0215] If the user is wearing smart glasses or a head-mounted display, the device transmits the story text to these devices and displays it visually in 3D space, allowing the user to experience the generated story in a virtual reality environment.

[0216] Input: Generated narrative text

[0217] Output: Visualization on smart glasses or head-mounted displays

[0218] This is the process flow of the system. Through this series of steps, users can enjoy a dynamic and immersive story experience based on their own input.

[0219] Furthermore, an emotion engine that estimates the emotion of the user may be combined. That is, the identification processing unit 290 may estimate the emotion of the user using the emotion identification model 59, and perform identification processing using the emotion of the user.

[0220] An embodiment for implementing the present invention includes the following elements.

[0221] 1. Server

[0222] A server is provided that executes a program for generating story development and character reactions.

[0223] It is combined with an emotion engine that recognizes the user's emotions and performs voice and image analysis to recognize the user's emotions.

[0224] The results of the generative AI and emotion engine are combined to generate story development and character reactions.

[0225] 2. Terminal

[0226] It provides an interface for users to make story suggestions and questions.

[0227] The user's input information is sent to the server.

[0228] Voice and image data is sent to the server to recognize the user's emotions.

[0229] 3. Robot

[0230] Act on developments and reactions sent from the server.

[0231] As a character in the story, it will act and express itself according to the generated developments and reactions.

[0232] The response is displayed on the terminal, enabling a dialogue with the user.

[0233] (Specific examples)

[0234] For example, if a user suggests through a device, "Create a story in which the main character goes on a journey into space," the device will send that information to the server. Based on the suggestion, the server will use generative AI and an emotion engine to generate a plot related to space travel. The generated plot will be sent to the robot, which will then act out the journey into space. At the same time, the emotion engine will recognize emotions from the user's speech and facial expressions, and adjust the robot's behavior and reactions based on those emotions.

[0235] Also, if the user asks, "What would happen if the main character met an alien?", the device sends that information to the server. The server uses generative AI and an emotion engine to generate the character's reaction to meeting an alien. The robot acts based on that reaction and expresses its reaction to meeting an alien. At the same time, the emotion engine recognizes the user's emotions and adjusts the robot's response to match those emotions.

[0236] The above is an example of an embodiment of the present invention. Story creation and dialogue combining an emotion engine are specifically carried out by the roles and cooperation of the server, terminal, and robot.

[0237] The process flow will be explained below.

[0238] Step 1: Receiving User Input

[0239] The terminal provides an interface for the user to suggest stories and ask questions, and receives user-input information (text, voice, images, etc.).

[0240] Step 2: Sending input information and emotion recognition

[0241] The device sends the received user input information to the server. The server uses an emotion engine to recognize the user's emotions based on the received information. The emotion engine analyzes voice and images to extract the user's emotions.

[0242] Step 3: Receiving and processing information

[0243] The server receives story suggestions and questions sent from the terminal, as well as emotion information from the emotion engine.

[0244] Based on the information received, generative AI is used to generate story development and character reactions.

[0245] By combining the results of generative AI and the emotion engine, we can generate more flexible, emotionally-responsive developments and responses.

[0246] Step 4: Development and reaction generation

[0247] The server combines the results of the generative AI and the emotion engine to generate story developments and character reactions based on the information it receives, which are then sent to the robot in the next step.

[0248] Step 5: Robot behavior

[0249] The robot receives the developments and reactions sent from the server. Based on the information it receives, it acts as a character in the story. The specific actions it takes vary depending on the development of the story and the character's reactions.

[0250] Step 6: Viewing the Robot's Response

[0251] The robot displays responses on the terminal based on its own actions and reactions. Responses include the development of the story and the reactions of the characters. At the same time, it uses emotional information from the emotion engine to adjust the robot's facial expressions and tone of voice to respond in accordance with the emotions.

[0252] The above is the specific flow of the program's processing. The terminal receives the user's input, the server processes the information and generates developments and reactions, the robot acts based on that, and finally the robot's response is displayed on the terminal. This is the flow of events that combines the emotion engine to create a story and carry out a dialogue.

[0253] Example 2

[0254] Next, a description will be given of Example 2. In the following description, the data processing device 12 is referred to as a "server" and the smart device 14 is referred to as a "terminal."

[0255] Conventional story creation and dialogue systems can generate stories based on user suggestions and questions, but they have the problem of lacking naturalness and realism in dialogue because they do not take the user's emotions into account. In addition, the actions and emotional expressions of characters that respond in real time to story suggestions and questions input by the user are insufficient, making it difficult to increase user satisfaction.

[0256] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for receiving a story suggestion or question from a user, a means for generating a story development or a character's reaction using a natural language processing engine based on the received suggestion or question, a means for the robot to act based on the generated development and reaction, a means for displaying the robot's behavior and realizing a dialogue with the user, an emotion engine for recognizing the emotion by analyzing the user's voice or image, and a means for adjusting the robot's behavior based on the result of the emotion engine. This makes it possible to realize a story generation and a dialogue in real time that takes the user's emotion into consideration.

[0257] "User" refers to a person or institution that inputs a story suggestion or question.

[0258] "Terminal" refers to a device through which a user inputs story suggestions and questions and transmits them to a server.

[0259] The "server" refers to a central processing unit that generates story development and character reactions based on information received from the user, as well as analyzes emotions.

[0260] A "generative AI model" refers to an artificial intelligence model that generates story development and character responses based on user suggestions and questions.

[0261] A "natural language processing engine" refers to software that analyzes natural language entered by the user and generates appropriate story development and character responses.

[0262] An "emotion engine" refers to software that analyzes the user's voice and image data to recognize emotions, and adjusts the character's response based on that information.

[0263] A "robot" refers to a mechanical device that physically moves and expresses itself based on the story development and character reactions transmitted from the server.

[0264] "Story suggestions" refer to instructions or ideas about the story scenario or development that a user inputs via a terminal.

[0265] "Narrative development" refers to a series of storylines or scenarios that a generative AI model generates based on user suggestions.

[0266] "Character reactions" refers to the actions and expressions of characters that are generated according to the development of the story.

[0267] "Action" refers to the physical movements the robot makes based on the story development and character reactions.

[0268] This invention is a system for creating stories and conducting dialogue, and is mainly composed of a server, a terminal, and a robot.

[0269] Specifically, a user inputs story suggestions and questions via a terminal. The terminal is provided with an interface for inputting suggestions and questions, and is equipped with a microphone and a camera for capturing the user's voice and image. When the user inputs a prompt such as "Create a story in which the main character travels into space," the information is sent to the server.

[0270] The server uses several software components to generate storylines and character responses: First, it uses a generative AI model (for example, a system known as a natural language generation engine) to generate storylines based on the user's prompts. The generative AI model could be OpenAI's ChatGPT or another natural language processing engine.

[0271] Next, the server uses an emotion engine (such as Microsoft's Azure Emotion API or Google's Cloud Vision API) to analyze the user's voice and image data. This allows the server to recognize the user's emotional state and reflect it in the development of the generated story. As a result of the analysis, the emotion engine provides emotional information such as whether the user is happy, surprised, or angry.

[0272] The server combines the story development results of the generative AI model with the analysis results of the emotion engine to generate the final story development and character reactions. This generation result is encoded in an appropriate format such as JSON or XML and sent to the robot.

[0273] The robot performs physical movements and outputs voice based on the generated results. For example, the robot imitates a rocket launch in a space travel scene. It also uses a voice synthesis engine to speak the characters' lines. In this way, it provides the user with a more realistic storytelling experience.

[0274] As a concrete example, consider the case where a user asks, "What would happen if the main character met an alien?" The device sends this question to a server, which uses a generative AI model to generate plot developments and character reactions. At the same time, an emotion engine analyzes the user's voice and facial expressions to identify the user's emotional state. Based on the results, the robot expresses its reaction to the encounter with the alien and communicates this to the user through voice and action.

[0275] As described above, the present invention provides a system for realizing story generation and real-time dialogue that takes into account the user's emotions.

[0276] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0277] Step 1:

[0278] The user inputs story suggestions and questions into the device. Using the device's interface, the user inputs prompts such as "Create a story in which the main character travels into space." The user can also input voice and images. The input text, voice, and image data are captured by the device.

[0279] Step 2:

[0280] The terminal sends the user's input information and voice and image data to the server. The terminal divides the input data into packets and transfers them to the server using a security protocol (e.g. HTTPS, SSL / TLS). The input data includes the text entered by the user, the captured voice data, and image data. The server receives this data.

[0281] Step 3:

[0282] The server generates a story development using a generative AI model. The server inputs the text suggestions received from the user into a generative AI model (e.g., a natural language generation engine). For example, to generate a "story about going on a journey into space," the generative AI model is sent a prompt sentence such as "Make a story about the protagonist going on a journey into space." The generative AI model generates a story development based on this prompt sentence and returns the result to the server. The generated story development is output to the server.

[0283] Step 4:

[0284] The server analyzes the user's emotions using an emotion engine. The server inputs the received voice and image data into the emotion engine. The emotion engine (e.g., a voice tone analysis engine or a facial expression analysis engine) analyzes the voice tone and facial expression to identify the user's emotional state (e.g., happiness, surprise, anger). The analysis results of the emotion engine are output to the server.

[0285] Step 5:

[0286] The server integrates the results of the generative AI model and the emotion engine to generate story development and character reactions. The server integrates the story development obtained from the generative AI model with the user's emotion information obtained from the emotion engine. Based on the results of this integration, it generates the reactions of the story characters and further story content. The generated character reactions and story development are output to the server.

[0287] Step 6:

[0288] The server sends the generated results to the robot. The server converts the story development and character reactions generated by the server into an appropriate format (e.g. JSON, XML) and sends it to the robot. The robot receives the data sent from the server.

[0289] Step 7:

[0290] The robot acts and expresses itself based on the content it receives. The robot performs physical movements and outputs voice based on the development of the story it receives and the reactions of the characters. For example, in a scene where the robot departs into space, the robot looks up at the sky and moves its hands to mimic a rocket launch sequence. It also uses a voice synthesis engine to vocalize the characters' lines. This provides a story experience for the user.

[0291] Step 8:

[0292] The terminal displays the robot's actions and enables dialogue with the user. Based on the data sent from the robot, the terminal displays the robot's actions and speech. The user can visually and audibly confirm the robot's actions and speech through the terminal. This allows the user to make further story suggestions or ask questions, continuing the interaction with the system.

[0293] (Application example 2)

[0294] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0295] Conventional story creation systems and dialogue systems have been unable to generate story developments and character reactions that reflect the user's emotions in real time, limiting the user experience. The present invention aims to provide a more interactive and personalized story experience by recognizing the user's emotions and dynamically adjusting the story development and character reactions based on those emotions.

[0296] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means. In this invention, the server includes a means for recognizing the user's emotions using an emotion engine that recognizes the user's emotions based on the user's voice and image, a means for receiving a story proposal or question from the user, a prompt sentence that instructs generating a story development or a character's reaction based on the received story proposal or question and the user's emotions, a means for generating the story development or the character's reaction using a generative AI, and a means for controlling the robot's behavior based on the generated story development or character reaction. This makes it possible to provide a more interactive and personalized story experience by dynamically adjusting the story development or the character's reaction by recognizing the emotion when the user makes a story proposal or question. In addition, the server further includes a means for displaying the behavior of the robot on the user's terminal, and repeats recognizing the user's emotions, receiving a story proposal or question from the user, generating the story development or the character's reaction, controlling the robot's behavior, and displaying on the user's terminal. The means for controlling the behavior of the robot controls the robot to perform actions that imitate the reactions of the character, and controls the robot to output the lines of the character as voice.

[0297] "Recognizing a user's emotions" means analyzing input audio and image data and identifying the user's emotional state (e.g., joy, sadness, surprise, etc.).

[0298] "Story suggestions and questions" refers to a user requesting content related to the progression of the story or asking a question about the development of the story.

[0299] "Generating story development and character reactions" refers to using generative AI models and emotion engines to create story progression and character actions and reactions based on user suggestions and questions.

[0300] "The robot acts" refers to the robot physically moving and carrying out set actions based on the story development and character reactions sent from the server.

[0301] "Displaying responses" means visually presenting the generated story development and the character's reactions on the terminal screen or the robot's display.

[0302] "Server" refers to a computer system that receives user input and executes programs to generate story development and character responses using a generative AI model and emotion engine.

[0303] A "generative AI model" refers to an artificial intelligence algorithm that generates story development and character responses based on user suggestions and questions.

[0304] An "emotion engine" is software or hardware that analyzes a user's voice and image data and recognizes emotions.

[0305] "Dynamic adjustment" refers to changing the story content and character reactions in real time based on the user's emotional recognition.

[0306] "Interactive and personalized narrative experience" means that the story progresses in response to user input and emotional state, providing a customized storytelling experience for each individual user.

[0307] As an embodiment of the present invention, a system for generating story development and character reactions includes the following elements: A user inputs story suggestions and questions through a terminal, and the server uses a generative AI model and an emotion engine based on that information to generate story development and character reactions. The generated development and reactions are sent to a robot, which operates based on the content and displays a response.

[0308] Specifically, the server is equipped with an emotion engine for analyzing the user's voice and image data, and a generative AI model that generates the story development and character reactions. When a user suggests through their device, "Make a story in which the main character goes on a journey into space," the device sends that information to the server. The server uses the emotion engine to recognize emotions from the user's voice and images. The server passes the recognized emotion data and the user's suggestions to the generative AI model to generate the story development.

[0309] The generated story development and character reactions are sent to the robot, and the robot acts accordingly. For example, if a user asks, "What happens if the main character meets an alien?", the server uses the information to generate a story development and character reactions regarding the encounter with the alien, using an emotion engine and generative AI model. The robot acts based on this reaction and displays the response on the terminal.

[0310] Examples of specific hardware and software used in servers, terminals, and robots include the emotion engine "EmotionEngine" and the generative AI model "StoryGenAI." This makes it possible to provide an interactive storytelling experience that dynamically adjusts the development of the story and the reactions of characters based on the user's emotions.

[0311] For example, if a user requests, "Make a story where the main character goes on a journey into space," the server will analyze the user's emotions and generate a positive development using a generative AI model. In addition, if a user asks, "What would happen if the main character met an alien?", the server will generate a character's reaction based on the emotional data, and the robot will act accordingly.

[0312] Example of a prompt that generates a storyline based on the story suggestions and the user's emotions: "The user requested, 'Make a story where the main character goes on a journey into space.' The user's emotions are bright and positive, so please generate a positive development accordingly."

[0313] An example prompt that generates a character's response based on the question and the user's emotions: "The user asked, 'What would happen if the main character met an alien?' The user's emotions are bright and positive, so generate a character reaction based on that emotion."

[0314] The flow of the specific process in the application example 2 will be described with reference to FIG.

[0315] Step 1:

[0316] The user inputs story suggestions and questions through the terminal. The input data includes the user's voice, text, and image data. For example, a request might be "Make a story where the main character goes on a journey into space." This input data is sent directly to the terminal.

[0317] Step 2:

[0318] The device sends the user's input information to the server. This includes voice and image data, and is sent in JSON format formatted for processing on the server side. The input data includes the user's request and voice and image data. The data received as input to the server is analyzed by the emotion engine.

[0319] Step 3:

[0320] The server uses the emotion engine to recognize the user's emotions from the received voice and image data. In this step, it performs voice and image analysis to convert the user's emotions into concrete tags (e.g., happiness, sadness, surprise, etc.). It obtains the emotion recognition results and prepares them to be passed to the generative AI model.

[0321] Step 4:

[0322] The server uses a generative AI model to generate a story development and character reactions based on the user's suggestions and questions. In this step, the user's input text and emotion recognition results are input as prompts to the generative AI model to obtain a generated story development and character reactions. Specifically, a story is generated based on the request "The protagonist goes on a space journey."

[0323] Step 5:

[0324] The server sends the generated story development and character reaction data to the robot. The output includes the story content and character actions, which are converted into robot movement instructions. The robot receives this data and starts moving according to the content.

[0325] Step 6:

[0326] The robot physically moves based on the received story development and the character's reaction. For example, the robot acts out a scene in which it travels into space. The robot's actions are displayed on the user's device. Specific examples of the actions include the robot making a specific gesture or speaking the character's lines.

[0327] Step 7:

[0328] If the user makes a new question or suggestion, the device again sends that information to the server, and the following steps 2 to 6 are executed in a loop. This cycle dynamically adjusts the development of the story, realizing an interactive story experience that responds to the user's emotions. For example, in response to the question, "What would happen if the main character met an alien?", the server generates a new development and the robot behaves based on it.

[0329] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires a voice indicating a user input for the result of the specific processing. The control unit 46A transmits the voice data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0330] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0331] In the above embodiment, an example was given in which the specific process was performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0332] [Second embodiment]

[0333] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0334] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0335] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).

[0336] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0337] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs the voice according to instructions from the processor 46.

[0338] Camera 42 is a small digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (e.g., an imaging range defined by an angle of view equivalent to the width of the field of vision of an average healthy person).

[0339] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54. The transmission and reception of various types of information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is performed in a secure state.

[0340] Fig. 4 shows an example of main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0341] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32, and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0342] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0343] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0344] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal".

[0345] An embodiment for implementing the present invention includes the following elements.

[0346] 1. Server

[0347] A server is provided that executes a program for generating story development and character reactions.

[0348] It receives story suggestions and questions from users, and uses generative AI to generate developments and responses based on them.

[0349] The generated developments and reactions are sent to the robot.

[0350] 2. Terminal

[0351] It provides an interface for users to make story suggestions and questions.

[0352] The user's input information is sent to the server.

[0353] 3. Robot

[0354] Act on developments and reactions sent from the server.

[0355] As a character in the story, it will act and express itself according to the generated developments and reactions.

[0356] The response is displayed on the terminal, enabling a dialogue with the user.

[0357] (Specific examples)

[0358] For example, if a user suggests through a device, "Make a story where the main character goes on a journey into space," the device will send that information to the server. The server will use generative AI based on the suggestion to generate a plot related to space travel. The generated plot will be sent to the robot, which will then act out the journey into space.

[0359] Also, if the user asks, "What would happen if the main character met an alien?", the device sends that information to the server. The server uses generative AI to generate the character's reaction to meeting an alien. The robot acts based on that reaction and expresses its reaction to meeting an alien.

[0360] The above is an example of an embodiment of the present invention. A system for creating a story and realizing a dialogue is specifically embodied by the roles and cooperation of a server, a terminal, and a robot.

[0361] The process flow will be explained below.

[0362] Step 1: Receiving User Input

[0363] The terminal provides an interface for the user to make story suggestions and questions, and receives information entered by the user.

[0364] Step 2: The input information sending terminal sends the received user input information to the server, including the contents of suggestions and questions.

[0365] Step 3: Receiving and processing information

[0366] The server receives the information sent from the device and uses generative AI to generate the story development and character reactions based on the received information.

[0367] Step 4: Development and reaction generation

[0368] The server uses generative AI to generate story developments and character reactions based on the information it receives, which are then sent to the robot in the next step.

[0369] Step 5: Robot behavior

[0370] The robot receives the developments and reactions sent from the server. Based on the information it receives, it acts as a character in the story. The specific actions it takes vary depending on the development of the story.

[0371] Step 6: Viewing the Robot's Response

[0372] The robot displays responses on the device based on its own actions and reactions, including story developments and character reactions.

[0373] The above is the specific flow of the program's processing. The terminal receives the user's input, the server processes the information and generates developments and reactions, the robot acts based on that, and finally the robot's response is displayed on the terminal. This is how the story is created and the dialogue is carried out.

[0374] Example 1

[0375] Next, a description will be given of Example 1. In the following description, the data processing device 12 is referred to as a "server" and the smart glasses 214 are referred to as a "terminal".

[0376] Conventional story generation systems have difficulty generating natural story developments and character reactions to user input in real time, and there is a lack of systems that can express them as real objects and characters. This limits interactive dialogue with the user, resulting in a poor quality experience.

[0377] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0378] In this invention, the server includes a means for receiving story suggestions and questions from a user, a means for generating a story development and a character's reaction using a generative AI model based on the received suggestions and questions, and a means for the robot to act based on the generated development and reaction. This enables the generation and expression of a dynamic and natural story development and character reaction in response to user input.

[0379] "User" refers to the person or user who inputs story suggestions and questions.

[0380] "Terminal" refers to an electronic device that provides an interface through which a user can enter story suggestions and questions and transmit them to a server.

[0381] The "server" refers to a computer system that receives suggestions and questions sent by users, uses a generative AI model to generate story developments and character responses, and transmits them to the robot.

[0382] A "generative AI model" is a model that uses artificial intelligence technology and is an algorithm that performs natural language processing based on user input to generate story development and character reactions.

[0383] A "prompt sentence" refers to a text sentence entered into a generative AI model to give it a specific instruction or request.

[0384] A "robot" refers to a physical device that acts as a character in a story based on the generation results sent from a server, performs actions and expressions, and displays its responses on a terminal.

[0385] "Story suggestions and questions" refer to instructions or questions that a user inputs via a terminal regarding the content of the story or the actions of a character and that are then transmitted to the server.

[0386] "Development and response" refers to the story progression and character responses that the generative AI model generates based on user suggestions and questions.

[0387] To implement this invention, hardware such as a server, terminal, and robot, as well as software such as a generative AI model, are required. This system generates story development and character reactions based on user input, and expresses them as real objects, thereby realizing interactive dialogue with the user.

[0388] System Configuration

[0389] 1. Server

[0390] Hardware: High-performance computer systems (e.g., a server rack with multiple processors, memory, storage, etc.)

[0391] Software: Generative AI models (e.g. ChatGPT by OpenAI), data processing algorithms, HTTP servers, etc.

[0392] Function: Receives story suggestions and questions sent by the user via the device, and uses a generative AI model to generate story development and character responses based on those suggestions. The generated data is then sent to the robot.

[0393] 2. Terminal

[0394] Hardware: User interface devices such as smartphones, tablets, and computers

[0395] Software: Web applications, mobile applications

[0396] Function: Provides an interface for users to enter story suggestions and questions, and sends the information to a server.

[0397] 3. Robots

[0398] Hardware: Physically shaped robots (e.g. humanoid robots, animal-like robots, etc.), sensors and actuators for movement and voice output

[0399] Software: Robot control program

[0400] Function: Acts and expresses itself based on the story development and character reactions sent from the server, and displays the results on the device.

[0401] How it works

[0402] When a user uses a device to input story suggestions or questions, the input data is sent from the device to a server. The server analyzes the received input data and uses a generative AI model to generate the story development and character reactions. The generated data is sent to the robot, which then operates based on that data. As a result, the robot behaves and speaks as a character in the story, and displays this on the device, realizing an interactive dialogue with the user.

[0403] Examples

[0404] For example, if a user uses a terminal to suggest, "Create a story where the main character goes on a journey into space," the process would go like this:

[0405] User: Type "Create a story where the main character goes on a space journey" into the terminal.

[0406] Terminal: Sends the entered information to the server.

[0407] Server: Sends the prompt "Create a story in which the protagonist goes on a journey into space" to the generative AI model and generates the story development.

[0408] Server: Sends the generated story development to the robot.

[0409] Robot: Performs movements that represent "the protagonist setting off into space."

[0410] Robot: Display the results on the terminal.

[0411] Also, if a user asks, "What happens if the main character meets an alien?"

[0412] User: Type into terminal, "What happens when the main character meets an alien?"

[0413] Terminal: Sends the entered information to the server.

[0414] Server: Sends this prompt to the generative AI model and generates the character's response.

[0415] Server: Sends the generated responses to the robot.

[0416] Robot: Performs actions that express the reaction of the protagonist when he meets the alien.

[0417] Robot: Display the results on the terminal.

[0418] In this manner, the present invention enables the generation and expression of dynamic and natural story development and character reactions to user input.

[0419] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0420] Step 1:

[0421] A user uses the terminal to input story suggestions and questions.

[0422] Specific operations: The user enters "Create a story in which the main character goes on a journey into space" into the input form displayed on the device screen.

[0423] Input: User suggestions or questions ("Create a story in which the main character travels into space")

[0424] Output: Input data that the device sends to the server

[0425] Step 2:

[0426] The terminal transmits the user's input information to the server.

[0427] Specific operation: The terminal sends the entered string data to the server as an HTTP request.

[0428] Input: String data entered by the user

[0429] Output: HTTP request to the server

[0430] Step 3:

[0431] The server receives and analyzes the input data sent from the terminal.

[0432] Specific operation: The server receives an HTTP request and extracts the user input contained in the request body.

[0433] Input: HTTP request from the terminal

[0434] Output: Extraction result of user input data

[0435] Step 4:

[0436] The server sends a prompt to the generative AI model based on the received input data.

[0437] Specific operation: The server generates a prompt such as "Create a story in which the protagonist goes on a journey into space" and sends it to the generative AI model as an API request.

[0438] Input: Data entered by the user

[0439] Output: API request to the generative AI model

[0440] Step 5:

[0441] The generative AI model generates a story development in response to the prompt sentence and returns the result to the server.

[0442] Specific operation: The generative AI model analyzes the prompt sentence, generates a story development, and returns it to the server as an API response.

[0443] Input: API request to the generative AI model (prompt text)

[0444] Output: Generated story development

[0445] Step 6:

[0446] The server receives and analyzes the story development returned by the generative AI model.

[0447] Specific operation: The server receives the API response returned from the generative AI model and parses the contents in JSON format.

[0448] Input: API response from the generative AI model

[0449] Output: Parsed story development

[0450] Step 7:

[0451] The server transmits the parsed story development to the robot.

[0452] Specific operation: The server sends the analysis results to the robot via a dedicated API endpoint.

[0453] Input: Parsed story arc

[0454] Output: API request to the robot

[0455] Step 8:

[0456] The robot acts based on the received story development and displays it to the user.

[0457] Specific behavior: The robot reproduces the movements of the "protagonist who sets off into space" and utters the relevant lines. In addition, it sends the generated story text to the terminal, which displays it on the screen.

[0458] Input: API request from the server (the story unfolds)

[0459] Output: Robot movement, display on terminal

[0460] (Application example 1)

[0461] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal".

[0462] In recent years, many story generation and interactive robot systems have appeared, but these systems can only provide static stories and have difficulty responding flexibly to users' interests and requests. In addition, conventional systems have limited story visualization, and more immersive experiences are required. Furthermore, the functionality that allows users to enjoy stories in 3D space using smart glasses or head-mounted displays has not been fully achieved. To address these challenges, there is a demand for more engaging and immersive story experiences that utilize user interaction.

[0463] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0464] In this invention, the server includes a means for receiving story suggestions and questions from a user, a means for generating a story development and a character's reaction based on the received suggestions and questions, a means for a robot to act based on the generated development and reaction, a means for displaying a response based on the robot's action and reaction, a means for transmitting the generated story to a user's terminal via a network and visualizing it in a 3D space, and a means for displaying the generated story using smart glasses or a head-mounted display. This allows the user to enjoy a dynamic and personalized story based on their suggestions and questions in a 3D space. Furthermore, by using smart glasses or a head-mounted display, a more realistic story experience can be provided, improving the user's sense of immersion.

[0465] The term "user" refers to a human subject who performs operations and inputs, and who makes story suggestions and asks questions to the system.

[0466] A "story suggestion or question" is a question about the story's plot or characters that a user provides to the system.

[0467] "Generative AI" refers to artificial intelligence that generates new story developments and character reactions based on previously learned data.

[0468] The "server" is a computer system that receives input information from users, uses generative AI to generate story development and character reactions, and sends the results to each terminal and robot.

[0469] A "robot" is a mechanical device that can act based on the story development and character reactions transmitted from the server.

[0470] A "terminal" is an electronic device that acts as an interface for users to input story suggestions and questions and that communicates with the server.

[0471] A "network" is a communications infrastructure for data communication between a server, a terminal, and a robot.

[0472] "Means of visualization in 3D space" refers to technology for presenting the generated story to the user as a three-dimensional stereoscopic image.

[0473] "Smart glasses" are glasses-type devices that are worn by a user to visually display the generated story in three-dimensional space.

[0474] A "head-mounted display" is a device that a user wears and can directly view three-dimensional images projected on a display in front of the user's eyes.

[0475] The "means for displaying responses" is a function for visually presenting the robot's actions and reactions to the user on the terminal.

[0476] "Story development" refers to the progression of events or scenes that occur within a story.

[0477] "Character response" refers to the actions or lines that characters in the story show in response to the user's suggestions or questions.

[0478] An embodiment of the present invention is a system including the following elements.

[0479] 1. Hardware and Software Configuration

[0480] The servers use Amazon Web Services (AWS) Lambda and EC2 servers.

[0481] The user's terminal is a smartphone or tablet, which provides an interface for suggesting stories and inputting questions.

[0482] The smart glasses use Google Glass® or Microsoft HoloLens® to visualize the story in 3D space.

[0483] The robots use programmable robots such as Pepper to perform story-based actions.

[0484] 2. Data processing and data calculation processing

[0485] User input:

[0486] Users enter questions about the story's plot and characters on their smartphone or tablet, and this data is sent to the server via an HTTP POST request.

[0487] Example: User: "Make a story where the protagonist saves a medieval kingdom."

[0488] On the server:

[0489] The server receives the user's input data and inputs it as a prompt sentence to a generative AI model such as ChatGPT. The generative AI model generates the development of the story based on the received prompt sentence.

[0490] Example prompt: "User input: A story about saving a medieval kingdom. Generated story:"

[0491] Delivery and visualization of generative narratives:

[0492] The generated story is sent over the network to the user's device, where it is visualized using a 3D rendering engine (e.g., Unity3D).

[0493] At the same time, operating instructions are sent to the robot, and the robot acts based on the story.

[0494] Smart Glasses and Head Mounted Displays Applications:

[0495] The generated story is also sent to smart glasses or a head-mounted display, through which the story is displayed to the user in 3D space.

[0496] 3. Specific application examples

[0497] For example, if a user suggests "create a story in which the protagonist saves a medieval kingdom," the server will generate a storyline based on the suggestion, send the generated story to the user's device, and visualize it in 3D space. The robot will also act out the story.

[0498] If the user then asks, "What happens if the main character meets a dragon?", the server uses generative AI to generate a dragon encounter scene in response to the question. The robot acts out the scene, which is then visually displayed on the smart glasses or head-mounted display.

[0499] The above is a specific embodiment for carrying out the present invention. With this configuration, the user can enjoy a dynamic and realistic story that responds to the user's input.

[0500] The flow of the specific process in the application example 1 will be described with reference to FIG.

[0501] Step 1:

[0502] Users enter suggestions or questions about the story's plot or characters on their smartphone or tablet, and that input (e.g., "Create a story where the hero saves a medieval kingdom") is sent to the server through the application's interface as an HTTP POST request.

[0503] Step 2:

[0504] The server receives a request sent by a user. Based on the received input data, the server generates a prompt to be given to a generative AI model (e.g. ChatGPT).

[0505] Example: Prompt: "User input: A story about saving a medieval kingdom. Generated story:"

[0506] Step 3:

[0507] The server uses a generative AI model to generate story development and character responses based on the prompt text. The result of this process is a new story text. The generated story text is stored on the server.

[0508] Input: prompt statement

[0509] Output: Generated narrative text

[0510] Step 4:

[0511] The server transmits the generated story text to the robot, and at the same time, the story text is also transmitted to the user's terminal via the network.

[0512] Input: Generated narrative text

[0513] Output: Data sent to the robot and the user's device

[0514] Step 5:

[0515] The user's device passes the received story text to a 3D rendering engine (e.g. Unity3D), which visualizes the story as a 3D space on the device and displays it to the user.

[0516] Input: Generated narrative text

[0517] Output: A narrative visualized in 3D space.

[0518] Step 6:

[0519] The robot acts based on the story text sent from the server. The robot executes the specified scenario and acts out the story with actions and sounds.

[0520] Input: Generated narrative text

[0521] Output: Robot behavior and audio output

[0522] Step 7:

[0523] If the user is wearing smart glasses or a head-mounted display, the device transmits the story text to these devices and displays it visually in 3D space, allowing the user to experience the generated story in a virtual reality environment.

[0524] Input: Generated narrative text

[0525] Output: Visualization on smart glasses or head-mounted displays

[0526] This is the process flow of the system. Through this series of steps, users can enjoy a dynamic and immersive story experience based on their own input.

[0527] In addition, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0528] An embodiment for implementing the present invention includes the following elements.

[0529] 1. Server

[0530] A server is provided that executes a program for generating story development and character reactions.

[0531] It is combined with an emotion engine that recognizes the user's emotions and performs voice and image analysis to recognize the user's emotions.

[0532] The results of the generative AI and emotion engine are combined to generate story development and character reactions.

[0533] 2. Terminal

[0534] It provides an interface for users to make story suggestions and questions.

[0535] The user's input information is sent to the server.

[0536] Voice and image data is sent to the server to recognize the user's emotions.

[0537] 3. Robot

[0538] Act on developments and reactions sent from the server.

[0539] As a character in the story, it will act and express itself according to the generated developments and reactions.

[0540] The response is displayed on the terminal, enabling a dialogue with the user.

[0541] (Specific examples)

[0542] For example, if a user suggests through a device, "Create a story in which the main character goes on a journey into space," the device will send that information to the server. Based on the suggestion, the server will use generative AI and an emotion engine to generate a plot related to space travel. The generated plot will be sent to the robot, which will then act out the journey into space. At the same time, the emotion engine will recognize emotions from the user's speech and facial expressions, and adjust the robot's behavior and reactions based on those emotions.

[0543] Also, if the user asks, "What would happen if the main character met an alien?", the device sends that information to the server. The server uses generative AI and an emotion engine to generate the character's reaction to meeting an alien. The robot acts based on that reaction and expresses its reaction to meeting an alien. At the same time, the emotion engine recognizes the user's emotions and adjusts the robot's response to match those emotions.

[0544] The above is an example of an embodiment of the present invention. Story creation and dialogue combining an emotion engine are specifically carried out by the roles and cooperation of the server, terminal, and robot.

[0545] The process flow will be explained below.

[0546] Step 1: Receiving User Input

[0547] The terminal provides an interface for the user to suggest stories and ask questions, and receives user-input information (text, voice, images, etc.).

[0548] Step 2: Sending input information and emotion recognition

[0549] The terminal transmits the received user input information to the server.

[0550] The server uses an emotion engine to recognize the user's emotions based on the received information.

[0551] The emotion engine analyzes audio and images to extract the user's emotions.

[0552] Step 3: Receiving and processing information

[0553] The server receives story suggestions and questions sent from the terminal, as well as emotion information from the emotion engine.

[0554] Based on the information received, generative AI is used to generate story development and character reactions.

[0555] By combining the results of generative AI and the emotion engine, we can generate more flexible, emotionally-responsive developments and responses.

[0556] Step 4: Development and reaction generation

[0557] The server combines the results of the generative AI and the emotion engine to generate story developments and character reactions based on the information it receives, which are then sent to the robot in the next step.

[0558] Step 5: Robot behavior

[0559] The robot receives the developments and reactions sent from the server. Based on the information it receives, it acts as a character in the story. The specific actions it takes vary depending on the development of the story and the character's reactions.

[0560] Step 6: Viewing the Robot's Response

[0561] The robot displays responses on the terminal based on its own actions and reactions. Responses include the development of the story and the reactions of the characters. At the same time, it uses emotional information from the emotion engine to adjust the robot's facial expressions and tone of voice to respond in accordance with the emotions.

[0562] The above is the specific flow of the program's processing. The terminal receives the user's input, the server processes the information and generates developments and reactions, the robot acts based on that, and finally the robot's response is displayed on the terminal. This is the flow of events that combines the emotion engine to create a story and carry out a dialogue.

[0563] Example 2

[0564] Next, a description will be given of Example 2. In the following description, the data processing device 12 is referred to as a "server" and the smart glasses 214 are referred to as a "terminal".

[0565] Conventional story creation and dialogue systems can generate stories based on user suggestions and questions, but they have the problem of lacking naturalness and realism in dialogue because they do not take the user's emotions into account. In addition, the actions and emotional expressions of characters that respond in real time to story suggestions and questions input by the user are insufficient, making it difficult to increase user satisfaction.

[0566] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for receiving a story suggestion or question from a user, a means for generating a story development or a character's reaction using a natural language processing engine based on the received suggestion or question, a means for the robot to act based on the generated development and reaction, a means for displaying the robot's behavior and realizing a dialogue with the user, an emotion engine for recognizing the emotion by analyzing the user's voice or image, and a means for adjusting the robot's behavior based on the result of the emotion engine. This makes it possible to realize a story generation and a dialogue in real time that takes the user's emotion into consideration.

[0567] "User" refers to a person or institution that inputs a story suggestion or question.

[0568] "Terminal" refers to a device through which a user inputs story suggestions and questions and transmits them to a server.

[0569] The "server" refers to a central processing unit that generates story development and character reactions based on information received from the user, as well as analyzes emotions.

[0570] A "generative AI model" refers to an artificial intelligence model that generates story development and character responses based on user suggestions and questions.

[0571] A "natural language processing engine" refers to software that analyzes natural language entered by the user and generates appropriate story development and character responses.

[0572] An "emotion engine" refers to software that analyzes the user's voice and image data to recognize emotions, and adjusts the character's response based on that information.

[0573] A "robot" refers to a mechanical device that physically moves and expresses itself based on the story development and character reactions transmitted from the server.

[0574] "Story suggestions" refer to instructions or ideas about the story scenario or development that a user inputs via a terminal.

[0575] "Narrative development" refers to a series of storylines or scenarios that a generative AI model generates based on user suggestions.

[0576] "Character reactions" refers to the actions and expressions of characters that are generated according to the development of the story.

[0577] "Action" refers to the physical movements the robot makes based on the story development and character reactions.

[0578] This invention is a system for creating stories and conducting dialogue, and is mainly composed of a server, a terminal, and a robot.

[0579] Specifically, a user inputs story suggestions and questions via a terminal. The terminal is provided with an interface for inputting suggestions and questions, and is equipped with a microphone and a camera for capturing the user's voice and image. When the user inputs a prompt such as "Create a story in which the main character travels into space," the information is sent to the server.

[0580] The server uses several software components to generate storylines and character responses: First, it uses a generative AI model (for example, a system known as a natural language generation engine) to generate storylines based on the user's prompts. The generative AI model could be OpenAI's ChatGPT or another natural language processing engine.

[0581] Next, the server uses an emotion engine (such as Microsoft's Azure Emotion API or Google's Cloud Vision API) to analyze the user's voice and image data. This allows the server to recognize the user's emotional state and reflect it in the development of the generated story. As a result of the analysis, the emotion engine provides emotional information such as whether the user is happy, surprised, or angry.

[0582] The server combines the story development results of the generative AI model with the analysis results of the emotion engine to generate the final story development and character reactions. This generation result is encoded in an appropriate format such as JSON or XML and sent to the robot.

[0583] The robot performs physical movements and outputs voice based on the generated results. For example, the robot imitates a rocket launch in a space travel scene. It also uses a voice synthesis engine to speak the characters' lines. In this way, it provides the user with a more realistic storytelling experience.

[0584] As a concrete example, consider the case where a user asks, "What would happen if the main character met an alien?" The device sends this question to a server, which uses a generative AI model to generate plot developments and character reactions. At the same time, an emotion engine analyzes the user's voice and facial expressions to identify the user's emotional state. Based on the results, the robot expresses its reaction to the encounter with the alien and communicates this to the user through voice and action.

[0585] As described above, the present invention provides a system for realizing story generation and real-time dialogue that takes into account the user's emotions.

[0586] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0587] Step 1:

[0588] The user inputs story suggestions and questions into the device. Using the device's interface, the user inputs prompts such as "Create a story in which the main character travels into space." The user can also input voice and images. The input text, voice, and image data are captured by the device.

[0589] Step 2:

[0590] The terminal sends the user's input information and voice and image data to the server. The terminal divides the input data into packets and transfers them to the server using a security protocol (e.g. HTTPS, SSL / TLS). The input data includes the text entered by the user, the captured voice data, and image data. The server receives this data.

[0591] Step 3:

[0592] The server generates a story development using a generative AI model. The server inputs the text suggestions received from the user into a generative AI model (e.g., a natural language generation engine). For example, to generate a "story about going on a journey into space," the generative AI model is sent a prompt sentence such as "Make a story about the protagonist going on a journey into space." The generative AI model generates a story development based on this prompt sentence and returns the result to the server. The generated story development is output to the server.

[0593] Step 4:

[0594] The server analyzes the user's emotions using an emotion engine. The server inputs the received voice and image data into the emotion engine. The emotion engine (e.g., a voice tone analysis engine or a facial expression analysis engine) analyzes the voice tone and facial expression to identify the user's emotional state (e.g., happiness, surprise, anger). The analysis results of the emotion engine are output to the server.

[0595] Step 5:

[0596] The server integrates the results of the generative AI model and the emotion engine to generate story development and character reactions. The server integrates the story development obtained from the generative AI model with the user's emotion information obtained from the emotion engine. Based on the results of this integration, it generates the reactions of the story characters and further story content. The generated character reactions and story development are output to the server.

[0597] Step 6:

[0598] The server sends the generated results to the robot. The server converts the story development and character reactions generated by the server into an appropriate format (e.g. JSON, XML) and sends it to the robot. The robot receives the data sent from the server.

[0599] Step 7:

[0600] The robot acts and expresses itself based on the content it receives. The robot performs physical movements and outputs voice based on the development of the story it receives and the reactions of the characters. For example, in a scene where the robot departs into space, the robot looks up at the sky and moves its hands to mimic a rocket launch sequence. It also uses a voice synthesis engine to vocalize the characters' lines. This provides a story experience for the user.

[0601] Step 8:

[0602] The terminal displays the robot's actions and enables dialogue with the user. Based on the data sent from the robot, the terminal displays the robot's actions and speech. The user can visually and audibly confirm the robot's actions and speech through the terminal. This allows the user to make further story suggestions or ask questions, continuing the interaction with the system.

[0603] (Application example 2)

[0604] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal".

[0605] Conventional story creation systems and dialogue systems have been unable to generate story developments and character reactions that reflect the user's emotions in real time, limiting the user experience. The present invention aims to provide a more interactive and personalized story experience by recognizing the user's emotions and dynamically adjusting the story development and character reactions based on those emotions.

[0606] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means. In this invention, the server includes a means for recognizing the user's emotions using an emotion engine that recognizes the user's emotions based on the user's voice and image, a means for receiving a story proposal or question from the user, a prompt sentence that instructs generating a story development or a character's reaction based on the received story proposal or question and the user's emotions, a means for generating the story development or the character's reaction using a generative AI, and a means for controlling the robot's behavior based on the generated story development or character reaction. This makes it possible to provide a more interactive and personalized story experience by dynamically adjusting the story development or the character's reaction by recognizing the emotion when the user makes a story proposal or question. In addition, the server further includes a means for displaying the behavior of the robot on the user's terminal, and repeats recognizing the user's emotions, receiving a story proposal or question from the user, generating the story development or the character's reaction, controlling the robot's behavior, and displaying on the user's terminal. The means for controlling the behavior of the robot controls the robot to perform actions that imitate the reactions of the character, and controls the robot to output the lines of the character as voice.

[0607] "Recognizing a user's emotions" means analyzing input audio and image data and identifying the user's emotional state (e.g., joy, sadness, surprise, etc.).

[0608] "Story suggestions and questions" refers to a user requesting content related to the progression of the story or asking a question about the development of the story.

[0609] "Generating story development and character reactions" refers to using generative AI models and emotion engines to create story progression and character actions and reactions based on user suggestions and questions.

[0610] "The robot acts" refers to the robot physically moving and carrying out set actions based on the story development and character reactions sent from the server.

[0611] "Displaying responses" means visually presenting the generated story development and the character's reactions on the terminal screen or the robot's display.

[0612] "Server" refers to a computer system that receives user input and executes programs to generate story development and character responses using a generative AI model and emotion engine.

[0613] A "generative AI model" refers to an artificial intelligence algorithm that generates story development and character responses based on user suggestions and questions.

[0614] An "emotion engine" is software or hardware that analyzes a user's voice and image data and recognizes emotions.

[0615] "Dynamic adjustment" refers to changing the story content and character reactions in real time based on the user's emotional recognition.

[0616] "Interactive and personalized narrative experience" means that the story progresses in response to user input and emotional state, providing a customized storytelling experience for each individual user.

[0617] As an embodiment of the present invention, a system for generating story development and character reactions includes the following elements: A user inputs story suggestions and questions through a terminal, and the server uses a generative AI model and an emotion engine based on that information to generate story development and character reactions. The generated development and reactions are sent to a robot, which operates based on the content and displays a response.

[0618] Specifically, the server is equipped with an emotion engine for analyzing the user's voice and image data, and a generative AI model that generates the story development and character reactions. When a user suggests through their device, "Make a story in which the main character goes on a journey into space," the device sends that information to the server. The server uses the emotion engine to recognize emotions from the user's voice and images. The server passes the recognized emotion data and the user's suggestions to the generative AI model to generate the story development.

[0619] The generated story development and character reactions are sent to the robot, and the robot acts accordingly. For example, if a user asks, "What happens if the main character meets an alien?", the server uses the information to generate a story development and character reactions regarding the encounter with the alien, using an emotion engine and generative AI model. The robot acts based on this reaction and displays the response on the terminal.

[0620] Examples of specific hardware and software used in servers, terminals, and robots include the emotion engine "EmotionEngine" and the generative AI model "StoryGenAI." This makes it possible to provide an interactive storytelling experience that dynamically adjusts the development of the story and the reactions of characters based on the user's emotions.

[0621] For example, if a user requests, "Make a story where the main character goes on a journey into space," the server will analyze the user's emotions and generate a positive development using a generative AI model. In addition, if a user asks, "What would happen if the main character met an alien?", the server will generate a character's reaction based on the emotional data, and the robot will act accordingly.

[0622] Example of a prompt that generates a storyline based on the story suggestions and the user's emotions: "The user requested, 'Make a story where the main character goes on a journey into space.' The user's emotions are bright and positive, so please generate a positive development accordingly."

[0623] An example prompt that generates a character's response based on the question and the user's emotions: "The user asked, 'What would happen if the main character met an alien?' The user's emotions are bright and positive, so generate a character reaction based on that emotion."

[0624] The flow of the specific process in the application example 2 will be described with reference to FIG.

[0625] Step 1:

[0626] The user inputs story suggestions and questions through the terminal. The input data includes the user's voice, text, and image data. For example, a request might be "Make a story where the main character goes on a journey into space." This input data is sent directly to the terminal.

[0627] Step 2:

[0628] The device sends the user's input information to the server. This includes voice and image data, and is sent in JSON format formatted for processing on the server side. The input data includes the user's request and voice and image data. The data received as input to the server is analyzed by the emotion engine.

[0629] Step 3:

[0630] The server uses the emotion engine to recognize the user's emotions from the received voice and image data. In this step, it performs voice and image analysis to convert the user's emotions into concrete tags (e.g., happiness, sadness, surprise, etc.). It obtains the emotion recognition results and prepares them to be passed to the generative AI model.

[0631] Step 4:

[0632] The server uses a generative AI model to generate a story development and character reactions based on the user's suggestions and questions. In this step, the user's input text and emotion recognition results are input as prompts to the generative AI model to obtain a generated story development and character reactions. Specifically, a story is generated based on the request "The protagonist goes on a space journey."

[0633] Step 5:

[0634] The server sends the generated story development and character reaction data to the robot. The output includes the story content and character actions, which are converted into robot movement instructions. The robot receives this data and starts moving according to the content.

[0635] Step 6:

[0636] The robot physically moves based on the received story development and the character's reaction. For example, the robot acts out a scene in which it travels into space. The robot's actions are displayed on the user's device. Specific examples of the actions include the robot making a specific gesture or speaking the character's lines.

[0637] Step 7:

[0638] If the user makes a new question or suggestion, the device again sends that information to the server, and the following steps 2 to 6 are executed in a loop. This cycle dynamically adjusts the development of the story, realizing an interactive story experience that responds to the user's emotions. For example, in response to the question, "What would happen if the main character met an alien?", the server generates a new development and the robot behaves based on it.

[0639] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires a voice indicating a user input for the result of the specific processing. The control unit 46A transmits the voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0640] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0641] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0642] [Third embodiment]

[0643] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0644] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0645] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).

[0646] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0647] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs the voice according to instructions from the processor 46.

[0648] Camera 42 is a small digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (e.g., an imaging range defined by an angle of view equivalent to the width of the field of vision of an average healthy person).

[0649] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54. The transmission and reception of various types of information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is performed in a secure state.

[0650] Fig. 6 shows an example of main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0651] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32, and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0652] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0653] In the headset type terminal 314, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0654] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server", and the headset type terminal 314 will be referred to as the "terminal".

[0655] An embodiment for implementing the present invention includes the following elements.

[0656] 1. Server

[0657] A server is provided that executes a program for generating story development and character reactions.

[0658] It receives story suggestions and questions from users, and uses generative AI to generate developments and responses based on them.

[0659] The generated developments and reactions are sent to the robot.

[0660] 2. Terminal

[0661] It provides an interface for users to make story suggestions and questions.

[0662] The user's input information is sent to the server.

[0663] 3. Robot

[0664] Act on developments and reactions sent from the server.

[0665] As a character in the story, it will act and express itself according to the generated developments and reactions.

[0666] The response is displayed on the terminal, enabling a dialogue with the user.

[0667] (Specific examples)

[0668] For example, if a user suggests through the terminal, "Make a story where the main character goes on a journey into space," the terminal will send that information to the server. The server will use ChatGPT based on the received suggestion to generate a plot related to space travel. The generated plot will be sent to the robot, which will then act out the journey into space.

[0669] Also, if the user asks, "What would happen if the main character met an alien?", the device sends that information to the server. The server uses generative AI to generate the character's reaction to meeting an alien. The robot acts based on that reaction and expresses its reaction to meeting an alien.

[0670] The above is an example of an embodiment of the present invention. A system for creating a story and realizing a dialogue is specifically embodied by the roles and cooperation of a server, a terminal, and a robot.

[0671] The process flow will be explained below.

[0672] Step 1: Receiving User Input

[0673] The terminal provides an interface for the user to make story suggestions and questions, and receives information entered by the user.

[0674] Step 2: Submit your input

[0675] The device sends the received user input information, including suggestions and questions, to the server.

[0676] Step 3: Receiving and processing information

[0677] The server receives the information sent from the device and uses generative AI to generate the story development and character reactions based on the received information.

[0678] Step 4: Development and reaction generation

[0679] The server uses generative AI to generate story developments and character reactions based on the information it receives, which are then sent to the robot in the next step.

[0680] Step 5: Robot behavior

[0681] The robot receives the developments and reactions sent from the server. Based on the information it receives, it acts as a character in the story. The specific actions it takes vary depending on the development of the story.

[0682] Step 6: Viewing the Robot's Response

[0683] The robot displays responses on the device based on its own actions and reactions, including story developments and character reactions.

[0684] The above is the specific flow of the program's processing. The terminal receives the user's input, the server processes the information and generates developments and reactions, the robot acts based on that, and finally the robot's response is displayed on the terminal. This is how the story is created and the dialogue is carried out.

[0685] Example 1

[0686] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal".

[0687] Conventional story generation systems have difficulty generating natural story developments and character reactions to user input in real time, and there is a lack of systems that can express them as real objects and characters. This limits interactive dialogue with the user, resulting in a poor quality experience.

[0688] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0689] In this invention, the server includes a means for receiving story suggestions and questions from a user, a means for generating a story development and a character's reaction using a generative AI model based on the received suggestions and questions, and a means for the robot to act based on the generated development and reaction. This enables the generation and expression of a dynamic and natural story development and character reaction in response to user input.

[0690] "User" refers to the person or user who inputs story suggestions and questions.

[0691] "Terminal" refers to an electronic device that provides an interface through which a user can enter story suggestions and questions and transmit them to a server.

[0692] The "server" refers to a computer system that receives suggestions and questions sent by users, uses a generative AI model to generate story developments and character responses, and transmits them to the robot.

[0693] A "generative AI model" is a model that uses artificial intelligence technology and is an algorithm that performs natural language processing based on user input to generate story development and character reactions.

[0694] A "prompt sentence" refers to a text sentence entered into a generative AI model to give it a specific instruction or request.

[0695] A "robot" refers to a physical device that acts as a character in a story based on the generation results sent from a server, performs actions and expressions, and displays its responses on a terminal.

[0696] "Story suggestions and questions" refer to instructions or questions that a user inputs via a terminal regarding the content of the story or the actions of a character and that are then transmitted to the server.

[0697] "Development and response" refers to the story progression and character responses that the generative AI model generates based on user suggestions and questions.

[0698] To implement this invention, hardware such as a server, terminal, and robot, as well as software such as a generative AI model, are required. This system generates story development and character reactions based on user input, and expresses them as real objects, thereby realizing interactive dialogue with the user.

[0699] System Configuration

[0700] 1. Server

[0701] Hardware: High-performance computer systems (e.g., a server rack with multiple processors, memory, storage, etc.)

[0702] Software: Generative AI models (e.g. ChatGPT by OpenAI), data processing algorithms, HTTP servers, etc.

[0703] Function: Receives story suggestions and questions sent by the user via the device, and uses a generative AI model to generate story development and character responses based on those suggestions. The generated data is then sent to the robot.

[0704] 2. Terminal

[0705] Hardware: User interface devices such as smartphones, tablets, and computers

[0706] Software: Web applications, mobile applications

[0707] Function: Provides an interface for users to enter story suggestions and questions, and sends the information to a server.

[0708] 3. Robots

[0709] Hardware: Physically shaped robots (e.g. humanoid robots, animal-like robots, etc.), sensors and actuators for movement and voice output

[0710] Software: Robot control program

[0711] Function: Acts and expresses itself based on the story development and character reactions sent from the server, and displays the results on the device.

[0712] How it works

[0713] When a user uses a device to input story suggestions or questions, the input data is sent from the device to a server. The server analyzes the received input data and uses a generative AI model to generate the story development and character reactions. The generated data is sent to the robot, which then operates based on that data. As a result, the robot behaves and speaks as a character in the story, and displays this on the device, realizing an interactive dialogue with the user.

[0714] Examples

[0715] For example, if a user uses a terminal to suggest, "Create a story where the main character goes on a journey into space," the process would go like this:

[0716] User: Type "Create a story where the main character goes on a space journey" into the terminal.

[0717] Terminal: Sends the entered information to the server.

[0718] Server: Sends the prompt "Create a story in which the protagonist goes on a journey into space" to the generative AI model and generates the story development.

[0719] Server: Sends the generated story development to the robot.

[0720] Robot: Performs movements that represent "the protagonist setting off into space."

[0721] Robot: Display the results on the terminal.

[0722] Also, if a user asks, "What happens if the main character meets an alien?"

[0723] User: Type into terminal, "What happens when the main character meets an alien?"

[0724] Terminal: Sends the entered information to the server.

[0725] Server: Sends this prompt to the generative AI model and generates the character's response.

[0726] Server: Sends the generated responses to the robot.

[0727] Robot: Performs actions that express the reaction of the protagonist when he meets the alien.

[0728] Robot: Display the results on the terminal.

[0729] In this manner, the present invention enables the generation and expression of dynamic and natural story development and character reactions to user input.

[0730] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0731] Step 1:

[0732] A user uses the terminal to input story suggestions and questions.

[0733] Specific operations: The user enters "Create a story in which the main character goes on a journey into space" into the input form displayed on the device screen.

[0734] Input: User suggestions or questions ("Create a story in which the main character travels into space")

[0735] Output: Input data that the device sends to the server

[0736] Step 2:

[0737] The terminal transmits the user's input information to the server.

[0738] Specific operation: The terminal sends the entered string data to the server as an HTTP request.

[0739] Input: String data entered by the user

[0740] Output: HTTP request to the server

[0741] Step 3:

[0742] The server receives and analyzes the input data sent from the terminal.

[0743] Specific operation: The server receives an HTTP request and extracts the user input contained in the request body.

[0744] Input: HTTP request from the terminal

[0745] Output: Extraction result of user input data

[0746] Step 4:

[0747] The server sends a prompt to the generative AI model based on the received input data.

[0748] Specific operation: The server generates a prompt such as "Create a story in which the protagonist goes on a journey into space" and sends it to the generative AI model as an API request.

[0749] Input: Data entered by the user

[0750] Output: API request to the generative AI model

[0751] Step 5:

[0752] The generative AI model generates a story development in response to the prompt sentence and returns the result to the server.

[0753] Specific operation: The generative AI model analyzes the prompt sentence, generates a story development, and returns it to the server as an API response.

[0754] Input: API request to the generative AI model (prompt text)

[0755] Output: Generated story development

[0756] Step 6:

[0757] The server receives and analyzes the story development returned by the generative AI model.

[0758] Specific operation: The server receives the API response returned from the generative AI model and parses the contents in JSON format.

[0759] Input: API response from the generative AI model

[0760] Output: Parsed story development

[0761] Step 7:

[0762] The server transmits the parsed story development to the robot.

[0763] Specific operation: The server sends the analysis results to the robot via a dedicated API endpoint.

[0764] Input: Parsed story arc

[0765] Output: API request to the robot

[0766] Step 8:

[0767] The robot acts based on the received story development and displays it to the user.

[0768] Specific behavior: The robot reproduces the movements of the "protagonist who sets off into space" and utters the relevant lines. In addition, it sends the generated story text to the terminal, which displays it on the screen.

[0769] Input: API request from the server (the story unfolds)

[0770] Output: Robot movement, display on terminal

[0771] (Application example 1)

[0772] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0773] In recent years, many story generation and interactive robot systems have appeared, but these systems can only provide static stories and have difficulty responding flexibly to users' interests and requests. In addition, conventional systems have limited story visualization, and more immersive experiences are required. Furthermore, the functionality that allows users to enjoy stories in 3D space using smart glasses or head-mounted displays has not been fully achieved. To address these challenges, there is a demand for more engaging and immersive story experiences that utilize user interaction.

[0774] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0775] In this invention, the server includes a means for receiving story suggestions and questions from a user, a means for generating a story development and a character's reaction based on the received suggestions and questions, a means for a robot to act based on the generated development and reaction, a means for displaying a response based on the robot's action and reaction, a means for transmitting the generated story to a user's terminal via a network and visualizing it in a 3D space, and a means for displaying the generated story using smart glasses or a head-mounted display. This allows the user to enjoy a dynamic and personalized story based on their suggestions and questions in a 3D space. Furthermore, by using smart glasses or a head-mounted display, a more realistic story experience can be provided, improving the user's sense of immersion.

[0776] The term "user" refers to a human subject who performs operations and inputs, and who makes story suggestions and asks questions to the system.

[0777] A "story suggestion or question" is a question about the story's plot or characters that a user provides to the system.

[0778] "Generative AI" refers to artificial intelligence that generates new story developments and character reactions based on previously learned data.

[0779] The "server" is a computer system that receives input information from users, uses generative AI to generate story development and character reactions, and sends the results to each terminal and robot.

[0780] A "robot" is a mechanical device that can act based on the story development and character reactions transmitted from the server.

[0781] A "terminal" is an electronic device that acts as an interface for users to input story suggestions and questions and that communicates with the server.

[0782] A "network" is a communications infrastructure for data communication between a server, a terminal, and a robot.

[0783] "Means of visualization in 3D space" refers to technology for presenting the generated story to the user as a three-dimensional stereoscopic image.

[0784] "Smart glasses" are glasses-type devices that are worn by a user to visually display the generated story in three-dimensional space.

[0785] A "head-mounted display" is a device that a user wears and can directly view three-dimensional images projected on a display in front of the user's eyes.

[0786] The "means for displaying responses" is a function for visually presenting the robot's actions and reactions to the user on the terminal.

[0787] "Story development" refers to the progression of events or scenes that occur within a story.

[0788] "Character response" refers to the actions or lines that characters in the story show in response to the user's suggestions or questions.

[0789] An embodiment of the present invention is a system including the following elements.

[0790] 1. Hardware and Software Configuration

[0791] The servers use Amazon Web Services (AWS) Lambda and EC2 servers.

[0792] The user's terminal is a smartphone or tablet, which provides an interface for suggesting stories and inputting questions.

[0793] The smart glasses use Google Glass or Microsoft HoloLens to visualize the story in 3D space.

[0794] The robots use programmable robots such as Pepper to perform story-based actions.

[0795] 2. Data processing and data calculation processing

[0796] User input:

[0797] Users enter questions about the story's plot and characters on their smartphone or tablet, and this data is sent to the server via an HTTP POST request.

[0798] Example: User: "Make a story where the protagonist saves a medieval kingdom."

[0799] On the server:

[0800] The server receives the user's input data and inputs it as a prompt sentence to a generative AI model such as ChatGPT. The generative AI model generates the development of the story based on the received prompt sentence.

[0801] Example prompt: "User input: A story about saving a medieval kingdom. Generated story:"

[0802] Delivery and visualization of generative narratives:

[0803] The generated story is sent over the network to the user's device, where it is visualized using a 3D rendering engine (e.g., Unity3D).

[0804] At the same time, operating instructions are sent to the robot, and the robot acts based on the story.

[0805] Smart Glasses and Head Mounted Displays Applications:

[0806] The generated story is also sent to smart glasses or a head-mounted display, through which the story is displayed to the user in 3D space.

[0807] 3. Specific application examples

[0808] For example, if a user suggests "create a story in which the protagonist saves a medieval kingdom," the server will generate a storyline based on the suggestion, send the generated story to the user's device, and visualize it in 3D space. The robot will also act out the story.

[0809] If the user then asks, "What happens if the main character meets a dragon?", the server uses generative AI to generate a dragon encounter scene in response to the question. The robot acts out the scene, which is then visually displayed on the smart glasses or head-mounted display.

[0810] The above is a specific embodiment for carrying out the present invention. This configuration enables the user to enjoy a dynamic and realistic story that responds to the user's input.

[0811] The flow of the specific process in the application example 1 will be described with reference to FIG.

[0812] Step 1:

[0813] Users enter suggestions or questions about the story's plot or characters on their smartphone or tablet, and that input (e.g., "Create a story where the protagonist saves a medieval kingdom") is sent to the server through the application's interface as an HTTP POST request.

[0814] Step 2:

[0815] The server receives requests sent by users. Based on the received input data, the server generates a prompt to be given to a generative AI model (e.g. ChatGPT).

[0816] Example: Prompt: "User input: A story about saving a medieval kingdom. Generated story:"

[0817] Step 3:

[0818] The server uses a generative AI model to generate story development and character responses based on the prompt text. The result of this process is a new story text. The generated story text is stored on the server.

[0819] Input: prompt statement

[0820] Output: Generated narrative text

[0821] Step 4:

[0822] The server transmits the generated story text to the robot, and at the same time, the story text is also transmitted to the user's terminal via the network.

[0823] Input: Generated narrative text

[0824] Output: Data sent to the robot and the user's device

[0825] Step 5:

[0826] The user's device passes the received story text to a 3D rendering engine (e.g. Unity3D), which visualizes the story as a 3D space on the device and displays it to the user.

[0827] Input: Generated narrative text

[0828] Output: A narrative visualized in 3D space.

[0829] Step 6:

[0830] The robot acts based on the story text sent from the server. The robot executes the specified scenario and acts out the story with actions and sounds.

[0831] Input: Generated narrative text

[0832] Output: Robot behavior and audio output

[0833] Step 7:

[0834] If the user is wearing smart glasses or a head-mounted display, the device transmits the story text to these devices and displays it visually in 3D space, allowing the user to experience the generated story in a virtual reality environment.

[0835] Input: Generated narrative text

[0836] Output: Visualization on smart glasses or head-mounted displays

[0837] This is the process flow of the system. Through this series of steps, users can enjoy a dynamic and immersive story experience based on their own input.

[0838] In addition, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0839] An embodiment for implementing the present invention includes the following elements.

[0840] 1. Server

[0841] A server is provided that executes a program for generating story development and character reactions.

[0842] It is combined with an emotion engine that recognizes the user's emotions and performs voice and image analysis to recognize the user's emotions.

[0843] The results of the generative AI and emotion engine are combined to generate story development and character reactions.

[0844] 2. Terminal

[0845] It provides an interface for users to make story suggestions and questions.

[0846] The user's input information is sent to the server.

[0847] Voice and image data is sent to the server to recognize the user's emotions.

[0848] 3. Robot

[0849] Act on developments and reactions sent from the server.

[0850] As a character in the story, it will act and express itself according to the generated developments and reactions.

[0851] The response is displayed on the terminal, enabling a dialogue with the user.

[0852] (Specific examples)

[0853] For example, if a user suggests through a device, "Create a story in which the main character goes on a journey into space," the device will send that information to the server. Based on the suggestion, the server will use generative AI and an emotion engine to generate a plot related to space travel. The generated plot will be sent to the robot, which will then act out the journey into space. At the same time, the emotion engine will recognize emotions from the user's speech and facial expressions, and adjust the robot's behavior and reactions based on those emotions.

[0854] Also, if the user asks, "What would happen if the main character met an alien?", the device sends that information to the server. The server uses generative AI and an emotion engine to generate the character's reaction to meeting an alien. The robot acts based on that reaction and expresses its reaction to meeting an alien. At the same time, the emotion engine recognizes the user's emotions and adjusts the robot's response to match those emotions.

[0855] The above is an example of an embodiment of the present invention. Story creation and dialogue combining an emotion engine are specifically carried out by the roles and cooperation of the server, terminal, and robot.

[0856] The process flow will be explained below.

[0857] Step 1: Receiving User Input

[0858] The terminal provides an interface for the user to suggest stories and ask questions, and receives user-input information (text, voice, images, etc.).

[0859] Step 2: Sending input information and emotion recognition

[0860] The terminal transmits the received user input information to the server.

[0861] The server uses an emotion engine to recognize the user's emotions based on the received information.

[0862] The emotion engine analyzes audio and images to extract the user's emotions.

[0863] Step 3: Receiving and processing information

[0864] The server receives story suggestions and questions sent from the terminal, as well as emotion information from the emotion engine.

[0865] Based on the information received, generative AI is used to generate story development and character reactions.

[0866] By combining the results of generative AI and the emotion engine, we can generate more flexible, emotionally-responsive developments and responses.

[0867] Step 4: Development and reaction generation

[0868] The server combines the results of the generative AI and the emotion engine to generate story developments and character reactions based on the information it receives, which are then sent to the robot in the next step.

[0869] Step 5: Robot behavior

[0870] The robot receives the developments and reactions sent from the server. Based on the information it receives, it acts as a character in the story. The specific actions it takes vary depending on the development of the story and the character's reactions.

[0871] Step 6: Viewing the Robot's Response

[0872] The robot displays responses on the terminal based on its own actions and reactions. Responses include the development of the story and the reactions of the characters. At the same time, it uses emotional information from the emotion engine to adjust the robot's facial expressions and tone of voice to respond in accordance with the emotions.

[0873] The above is the specific flow of the program's processing. The terminal receives the user's input, the server processes the information and generates developments and reactions, the robot acts based on that, and finally the robot's response is displayed on the terminal. This is the flow of events that combines the emotion engine to create a story and carry out a dialogue.

[0874] Example 2

[0875] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal".

[0876] Conventional story creation and dialogue systems can generate stories based on user suggestions and questions, but they have the problem of lacking naturalness and realism in dialogue because they do not take the user's emotions into account. In addition, the actions and emotional expressions of characters that respond in real time to story suggestions and questions input by the user are insufficient, making it difficult to increase user satisfaction.

[0877] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for receiving a story suggestion or question from a user, a means for generating a story development or a character's reaction using a natural language processing engine based on the received suggestion or question, a means for the robot to act based on the generated development and reaction, a means for displaying the robot's behavior and realizing a dialogue with the user, an emotion engine for recognizing the emotion by analyzing the user's voice or image, and a means for adjusting the robot's behavior based on the result of the emotion engine. This makes it possible to realize a story generation and a dialogue in real time that takes the user's emotion into consideration.

[0878] "User" refers to a person or institution that inputs a story suggestion or question.

[0879] "Terminal" refers to a device through which a user inputs story suggestions and questions and transmits them to a server.

[0880] The "server" refers to a central processing unit that generates story development and character reactions based on information received from the user, as well as analyzes emotions.

[0881] A "generative AI model" refers to an artificial intelligence model that generates story development and character responses based on user suggestions and questions.

[0882] A "natural language processing engine" refers to software that analyzes natural language entered by the user and generates appropriate story development and character responses.

[0883] An "emotion engine" refers to software that analyzes the user's voice and image data to recognize emotions, and adjusts the character's response based on that information.

[0884] A "robot" refers to a mechanical device that physically moves and expresses itself based on the story development and character reactions transmitted from the server.

[0885] "Story suggestions" refer to instructions or ideas about the story scenario or development that a user inputs via a terminal.

[0886] "Narrative development" refers to a series of storylines or scenarios that a generative AI model generates based on user suggestions.

[0887] "Character reactions" refers to the actions and expressions of characters that are generated according to the development of the story.

[0888] "Action" refers to the physical movements the robot makes based on the story development and character reactions.

[0889] This invention is a system for creating stories and conducting dialogue, and is mainly composed of a server, a terminal, and a robot.

[0890] Specifically, a user inputs story suggestions and questions via a terminal. The terminal is provided with an interface for inputting suggestions and questions, and is equipped with a microphone and a camera for capturing the user's voice and image. When the user inputs a prompt such as "Create a story in which the main character travels into space," the information is sent to the server.

[0891] The server uses several software components to generate storylines and character responses: First, it uses a generative AI model (for example, a system known as a natural language generation engine) to generate storylines based on the user's prompts. The generative AI model could be OpenAI's ChatGPT or another natural language processing engine.

[0892] Next, the server uses an emotion engine (such as Microsoft's Azure Emotion API or Google's Cloud Vision API) to analyze the user's voice and image data. This allows the server to recognize the user's emotional state and reflect it in the development of the generated story. As a result of the analysis, the emotion engine provides emotional information such as whether the user is happy, surprised, or angry.

[0893] The server combines the story development results of the generative AI model with the analysis results of the emotion engine to generate the final story development and character reactions. This generation result is encoded in an appropriate format such as JSON or XML and sent to the robot.

[0894] The robot performs physical movements and outputs voice based on the generated results. For example, the robot imitates a rocket launch in a space travel scene. It also uses a voice synthesis engine to speak the characters' lines. In this way, it provides the user with a more realistic storytelling experience.

[0895] As a concrete example, consider the case where a user asks, "What would happen if the main character met an alien?" The device sends this question to a server, which uses a generative AI model to generate plot developments and character reactions. At the same time, an emotion engine analyzes the user's voice and facial expressions to identify the user's emotional state. Based on the results, the robot expresses its reaction to the encounter with the alien and communicates this to the user through voice and action.

[0896] As described above, the present invention provides a system for realizing story generation and real-time dialogue that takes into account the user's emotions.

[0897] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0898] Step 1:

[0899] The user inputs story suggestions and questions into the device. Using the device's interface, the user inputs prompts such as "Create a story in which the main character travels into space." The user can also input voice and images. The input text, voice, and image data are captured by the device.

[0900] Step 2:

[0901] The terminal sends the user's input information and voice and image data to the server. The terminal divides the input data into packets and transfers them to the server using a security protocol (e.g. HTTPS, SSL / TLS). The input data includes the text entered by the user, the captured voice data, and image data. The server receives this data.

[0902] Step 3:

[0903] The server generates a story development using a generative AI model. The server inputs the text suggestions received from the user into a generative AI model (e.g., a natural language generation engine). For example, to generate a "story about going on a journey into space," the generative AI model is sent a prompt sentence such as "Make a story about the protagonist going on a journey into space." The generative AI model generates a story development based on this prompt sentence and returns the result to the server. The generated story development is output to the server.

[0904] Step 4:

[0905] The server analyzes the user's emotions using an emotion engine. The server inputs the received voice and image data into the emotion engine. The emotion engine (e.g., a voice tone analysis engine or a facial expression analysis engine) analyzes the voice tone and facial expression to identify the user's emotional state (e.g., happiness, surprise, anger). The analysis results of the emotion engine are output to the server.

[0906] Step 5:

[0907] The server integrates the results of the generative AI model and the emotion engine to generate story development and character reactions. The server integrates the story development obtained from the generative AI model with the user's emotion information obtained from the emotion engine. Based on the results of this integration, it generates the reactions of the story characters and further story content. The generated character reactions and story development are output to the server.

[0908] Step 6:

[0909] The server sends the generated results to the robot. The server converts the story development and character reactions generated by the server into an appropriate format (e.g. JSON, XML) and sends it to the robot. The robot receives the data sent from the server.

[0910] Step 7:

[0911] The robot acts and expresses itself based on the content it receives. The robot performs physical movements and outputs voice based on the development of the story it receives and the reactions of the characters. For example, in a scene where the robot departs into space, the robot looks up at the sky and moves its hands to mimic a rocket launch sequence. It also uses a voice synthesis engine to vocalize the characters' lines. This provides a story experience for the user.

[0912] Step 8:

[0913] The terminal displays the robot's actions and enables dialogue with the user. Based on the data sent from the robot, the terminal displays the robot's actions and speech. The user can visually and audibly confirm the robot's actions and speech through the terminal. This allows the user to make further story suggestions or ask questions, continuing the interaction with the system.

[0914] (Application example 2)

[0915] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server", and the headset type terminal 314 will be referred to as a "terminal".

[0916] Conventional story creation systems and dialogue systems have been unable to generate story developments and character reactions that reflect the user's emotions in real time, limiting the user experience. The present invention aims to provide a more interactive and personalized story experience by recognizing the user's emotions and dynamically adjusting the story development and character reactions based on those emotions.

[0917] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means. In this invention, the server includes a means for recognizing the user's emotions using an emotion engine that recognizes the user's emotions based on the user's voice and image, a means for receiving a story proposal or question from the user, a prompt sentence that instructs generating a story development or a character's reaction based on the received story proposal or question and the user's emotions, a means for generating the story development or the character's reaction using a generative AI, and a means for controlling the robot's behavior based on the generated story development or character reaction. This makes it possible to provide a more interactive and personalized story experience by dynamically adjusting the story development or the character's reaction by recognizing the emotion when the user makes a story proposal or question. In addition, the server further includes a means for displaying the behavior of the robot on the user's terminal, and repeats recognizing the user's emotions, receiving a story proposal or question from the user, generating the story development or the character's reaction, controlling the robot's behavior, and displaying on the user's terminal. The means for controlling the behavior of the robot controls the robot to perform actions that imitate the reactions of the character, and controls the robot to output the lines of the character as voice.

[0918] "Recognizing a user's emotions" means analyzing input audio and image data and identifying the user's emotional state (e.g., joy, sadness, surprise, etc.).

[0919] "Story suggestions and questions" refers to a user requesting content related to the progression of the story or asking a question about the development of the story.

[0920] "Generating story development and character reactions" refers to using generative AI models and emotion engines to create story progression and character actions and reactions based on user suggestions and questions.

[0921] "The robot acts" refers to the robot physically moving and carrying out set actions based on the story development and character reactions sent from the server.

[0922] "Displaying responses" means visually presenting the generated story development and the character's reactions on the terminal screen or the robot's display.

[0923] "Server" refers to a computer system that receives user input and executes programs to generate story development and character responses using a generative AI model and emotion engine.

[0924] A "generative AI model" refers to an artificial intelligence algorithm that generates story development and character responses based on user suggestions and questions.

[0925] An "emotion engine" is software or hardware that analyzes a user's voice and image data and recognizes emotions.

[0926] "Dynamic adjustment" refers to changing the story content and character reactions in real time based on the user's emotional recognition.

[0927] "Interactive and personalized narrative experience" means that the story progresses in response to user input and emotional state, providing a customized storytelling experience for each individual user.

[0928] As an embodiment of the present invention, a system for generating story development and character reactions includes the following elements: A user inputs story suggestions and questions through a terminal, and the server uses a generative AI model and an emotion engine based on that information to generate story development and character reactions. The generated development and reactions are sent to a robot, which operates based on the content and displays a response.

[0929] Specifically, the server is equipped with an emotion engine for analyzing the user's voice and image data, and a generative AI model that generates the story development and character reactions. When a user suggests through their device, "Make a story in which the main character goes on a journey into space," the device sends that information to the server. The server uses the emotion engine to recognize emotions from the user's voice and images. The server passes the recognized emotion data and the user's suggestions to the generative AI model to generate the story development.

[0930] The generated story development and character reactions are sent to the robot, and the robot acts accordingly. For example, if a user asks, "What happens if the main character meets an alien?", the server uses the information to generate a story development and character reactions regarding the encounter with the alien, using an emotion engine and generative AI model. The robot acts based on this reaction and displays the response on the terminal.

[0931] Examples of specific hardware and software used in servers, terminals, and robots include the emotion engine "EmotionEngine" and the generative AI model "StoryGenAI." This makes it possible to provide an interactive storytelling experience that dynamically adjusts the development of the story and the reactions of characters based on the user's emotions.

[0932] For example, if a user requests, "Make a story where the main character goes on a journey into space," the server will analyze the user's emotions and generate a positive development using a generative AI model. In addition, if a user asks, "What would happen if the main character met an alien?", the server will generate a character's reaction based on the emotional data, and the robot will act accordingly.

[0933] Example of a prompt that generates a storyline based on the story suggestions and the user's emotions: "The user requested, 'Make a story where the main character goes on a journey into space.' The user's emotions are bright and positive, so please generate a positive development accordingly."

[0934] An example prompt that generates a character's response based on the question and the user's emotions: "The user asked, 'What would happen if the main character met an alien?' The user's emotions are bright and positive, so generate a character reaction based on that emotion."

[0935] The flow of the specific process in the application example 2 will be described with reference to FIG.

[0936] Step 1:

[0937] The user inputs story suggestions and questions through the terminal. The input data includes the user's voice, text, and image data. For example, a request might be "Make a story where the main character goes on a journey into space." This input data is sent directly to the terminal.

[0938] Step 2:

[0939] The device sends the user's input information to the server. This includes voice and image data, and is sent in JSON format formatted for processing on the server side. The input data includes the user's request and voice and image data. The data received as input to the server is analyzed by the emotion engine.

[0940] Step 3:

[0941] The server uses the emotion engine to recognize the user's emotions from the received voice and image data. In this step, it performs voice and image analysis to convert the user's emotions into concrete tags (e.g., happiness, sadness, surprise, etc.). It obtains the emotion recognition results and prepares them to be passed to the generative AI model.

[0942] Step 4:

[0943] The server uses a generative AI model to generate a story development and character reactions based on the user's suggestions and questions. In this step, the user's input text and emotion recognition results are input as prompts to the generative AI model to obtain a generated story development and character reactions. Specifically, a story is generated based on the request "The protagonist goes on a space journey."

[0944] Step 5:

[0945] The server sends the generated story development and character reaction data to the robot. The output includes the story content and character actions, which are converted into robot movement instructions. The robot receives this data and starts moving according to the content.

[0946] Step 6:

[0947] The robot physically moves based on the received story development and the character's reaction. For example, the robot acts out a scene in which it travels into space. The robot's actions are displayed on the user's device. Specific examples of the actions include the robot making a specific gesture or speaking the character's lines.

[0948] Step 7:

[0949] If the user makes a new question or suggestion, the device again sends that information to the server, and the following steps 2 to 6 are executed in a loop. This cycle dynamically adjusts the development of the story, realizing an interactive story experience that responds to the user's emotions. For example, in response to the question, "What would happen if the main character met an alien?", the server generates a new development and the robot behaves based on it.

[0950] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input for the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0951] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0952] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[0953] [Fourth embodiment]

[0954] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[0955] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0956] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).

[0957] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. In addition, the microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[0958] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs the voice according to instructions from the processor 46.

[0959] Camera 42 is a small digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (e.g., an imaging range defined by an angle of view equivalent to the width of the field of vision of an average healthy person).

[0960] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54. The transmission and reception of various types of information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is performed in a secure state.

[0961] The control target 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, legs, etc. The posture and behavior of the robot 414 are controlled by controlling the motors of the arms, hands, legs, etc. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[0962] Fig. 8 shows an example of main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0963] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32, and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0964] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0965] In the robot 414, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0966] Next, a description will be given of the specific processing by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal".

[0967] An embodiment for implementing the present invention includes the following elements.

[0968] 1. Server

[0969] A server is provided that executes a program for generating story development and character reactions.

[0970] It receives story suggestions and questions from users, and uses generative AI to generate developments and responses based on them.

[0971] The generated developments and reactions are sent to the robot.

[0972] 2. Terminal

[0973] It provides an interface for users to make story suggestions and questions.

[0974] The user's input information is sent to the server.

[0975] 3. Robot

[0976] Act on developments and reactions sent from the server.

[0977] As a character in the story, it will act and express itself according to the generated developments and reactions.

[0978] The response is displayed on the terminal, enabling a dialogue with the user.

[0979] (Specific examples)

[0980] For example, if a user suggests through a device, "Make a story where the main character goes on a journey into space," the device will send that information to the server. The server will use generative AI based on the suggestion to generate a plot related to space travel. The generated plot will be sent to the robot, which will then act out the journey into space.

[0981] Also, if the user asks, "What would happen if the main character met an alien?", the device sends that information to the server. The server uses generative AI to generate the character's reaction to meeting an alien. The robot acts based on that reaction and expresses its reaction to meeting an alien.

[0982] The above is an example of an embodiment of the present invention. A system for creating a story and realizing a dialogue is specifically embodied by the roles and cooperation of a server, a terminal, and a robot.

[0983] The process flow will be explained below.

[0984] Step 1: Receiving User Input

[0985] The terminal provides an interface for the user to make story suggestions and questions, and receives information entered by the user.

[0986] Step 2: Submit your input

[0987] The device sends the received user input information, including suggestions and questions, to the server.

[0988] Step 3: Receiving and processing information

[0989] The server receives the information sent from the device and uses generative AI to generate the story development and character reactions based on the received information.

[0990] Step 4: Development and reaction generation

[0991] The server uses generative AI to generate story developments and character reactions based on the information it receives, which are then sent to the robot in the next step.

[0992] Step 5: Robot behavior

[0993] The robot receives the developments and reactions sent from the server. Based on the information it receives, it acts as a character in the story. The specific actions it takes vary depending on the development of the story.

[0994] Step 6: Viewing the Robot's Response

[0995] The robot displays responses on the device based on its own actions and reactions, including story developments and character reactions.

[0996] The above is the specific flow of the program's processing. The terminal receives the user's input, the server processes the information and generates developments and reactions, the robot acts based on that, and finally the robot's response is displayed on the terminal. This is how the story is created and the dialogue is carried out.

[0997] Example 1

[0998] Next, a description will be given of Example 1. In the following description, the data processing device 12 is referred to as a "server" and the robot 414 is referred to as a "terminal."

[0999] Conventional story generation systems have difficulty generating natural story developments and character reactions to user input in real time, and there is a lack of systems that can express them as real objects and characters. This limits interactive dialogue with the user, resulting in a poor quality experience.

[1000] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1001] In this invention, the server includes a means for receiving story suggestions and questions from a user, a means for generating a story development and a character's reaction using a generative AI model based on the received suggestions and questions, and a means for the robot to act based on the generated development and reaction. This enables the generation and expression of a dynamic and natural story development and character reaction in response to user input.

[1002] "User" refers to the person or user who inputs story suggestions and questions.

[1003] "Terminal" refers to an electronic device that provides an interface through which a user can enter story suggestions and questions and transmit them to a server.

[1004] The "server" refers to a computer system that receives suggestions and questions sent by users, uses a generative AI model to generate story developments and character responses, and transmits them to the robot.

[1005] A "generative AI model" is a model that uses artificial intelligence technology and is an algorithm that performs natural language processing based on user input to generate story development and character reactions.

[1006] A "prompt sentence" refers to a text sentence entered into a generative AI model to give it a specific instruction or request.

[1007] A "robot" refers to a physical device that acts as a character in a story based on the generation results sent from a server, performs actions and expressions, and displays its responses on a terminal.

[1008] "Story suggestions and questions" refer to instructions or questions that a user inputs via a terminal regarding the content of the story or the actions of a character and that are then transmitted to the server.

[1009] "Development and response" refers to the story progression and character responses that the generative AI model generates based on user suggestions and questions.

[1010] To implement this invention, hardware such as a server, terminal, and robot, as well as software such as a generative AI model, are required. This system generates story development and character reactions based on user input, and expresses them as real objects, thereby realizing interactive dialogue with the user.

[1011] System Configuration

[1012] 1. Server

[1013] Hardware: High-performance computer systems (e.g., a server rack with multiple processors, memory, storage, etc.)

[1014] Software: Generative AI models (e.g. ChatGPT by OpenAI), data processing algorithms, HTTP servers, etc.

[1015] Function: Receives story suggestions and questions sent by the user via the device, and uses a generative AI model to generate story development and character responses based on those suggestions. The generated data is then sent to the robot.

[1016] 2. Terminal

[1017] Hardware: User interface devices such as smartphones, tablets, and computers

[1018] Software: Web applications, mobile applications

[1019] Function: Provides an interface for users to enter story suggestions and questions, and sends the information to a server.

[1020] 3. Robots

[1021] Hardware: Physically shaped robots (e.g. humanoid robots, animal-like robots, etc.), sensors and actuators for movement and voice output

[1022] Software: Robot control program

[1023] Function: Acts and expresses itself based on the story development and character reactions sent from the server, and displays the results on the device.

[1024] How it works

[1025] When a user uses a device to input story suggestions or questions, the input data is sent from the device to a server. The server analyzes the received input data and uses a generative AI model to generate the story development and character reactions. The generated data is sent to the robot, which then operates based on that data. As a result, the robot behaves and speaks as a character in the story, and displays this on the device, realizing an interactive dialogue with the user.

[1026] Examples

[1027] For example, if a user uses a terminal to suggest, "Create a story where the main character goes on a journey into space," the process would go like this:

[1028] User: Type "Create a story where the main character goes on a space journey" into the terminal.

[1029] Terminal: Sends the entered information to the server.

[1030] Server: Sends the prompt "Create a story in which the protagonist goes on a journey into space" to the generative AI model and generates the story development.

[1031] Server: Sends the generated story development to the robot.

[1032] Robot: Performs movements that represent "the protagonist setting off into space."

[1033] Robot: Display the results on the terminal.

[1034] Also, if a user asks, "What happens if the main character meets an alien?"

[1035] User: Type into terminal, "What happens when the main character meets an alien?"

[1036] Terminal: Sends the entered information to the server.

[1037] Server: Sends this prompt to the generative AI model and generates the character's response.

[1038] Server: Sends the generated responses to the robot.

[1039] Robot: Performs actions that express the reaction of the protagonist when he meets the alien.

[1040] Robot: Display the results on the terminal.

[1041] In this manner, the present invention enables the generation and expression of dynamic and natural story development and character reactions to user input.

[1042] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1043] Step 1:

[1044] A user uses the terminal to input story suggestions and questions.

[1045] Specific operations: The user enters "Create a story in which the main character goes on a journey into space" into the input form displayed on the device screen.

[1046] Input: User suggestions or questions ("Create a story in which the main character travels into space")

[1047] Output: Input data that the device sends to the server

[1048] Step 2:

[1049] The terminal transmits the user's input information to the server.

[1050] Specific operation: The terminal sends the entered string data to the server as an HTTP request.

[1051] Input: String data entered by the user

[1052] Output: HTTP request to the server

[1053] Step 3:

[1054] The server receives and analyzes the input data sent from the terminal.

[1055] Specific operation: The server receives an HTTP request and extracts the user input contained in the request body.

[1056] Input: HTTP request from the terminal

[1057] Output: Extraction result of user input data

[1058] Step 4:

[1059] The server sends a prompt to the generative AI model based on the received input data.

[1060] Specific operation: The server generates a prompt such as "Create a story in which the protagonist goes on a journey into space" and sends it to the generative AI model as an API request.

[1061] Input: Data entered by the user

[1062] Output: API request to the generative AI model

[1063] Step 5:

[1064] The generative AI model generates a story development in response to the prompt sentence and returns the result to the server.

[1065] Specific operation: The generative AI model analyzes the prompt sentence, generates a story development, and returns it to the server as an API response.

[1066] Input: API request to the generative AI model (prompt text)

[1067] Output: Generated story development

[1068] Step 6:

[1069] The server receives and analyzes the story development returned by the generative AI model.

[1070] Specific operation: The server receives the API response returned from the generative AI model and parses the contents in JSON format.

[1071] Input: API response from the generative AI model

[1072] Output: Parsed story development

[1073] Step 7:

[1074] The server transmits the parsed story development to the robot.

[1075] Specific operation: The server sends the analysis results to the robot via a dedicated API endpoint.

[1076] Input: Parsed story arc

[1077] Output: API request to the robot

[1078] Step 8:

[1079] The robot acts based on the received story development and displays it to the user.

[1080] Specific behavior: The robot reproduces the movements of the "protagonist who sets off into space" and utters the relevant lines. In addition, it sends the generated story text to the terminal, which displays it on the screen.

[1081] Input: API request from the server (the story unfolds)

[1082] Output: Robot movement, display on terminal

[1083] (Application example 1)

[1084] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1085] In recent years, many story generation and interactive robot systems have appeared, but these systems can only provide static stories and have difficulty responding flexibly to users' interests and requests. In addition, conventional systems have limited story visualization, and more immersive experiences are required. Furthermore, the functionality that allows users to enjoy stories in 3D space using smart glasses or head-mounted displays has not been fully achieved. To address these challenges, there is a demand for more engaging and immersive story experiences that utilize user interaction.

[1086] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1087] In this invention, the server includes a means for receiving story suggestions and questions from a user, a means for generating a story development and a character's reaction based on the received suggestions and questions, a means for a robot to act based on the generated development and reaction, a means for displaying a response based on the robot's action and reaction, a means for transmitting the generated story to a user's terminal via a network and visualizing it in a 3D space, and a means for displaying the generated story using smart glasses or a head-mounted display. This allows the user to enjoy a dynamic and personalized story based on their suggestions and questions in a 3D space. Furthermore, by using smart glasses or a head-mounted display, a more realistic story experience can be provided, improving the user's sense of immersion.

[1088] The term "user" refers to a human subject who performs operations and inputs, and who makes story suggestions and asks questions to the system.

[1089] A "story suggestion or question" is a question about the story's plot or characters that a user provides to the system.

[1090] "Generative AI" refers to artificial intelligence that generates new story developments and character reactions based on previously learned data.

[1091] The "server" is a computer system that receives input information from users, uses generative AI to generate story development and character reactions, and sends the results to each terminal and robot.

[1092] A "robot" is a mechanical device that can act based on the story development and character reactions transmitted from the server.

[1093] A "terminal" is an electronic device that acts as an interface for users to input story suggestions and questions and that communicates with the server.

[1094] A "network" is a communications infrastructure for data communication between a server, a terminal, and a robot.

[1095] "Means of visualization in 3D space" refers to technology for presenting the generated story to the user as a three-dimensional stereoscopic image.

[1096] "Smart glasses" are glasses-type devices that are worn by a user to visually display the generated story in three-dimensional space.

[1097] A "head-mounted display" is a device that a user wears and can directly view three-dimensional images projected on a display in front of the user's eyes.

[1098] The "means for displaying responses" is a function for visually presenting the robot's actions and reactions to the user on the terminal.

[1099] "Story development" refers to the progression of events or scenes that occur within a story.

[1100] "Character response" refers to the actions or lines that characters in the story show in response to the user's suggestions or questions.

[1101] An embodiment of the present invention is a system including the following elements.

[1102] 1. Hardware and Software Configuration

[1103] The servers use Amazon Web Services (AWS) Lambda and EC2 servers.

[1104] The user's terminal is a smartphone or tablet, which provides an interface for suggesting stories and inputting questions.

[1105] The smart glasses use Google Glass or Microsoft HoloLens to visualize the story in 3D space.

[1106] The robots use programmable robots such as Pepper to perform story-based actions.

[1107] 2. Data processing and data calculation processing

[1108] User input:

[1109] Users enter questions about the story's plot and characters on their smartphone or tablet, and this data is sent to the server via an HTTP POST request.

[1110] Example: User: "Make a story where the protagonist saves a medieval kingdom."

[1111] On the server:

[1112] The server receives the user's input data and inputs it as a prompt sentence to a generative AI model such as ChatGPT. The generative AI model generates the development of the story based on the received prompt sentence.

[1113] Example prompt: "User input: A story about saving a medieval kingdom. Generated story:"

[1114] Delivery and visualization of generative narratives:

[1115] The generated story is sent over the network to the user's device, where it is visualized using a 3D rendering engine (e.g., Unity3D).

[1116] At the same time, operating instructions are sent to the robot, and the robot acts based on the story.

[1117] Smart Glasses and Head Mounted Displays Applications:

[1118] The generated story is also sent to smart glasses or a head-mounted display, through which the story is displayed to the user in 3D space.

[1119] 3. Specific application examples

[1120] For example, if a user suggests "create a story in which the protagonist saves a medieval kingdom," the server will generate a storyline based on the suggestion, send the generated story to the user's device, and visualize it in 3D space. The robot will also act out the story.

[1121] If the user then asks, "What happens if the main character meets a dragon?", the server uses generative AI to generate a dragon encounter scene in response to the question. The robot acts out the scene, which is then visually displayed on the smart glasses or head-mounted display.

[1122] The above is a specific embodiment for carrying out the present invention. This configuration enables the user to enjoy a dynamic and realistic story that responds to the user's input.

[1123] The flow of the specific process in the application example 1 will be described with reference to FIG.

[1124] Step 1:

[1125] Users enter suggestions or questions about the story's plot or characters on their smartphone or tablet, and that input (e.g., "Create a story where the protagonist saves a medieval kingdom") is sent to the server through the application's interface as an HTTP POST request.

[1126] Step 2:

[1127] The server receives requests sent by users. Based on the received input data, the server generates a prompt to be given to a generative AI model (e.g. ChatGPT).

[1128] Example: Prompt: "User input: A story about saving a medieval kingdom. Generated story:"

[1129] Step 3:

[1130] The server uses a generative AI model to generate story development and character responses based on the prompt text. The result of this process is a new story text. The generated story text is stored on the server.

[1131] Input: prompt statement

[1132] Output: Generated narrative text

[1133] Step 4:

[1134] The server transmits the generated story text to the robot, and at the same time, the story text is also transmitted to the user's terminal via the network.

[1135] Input: Generated narrative text

[1136] Output: Data sent to the robot and the user's device

[1137] Step 5:

[1138] The user's device passes the received story text to a 3D rendering engine (e.g. Unity3D), which visualizes the story as a 3D space on the device and displays it to the user.

[1139] Input: Generated narrative text

[1140] Output: A narrative visualized in 3D space.

[1141] Step 6:

[1142] The robot acts based on the story text sent from the server. The robot executes the specified scenario and acts out the story with actions and sounds.

[1143] Input: Generated narrative text

[1144] Output: Robot behavior and audio output

[1145] Step 7:

[1146] If the user is wearing smart glasses or a head-mounted display, the device transmits the story text to these devices and displays it visually in 3D space, allowing the user to experience the generated story in a virtual reality environment.

[1147] Input: Generated narrative text

[1148] Output: Visualization on smart glasses or head-mounted displays

[1149] This is the process flow of the system. Through this series of steps, users can enjoy a dynamic and immersive story experience based on their own input.

[1150] In addition, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1151] An embodiment for implementing the present invention includes the following elements.

[1152] 1. Server

[1153] A server is provided that executes a program for generating story development and character reactions.

[1154] It is combined with an emotion engine that recognizes the user's emotions and performs voice and image analysis to recognize the user's emotions.

[1155] The results of the generative AI and emotion engine are combined to generate story development and character reactions.

[1156] 2. Terminal

[1157] It provides an interface for users to make story suggestions and questions.

[1158] The user's input information is sent to the server.

[1159] Voice and image data is sent to the server to recognize the user's emotions.

[1160] 3. Robot

[1161] Act on developments and reactions sent from the server.

[1162] As a character in the story, it will act and express itself according to the generated developments and reactions.

[1163] The response is displayed on the terminal, enabling a dialogue with the user.

[1164] (Specific examples)

[1165] For example, if a user suggests through a device, "Create a story in which the main character goes on a journey into space," the device will send that information to the server. Based on the suggestion, the server will use generative AI and an emotion engine to generate a plot related to space travel. The generated plot will be sent to the robot, which will then act out the journey into space. At the same time, the emotion engine will recognize emotions from the user's speech and facial expressions, and adjust the robot's behavior and reactions based on those emotions.

[1166] Also, if the user asks, "What would happen if the main character met an alien?", the device sends that information to the server. The server uses generative AI and an emotion engine to generate the character's reaction to meeting an alien. The robot acts based on that reaction and expresses its reaction to meeting an alien. At the same time, the emotion engine recognizes the user's emotions and adjusts the robot's response to match those emotions.

[1167] The above is an example of an embodiment of the present invention. Story creation and dialogue combining an emotion engine are specifically carried out by the roles and cooperation of the server, terminal, and robot.

[1168] The process flow will be explained below.

[1169] Step 1: Receiving User Input

[1170] The terminal provides an interface for the user to suggest stories and ask questions, and receives user-input information (text, voice, images, etc.).

[1171] Step 2: Sending input information and emotion recognition

[1172] The terminal transmits the received user input information to the server.

[1173] The server uses an emotion engine to recognize the user's emotions based on the received information.

[1174] The emotion engine analyzes audio and images to extract the user's emotions.

[1175] Step 3: Receiving and processing information

[1176] The server receives story suggestions and questions sent from the terminal, as well as emotion information from the emotion engine.

[1177] Based on the information received, generative AI is used to generate story development and character reactions.

[1178] By combining the results of generative AI and the emotion engine, we can generate more flexible, emotionally-responsive developments and responses.

[1179] Step 4: Development and reaction generation

[1180] The server combines the results of the generative AI and the emotion engine to generate story developments and character reactions based on the information it receives, which are then sent to the robot in the next step.

[1181] Step 5: Robot behavior

[1182] The robot receives the developments and reactions sent from the server. Based on the information it receives, it acts as a character in the story. The specific actions it takes vary depending on the development of the story and the character's reactions.

[1183] Step 6: Viewing the Robot's Response

[1184] The robot displays responses on the terminal based on its own actions and reactions. Responses include the development of the story and the reactions of the characters. At the same time, it uses emotional information from the emotion engine to adjust the robot's facial expressions and tone of voice to respond in accordance with the emotions.

[1185] The above is the specific flow of the program's processing. The terminal receives the user's input, the server processes the information and generates developments and reactions, the robot acts based on that, and finally the robot's response is displayed on the terminal. This is the flow of events that combines the emotion engine to create a story and carry out a dialogue.

[1186] Example 2

[1187] Next, a description will be given of Example 2. In the following description, the data processing device 12 is referred to as a "server" and the robot 414 is referred to as a "terminal."

[1188] Conventional story creation and dialogue systems can generate stories based on user suggestions and questions, but they have the problem of lacking naturalness and realism in dialogue because they do not take the user's emotions into account. In addition, the actions and emotional expressions of characters that respond in real time to story suggestions and questions input by the user are insufficient, making it difficult to increase user satisfaction.

[1189] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for receiving a story suggestion or question from a user, a means for generating a story development or a character's reaction using a natural language processing engine based on the received suggestion or question, a means for the robot to act based on the generated development and reaction, a means for displaying the robot's behavior and realizing a dialogue with the user, an emotion engine for recognizing the emotion by analyzing the user's voice or image, and a means for adjusting the robot's behavior based on the result of the emotion engine. This makes it possible to realize a story generation and a dialogue in real time that takes the user's emotion into consideration.

[1190] "User" refers to a person or institution that inputs a story suggestion or question.

[1191] "Terminal" refers to a device through which a user inputs story suggestions and questions and transmits them to a server.

[1192] The "server" refers to a central processing unit that generates story development and character reactions based on information received from the user, as well as analyzes emotions.

[1193] A "generative AI model" refers to an artificial intelligence model that generates story development and character responses based on user suggestions and questions.

[1194] A "natural language processing engine" refers to software that analyzes natural language entered by the user and generates appropriate story development and character responses.

[1195] An "emotion engine" refers to software that analyzes the user's voice and image data to recognize emotions, and adjusts the character's response based on that information.

[1196] A "robot" refers to a mechanical device that physically moves and expresses itself based on the story development and character reactions transmitted from the server.

[1197] "Story suggestions" refer to instructions or ideas about the story scenario or development that a user inputs via a terminal.

[1198] "Narrative development" refers to a series of storylines or scenarios that a generative AI model generates based on user suggestions.

[1199] "Character reactions" refers to the actions and expressions of characters that are generated according to the development of the story.

[1200] "Action" refers to the physical movements the robot makes based on the story development and character reactions.

[1201] This invention is a system for creating stories and conducting dialogue, and is mainly composed of a server, a terminal, and a robot.

[1202] Specifically, a user inputs story suggestions and questions via a terminal. The terminal is provided with an interface for inputting suggestions and questions, and is equipped with a microphone and a camera for capturing the user's voice and image. When the user inputs a prompt such as "Create a story in which the main character travels into space," the information is sent to the server.

[1203] The server uses several software components to generate storylines and character responses: First, it uses a generative AI model (for example, a system known as a natural language generation engine) to generate storylines based on the user's prompts. The generative AI model could be OpenAI's ChatGPT or another natural language processing engine.

[1204] Next, the server uses an emotion engine (such as Microsoft's Azure Emotion API or Google's Cloud Vision API) to analyze the user's voice and image data. This allows the server to recognize the user's emotional state and reflect it in the development of the generated story. As a result of the analysis, the emotion engine provides emotional information such as whether the user is happy, surprised, or angry.

[1205] The server combines the story development results of the generative AI model with the analysis results of the emotion engine to generate the final story development and character reactions. This generation result is encoded in an appropriate format such as JSON or XML and sent to the robot.

[1206] The robot performs physical movements and outputs voice based on the generated results. For example, the robot imitates a rocket launch in a space travel scene. It also uses a voice synthesis engine to speak the characters' lines. In this way, it provides the user with a more realistic storytelling experience.

[1207] As a concrete example, consider the case where a user asks, "What would happen if the main character met an alien?" The device sends this question to a server, which uses a generative AI model to generate plot developments and character reactions. At the same time, an emotion engine analyzes the user's voice and facial expressions to identify the user's emotional state. Based on the results, the robot expresses its reaction to the encounter with the alien and communicates this to the user through voice and action.

[1208] As described above, the present invention provides a system for realizing story generation and real-time dialogue that takes into account the user's emotions.

[1209] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1210] Step 1:

[1211] The user inputs story suggestions and questions into the device. Using the device's interface, the user inputs prompts such as "Create a story in which the main character travels into space." The user can also input voice and images. The input text, voice, and image data are captured by the device.

[1212] Step 2:

[1213] The terminal sends the user's input information and voice and image data to the server. The terminal divides the input data into packets and transfers them to the server using a security protocol (e.g. HTTPS, SSL / TLS). The input data includes the text entered by the user, the captured voice data, and image data. The server receives this data.

[1214] Step 3:

[1215] The server generates a story development using a generative AI model. The server inputs the text suggestions received from the user into a generative AI model (e.g., a natural language generation engine). For example, to generate a "story about going on a journey into space," the generative AI model is sent a prompt sentence such as "Make a story about the protagonist going on a journey into space." The generative AI model generates a story development based on this prompt sentence and returns the result to the server. The generated story development is output to the server.

[1216] Step 4:

[1217] The server analyzes the user's emotions using an emotion engine. The server inputs the received voice and image data into the emotion engine. The emotion engine (e.g., a voice tone analysis engine or a facial expression analysis engine) analyzes the voice tone and facial expression to identify the user's emotional state (e.g., happiness, surprise, anger). The analysis results of the emotion engine are output to the server.

[1218] Step 5:

[1219] The server integrates the results of the generative AI model and the emotion engine to generate story development and character reactions. The server integrates the story development obtained from the generative AI model with the user's emotion information obtained from the emotion engine. Based on the results of this integration, it generates the reactions of the story characters and further story content. The generated character reactions and story development are output to the server.

[1220] Step 6:

[1221] The server sends the generated results to the robot. The server converts the story development and character reactions generated by the server into an appropriate format (e.g. JSON, XML) and sends it to the robot. The robot receives the data sent from the server.

[1222] Step 7:

[1223] The robot acts and expresses itself based on the content it receives. The robot performs physical movements and outputs voice based on the development of the story it receives and the reactions of the characters. For example, in a scene where the robot departs into space, the robot looks up at the sky and moves its hands to mimic a rocket launch sequence. It also uses a voice synthesis engine to vocalize the characters' lines. This provides a story experience for the user.

[1224] Step 8:

[1225] The terminal displays the robot's actions and enables dialogue with the user. Based on the data sent from the robot, the terminal displays the robot's actions and speech. The user can visually and audibly confirm the robot's actions and speech through the terminal. This allows the user to make further story suggestions or ask questions, continuing the interaction with the system.

[1226] (Application example 2)

[1227] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal".

[1228] Conventional story creation systems and dialogue systems have been unable to generate story developments and character reactions that reflect the user's emotions in real time, limiting the user experience. The present invention aims to provide a more interactive and personalized story experience by recognizing the user's emotions and dynamically adjusting the story development and character reactions based on those emotions.

[1229] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means. In this invention, the server includes a means for recognizing the user's emotions using an emotion engine that recognizes the user's emotions based on the user's voice and image, a means for receiving a story proposal or question from the user, a prompt sentence that instructs generating a story development or a character's reaction based on the received story proposal or question and the user's emotions, a means for generating the story development or the character's reaction using a generative AI, and a means for controlling the robot's behavior based on the generated story development or character reaction. This makes it possible to provide a more interactive and personalized story experience by dynamically adjusting the story development or the character's reaction by recognizing the emotion when the user makes a story proposal or question. In addition, the server further includes a means for displaying the behavior of the robot on the user's terminal, and repeats recognizing the user's emotions, receiving a story proposal or question from the user, generating the story development or the character's reaction, controlling the robot's behavior, and displaying on the user's terminal. The means for controlling the behavior of the robot controls the robot to perform actions that imitate the reactions of the character, and controls the robot to output the lines of the character as voice.

[1230] "Recognizing a user's emotions" means analyzing input audio and image data and identifying the user's emotional state (e.g., joy, sadness, surprise, etc.).

[1231] "Story suggestions and questions" refers to a user requesting content related to the progression of the story or asking a question about the development of the story.

[1232] "Generating story development and character reactions" refers to using generative AI models and emotion engines to create story progression and character actions and reactions based on user suggestions and questions.

[1233] "The robot acts" refers to the robot physically moving and carrying out set actions based on the story development and character reactions sent from the server.

[1234] "Displaying responses" means visually presenting the generated story development and the character's reactions on the terminal screen or the robot's display.

[1235] "Server" refers to a computer system that receives user input and executes programs to generate story development and character responses using a generative AI model and emotion engine.

[1236] A "generative AI model" refers to an artificial intelligence algorithm that generates story development and character responses based on user suggestions and questions.

[1237] An "emotion engine" is software or hardware that analyzes a user's voice and image data and recognizes emotions.

[1238] "Dynamic adjustment" refers to changing the story content and character reactions in real time based on the user's emotional recognition.

[1239] "Interactive and personalized narrative experience" means that the story progresses in response to user input and emotional state, providing a customized storytelling experience for each individual user.

[1240] As an embodiment of the present invention, a system for generating story development and character reactions includes the following elements: A user inputs story suggestions and questions through a terminal, and the server uses a generative AI model and an emotion engine based on that information to generate story development and character reactions. The generated development and reactions are sent to a robot, which operates based on the content and displays a response.

[1241] Specifically, the server is equipped with an emotion engine for analyzing the user's voice and image data, and a generative AI model that generates the story development and character reactions. When a user suggests through their device, "Make a story in which the main character goes on a journey into space," the device sends that information to the server. The server uses the emotion engine to recognize emotions from the user's voice and images. The server passes the recognized emotion data and the user's suggestions to the generative AI model to generate the story development.

[1242] The generated story development and character reactions are sent to the robot, and the robot acts accordingly. For example, if a user asks, "What happens if the main character meets an alien?", the server uses the information to generate a story development and character reactions regarding the encounter with the alien, using an emotion engine and generative AI model. The robot acts based on this reaction and displays the response on the terminal.

[1243] Examples of specific hardware and software used in servers, terminals, and robots include the emotion engine "EmotionEngine" and the generative AI model "StoryGenAI." This makes it possible to provide an interactive storytelling experience that dynamically adjusts the development of the story and the reactions of characters based on the user's emotions.

[1244] For example, if a user requests, "Make a story where the main character goes on a journey into space," the server will analyze the user's emotions and generate a positive development using a generative AI model. In addition, if a user asks, "What would happen if the main character met an alien?", the server will generate a character's reaction based on the emotional data, and the robot will act accordingly.

[1245] Example of a prompt that generates a storyline based on the story suggestions and the user's emotions: "The user requested, 'Make a story where the main character goes on a journey into space.' The user's emotions are bright and positive, so please generate a positive development accordingly."

[1246] An example prompt that generates a character's response based on the question and the user's emotions: "The user asked, 'What would happen if the main character met an alien?' The user's emotions are bright and positive, so generate a character reaction based on that emotion."

[1247] The flow of the specific process in the application example 2 will be described with reference to FIG.

[1248] Step 1:

[1249] The user inputs story suggestions and questions through the terminal. The input data includes the user's voice, text, and image data. For example, a request might be "Make a story where the main character goes on a journey into space." This input data is sent directly to the terminal.

[1250] Step 2:

[1251] The device sends the user's input information to the server. This includes voice and image data, and is sent in JSON format formatted for processing on the server side. The input data includes the user's request and voice and image data. The data received as input to the server is analyzed by the emotion engine.

[1252] Step 3:

[1253] The server uses the emotion engine to recognize the user's emotions from the received voice and image data. In this step, it performs voice and image analysis to convert the user's emotions into concrete tags (e.g., happiness, sadness, surprise, etc.). It obtains the emotion recognition results and prepares them to be passed to the generative AI model.

[1254] Step 4:

[1255] The server uses a generative AI model to generate a story development and character reactions based on the user's suggestions and questions. In this step, the user's input text and emotion recognition results are input as prompts to the generative AI model to obtain a generated story development and character reactions. Specifically, a story is generated based on the request "The protagonist goes on a space journey."

[1256] Step 5:

[1257] The server sends the generated story development and character reaction data to the robot. The output includes the story content and character actions, which are converted into robot movement instructions. The robot receives this data and starts moving according to the content.

[1258] Step 6:

[1259] The robot physically moves based on the received story development and the character's reaction. For example, the robot acts out a scene in which it travels into space. The robot's actions are displayed on the user's device. Specific examples of the actions include the robot making a specific gesture or speaking the character's lines.

[1260] Step 7:

[1261] If the user makes a new question or suggestion, the device again sends that information to the server, and the following steps 2 to 6 are executed in a loop. This cycle dynamically adjusts the development of the story, realizing an interactive story experience that responds to the user's emotions. For example, in response to the question, "What would happen if the main character met an alien?", the server generates a new development and the robot behaves based on it.

[1262] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires a voice indicating a user input for the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1263] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1264] In the above embodiment, an example was given in which the specific process was performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the robot 414.

[1265] The emotion identification model 59 as an emotion engine may determine the emotion of the user according to a specific mapping. Specifically, the emotion identification model 59 may determine the emotion of the user according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the emotion of the robot, and the identification processing unit 290 may perform identification processing using the emotion of the robot.

[1266] FIG. 9 is a diagram showing an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive emotions are arranged. The more outside the concentric circles, the more emotions that represent states and actions that arise from a state of mind are arranged. Emotions are a concept that includes emotions and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions that occur in the brain are arranged. On the right side of the concentric circles, emotions that are generally induced by situational judgment are arranged. On the upper and lower sides of the concentric circles, emotions that are generally generated from reactions that occur in the brain and are induced by situational judgment are arranged. In addition, on the upper side of the concentric circles, emotions of "pleasure" are arranged, and on the lower side, emotions of "discomfort" are arranged. In this way, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1267] These emotions are distributed in the 3 o'clock direction of emotion map 400 and usually fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1268] The inside of emotion map 400 represents what is going on inside one's mind, and the outside of emotion map 400 represents behavior, so the further out you go on emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1269] Here, human emotions are based on various balances such as posture and blood sugar level, and when these balances are far from the ideal, it indicates an unpleasant state, and when they are close to the ideal, it indicates a pleasant state. Emotions can also be created for robots, cars, motorcycles, etc., based on various balances such as posture and battery level, so that when these balances are far from the ideal, it indicates an unpleasant state, and when they are close to the ideal, it indicates a pleasant state. The emotion map may be generated, for example, based on the emotion map of Dr. Mitsuyoshi (Research on speech emotion recognition and emotion brain physiological signal analysis system, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). On the left half of the emotion map, emotions belonging to an area called "reaction" where sensation is dominant are lined up. On the right half of the emotion map, emotions belonging to an area called "situation" where situation recognition is dominant are lined up.

[1270] The emotion map defines two emotions that promote learning. The first is the negative emotion around the middle of "repentance" or "remorse" on the situation side. In other words, this is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the positive emotion around "desire" on the response side. In other words, this is when the robot has positive feelings such as "I want more" or "I want to know more."

[1271] The emotion identification model 59 inputs the user input to a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the emotion of the user. This neural network is pre-trained based on multiple learning data that are combinations of the user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in Fig. 10. Fig. 10 shows an example in which multiple emotions, "relief," "calm," and "encouraging," have similar emotion values.

[1272] Although the system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, the system according to the present disclosure is not necessarily implemented in a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program that runs on a personal computer, or an application that runs on a smartphone or the like. The method according to the present disclosure may be provided to a user in the form of SaaS (Software as a Service).

[1273] In the above embodiment, an example is given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to input data.

[1274] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable computer-readable non-transitory storage medium such as a Universal Serial Bus (USB) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1275] In addition, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 upon request from the data processing device 12.

[1276] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1277] As the hardware resource for executing the specific process, various processors as shown below can be used. An example of the processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing the specific process by executing software, i.e., a program. Another example of the processor is a dedicated electric circuit, which is a processor having a circuit configuration designed exclusively for executing the specific process, such as a Field-Programmable Gate Array (FPGA), a Programmable Logic Device (PLD), or an Application Specific Integrated Circuit (ASIC). Each processor has a built-in or connected memory, and each processor executes the specific process by using the memory.

[1278] The hardware resource that executes the specific process may be one of these various processors, or may be a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific process may be a single processor.

[1279] As an example of a configuration using one processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a configuration using a processor that realizes the functions of the entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1280] Furthermore, more specifically, the hardware structure of these various processors can be an electric circuit that combines circuit elements such as semiconductor elements. The specific processes described above are merely examples. It goes without saying that unnecessary steps may be deleted, new steps may be added, or the order of processes may be changed without departing from the spirit of the invention.

[1281] The above description and illustrations are detailed descriptions of the parts related to the technology of the present disclosure, and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, function, action, and effect is an example of the configuration, function, action, and effect of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above description and illustrations, within the scope of the gist of the technology of the present disclosure. In addition, in order to avoid confusion and to facilitate understanding of the parts related to the technology of the present disclosure, the above description and illustrations omit explanations of technical common sense that do not require explanation in order to enable the implementation of the technology of the present disclosure.

[1282] All publications, patent applications, and standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or standard was specifically and individually indicated to be incorporated by reference.

[1283] The following is further disclosed regarding the above embodiment.

[1284] (Claim 1)

[1285] A robot application system for creating stories and conducting dialogue, comprising: a means for receiving story suggestions and questions from users; A means of generating story developments and character responses based on suggestions and questions received; A means for the robot to act based on the generated developments and reactions; and and means for displaying a response based on the robot's actions and reactions.

[1286] (Claim 2)

[1287] 2. The system of claim 1, A means for receiving story suggestions and questions from users via the terminal; A means to generate developments and responses using generative AI based on the suggestions and questions received by the server; A means for the robot to act based on the deployment and responses sent from the server; and and means for the robot to display the response on a terminal.

[1288] (Claim 3)

[1289] A system according to claim 1 or 2, A means for a user to input story suggestions or questions via the terminal; A means for the server to generate developments and reactions based on the information it receives using generative AI; A means for the robot to act based on the generated deployments and responses; and and a means for including the development of the story and the character's reaction when the robot displays the response on the terminal.

[1290] (Claim 4)

[1291] 2. The system of claim 1 further comprising an emotion engine for recognizing user emotions.

[1292] (Claim 5)

[1293] The system according to claim 1, characterized in that the emotion engine recognizes emotions from the user's speech, facial expressions, etc., and generates story development and character reactions based on the emotions.

[1294] (Claim 6)

[1295] 6. The system according to claim 1, wherein the emotion engine analyzes voice and images to recognize the user's emotion and transmits the results to the server.

[1296] "Example 1"

[1297] (Claim 1)

[1298] a means for receiving story suggestions and questions from users; A means to use generative AI models to generate story developments and character responses based on suggestions and questions received; and A means for the robot to act based on the generated developments and reactions; and A means for displaying a response on the terminal based on the actions and reactions of the robot; A system including:

[1299] (Claim 2)

[1300] a means for receiving story suggestions and questions from a user via the terminal and transmitting the same to a server; A means for generating a prompt sentence using a generative AI model based on the suggestions and questions received by the server, and acquiring the generated result; means for transmitting the generated results to a robot and for the robot to act on the results; A means for the robot to display the response on a terminal; 2. The system of claim 1, comprising:

[1301] (Claim 3)

[1302] A means for a user to input via a terminal and transmit the input to a server; A means for the server to generate prompts based on the received information using a generative AI model, and generate expansions and responses; A means for the robot to act based on the generated deployment and reaction and display the response on a terminal; a means for displaying responses including story developments and character reactions; 2. The system of claim 1, comprising:

[1303] "Application example 1"

[1304] (Claim 1)

[1305] a means for receiving story suggestions and questions from users; A means of generating story developments and character responses based on suggestions and questions received; A means for the robot to act based on the generated developments and reactions; and a means for displaying a response based on the robot's actions and reactions; A means for transmitting the generated story to a user's terminal via a network and visualizing it in 3D space; A system including a means for displaying the generated story using smart glasses or a head-mounted display.

[1306] (Claim 2)

[1307] A means for receiving story suggestions and questions from users via the terminal; A means to generate developments and responses using generative AI based on the suggestions and questions received by the server; A means for the robot to act based on the deployment and responses sent from the server; and A means for transmitting the generated story to a user's terminal via a network and visualizing it in 3D space; 13. The system of claim 1, further comprising means for displaying the generated narrative using smart glasses or a head mounted display.

[1308] (Claim 3)

[1309] A means for a user to input story suggestions or questions via the terminal; A means for the server to generate developments and reactions based on the information it receives using generative AI; A means for the robot to act based on the generated deployments and responses; and A means for transmitting the generated story to a user's terminal via a network and visualizing it in 3D space; 3. The system of claim 1 or 2, further comprising means for displaying the generated story development and character reactions using smart glasses or a head mounted display.

[1310] "Example 2 of combining emotion engines"

[1311] (Claim 1)

[1312] a means for receiving story suggestions and questions from users; A means to generate story developments and character responses using a natural language processing engine based on suggestions and questions received; and A means for the robot to act based on the generated developments and reactions; and A means for displaying the robot's actions and realizing a dialogue with a user; An emotion engine that recognizes emotions by analyzing the user's voice and images; means for adjusting the behavior of the robot based on the results of the emotion engine; A system including:

[1313] (Claim 2)

[1314] a means for a user to input story suggestions and questions via the terminal; A means for the terminal to transmit user input information and voice and image data to a server; The server uses a generative AI model to generate story development and character reactions, A means for the server to analyze the user's emotions using an emotion engine; A means of integrating the results of the generative AI model and the emotion engine to generate story development and character reactions; and A means for transmitting the generated result to a robot; A means for the robot to act based on the generated deployment and reactions and provide the action to the user visually and audibly; 2. The system of claim 1, comprising:

[1315] (Claim 3)

[1316] A means for the robot to act based on the development of the story and the reactions of the characters transmitted from the server, and to display them on the terminal and realize a dialogue with the user; An emotion engine that recognizes the user's emotional state and adjusts the robot's behavior accordingly; 2. The system of claim 1, comprising:

[1317] "Application example 2 when combining emotion engines"

[1318] (Claim 1)

[1319] A means for performing voice and image analysis to recognize the user's emotions; a means for receiving story suggestions and questions from users; A means of generating story developments and character responses based on suggestions and questions received; A means for the robot to act based on the generated developments and reactions; and a means for displaying a response based on the robot's actions and reactions; A means for dynamically adjusting the development of a story and the reactions of characters based on the user's emotions; A method using a generative AI model combined with an emotion engine that recognizes the user's emotions; A system including:

[1320] (Claim 2)

[1321] A means for receiving story suggestions and questions from users via the terminal; A means to generate developments and responses using generative AI based on the suggestions and questions received by the server; A means for the robot to act based on the deployment and responses sent from the server; and A means for the robot to display the response on a terminal; A means for recognizing emotions through a user's voice and image; A means for dynamically adjusting the development of the story based on the recognition of the user's emotions; 2. The system of claim 1.

[1322] (Claim 3)

[1323] A means for a user to input story suggestions or questions via the terminal; A means for the server to generate developments and reactions based on the information it receives using generative AI; A means for the robot to act based on the generated deployments and responses; and A means for the robot to display responses on a terminal, including the development of the story and the reactions of the characters; A means for transmitting analysis data for recognizing an emotion of a user; 2. The system of claim 1. [Explanation of symbols]

[1324] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for recognizing a user's emotion using an emotion engine for recognizing the emotion of the user based on the user's voice and image; means for receiving story suggestions or questions from a user; A prompt sentence instructing the generation of a story development or a character's reaction based on the received story suggestion or question and the user's emotion, and a generative AI to generate the story development or the character's reaction; A means for controlling the behavior of a robot based on the generated development of the story or the reaction of a character; A system including:

2. The method further includes: displaying the behavior of the robot on a terminal of the user; The system of claim 1 , which repeats the steps of recognizing the user's emotions, receiving story suggestions or questions from the user, generating the story development or character's responses, controlling the robot's behavior, and displaying on the user's terminal.

3. 2. The system according to claim 1, wherein the means for controlling the behavior of the robot controls the robot to perform actions that imitate the reactions of the character and controls the robot to output the lines of the character as voice.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A