System

The system addresses the challenge of providing emotional education by automating the generation and display of stories and images on electronic paper, ensuring an easy-to-use interface for parents to support children's emotional and interpersonal development.

JP2026025728APending Publication Date: 2026-02-16SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024128540
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-02
Publication Date
2026-02-16

AI Technical Summary

Technical Problem

Parents find it difficult to provide appropriate emotional education for their children due to a lack of suitable picture books that support emotional development and interpersonal relationships, especially in busy daily lives, and existing systems lack an easy-to-use interface for generating and displaying such content on devices that are gentle on children's eyes.

Method used

A system that includes a selection means for choosing emotional areas, a story generation means using generative AI, an image generation means, and a display means on electronic paper, allowing for the automatic creation and gentle visual presentation of stories and images tailored to emotional education.

Benefits of technology

Facilitates easy provision of emotionally educational content for children through an intuitive interface, supporting emotional and interpersonal development with a visual experience that is gentle on children's eyes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026025728000001_ABST
    Figure 2026025728000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: generating means for generating a story for each emotion; image generating means for generating an image based on the generated story; and display means for displaying the generated story and image.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Today's parents find it difficult to find picture books that are suitable for their children and to provide appropriate emotional education in the midst of their busy daily lives. Furthermore, there are not enough picture books that support the development of emotions and interpersonal relationships at the appropriate time as children grow. Therefore, a system is needed that automatically generates picture books useful for emotional education and supports the development of children's emotions and interpersonal relationships through reading aloud. [Means for solving the problem]

[0005] The present invention solves this problem by providing a system including a selection means for selecting an emotional area based on a user's input, a story generation means for generating a story based on the selected emotional area, an image generation means for generating images based on the generated story, and a display means for displaying the generated story and images. Furthermore, the display means uses electronic paper without a backlight, providing a visual experience that is gentle on children's eyes. This allows parents to easily provide emotional education that is appropriate for their children.

[0006] "Generation means" refers to a device or program for generating stories for each emotion.

[0007] "Image generation means" refers to a device or program that generates appropriate images based on the generated story.

[0008] A "display means" is a device for displaying the generated story and images to a user.

[0009] The "selection means" is a device or program that selects a specific emotional area based on a user's input.

[0010] The "story generation means" is a device or program that automatically generates a story based on the selected emotional domain.

[0011] "Electronic paper" is a technology used as a display means without backlighting, providing a visual experience that is easy on the eyes.

[0012] The "emotional domain" refers to the emotional education theme selected by the user, such as self-development, empathy, cooperation, etc. [Brief explanation of the drawings]

[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2]1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0021] [First embodiment]

[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0034] This invention is a system that automatically generates picture books for children and supports the development of their emotions and interpersonal relationships through reading aloud, with the aim of emotional education. The system provides an interface that is easy for users to operate, and generates and displays stories and images based on emotional domains.

[0035] System Overview

[0036] 1. User operations

[0037] The user launches the application and selects an emotional area.

[0038] Example: A user selects "empathy" within an application.

[0039] The terminal transmits the user's selection data to the server.

[0040] 2. Story generation based on emotional domains

[0041] The server analyzes the emotion area data received from the device.

[0042] Example: The server checks that the emotion domain is "empathy."

[0043] Emotional domains are input into the generative AI model on the server to generate a story.

[0044] Example: An AI model generates a story about a little rabbit character named Ravi and his friend Nico, with a theme of empathy.

[0045] Some of the stories generated:

[0046] One day, Ravi, a little rabbit, finds his friend Niko sad, and he wants to know why Niko is sad.

[0047] 3. Narrative-based image generation

[0048] The server generates image prompts based on the generated story.

[0049] Example: Generate a prompt for "A scene with Ravi and Nico together."

[0050] The server passes the prompt to an image generation AI model, which generates an appropriate image.

[0051] Example: An image generation AI generates an image depicting a scene in which "Ravi and Nico are sitting together."

[0052] 4. Data transmission and display

[0053] The server compiles the generated story and images and sends them to the device.

[0054] Example: Sending story text and image files to a device.

[0055] The terminal analyzes the received data and displays it on the electronic paper device.

[0056] Example: An e-paper device displays "a story and corresponding images that show the relatable nature of Ravi and Nico."

[0057] 5. Reading aloud

[0058] A user reads a story to a child using an e-paper device.

[0059] Example: A parent reads part of a story: "Ravi looks at his friend Niko and realizes he is sad..."

[0060] The user interacts with the child to discuss the content of the story and develop the child's emotions.

[0061] Example: "How did Ravi know Nico was sad?" helps children develop an understanding of emotions.

[0062] This system allows parents to easily create picture books suitable for emotional education and read them to their children. By using an e-paper device, it provides a visual experience that is gentle on children's eyes while supporting the development of emotions and interpersonal relationships.

[0063] The processing flow will be explained below.

[0064] Step 1:

[0065] The user launches the application and selects an emotional area.

[0066] Example: A user selects "empathy" within an application.

[0067] Step 2:

[0068] The terminal transmits the user's selection data to the server.

[0069] Example: Send data from the emotion field "empathy" to the server.

[0070] Step 3:

[0071] The server analyzes the emotion area data received from the device.

[0072] Example: The server checks that the emotion domain is "empathy."

[0073] Step 4:

[0074] The emotional domain of "empathy" is input into the generative AI model on the server to generate a story.

[0075] Example: An AI model generates stories with empathy as the theme, using a rabbit character named "Rabi."

[0076] Some of the stories generated:

[0077] One day, Ravi, a little rabbit, finds his friend Niko sad, and he wants to know why Niko is sad.

[0078] Step 5:

[0079] The server generates image prompts based on the generated story.

[0080] Example: Generate a prompt for "A scene with Ravi and Nico together."

[0081] Step 6:

[0082] The server passes the prompt to an image generation AI model, which generates an appropriate image.

[0083] Example: An image generation AI generates an image of a scene where "Ravi and Nico are sitting together."

[0084] Step 7:

[0085] The server compiles the generated story and images into a single document.

[0086] Example: Combining generated narrative text and images.

[0087] Examples of compiled data:

[0088] Story text:

[0089] One day, a little rabbit named Ravi found his friend Niko sad.

[0090] Story Image:

[0091] Image file of Ravi and Nico sitting together

[0092] Step 8:

[0093] The server sends the compiled data to the terminal.

[0094] Example: Sending a story and image file to a device.

[0095] Step 9:

[0096] The terminal analyzes the received data and displays it on the electronic paper device.

[0097] Example: An e-paper device displays "a story and corresponding images that show the relatable nature of Ravi and Nico."

[0098] Step 10:

[0099] A user reads a story to a child using an e-paper device.

[0100] Example: A parent reads part of a story: "Ravi looks at his friend Niko and realizes he is sad..."

[0101] Step 11:

[0102] The user interacts with the child to discuss the content of the story and develop the child's emotions.

[0103] Example: "How did Ravi know Nico was sad?" helps children develop an understanding of emotions.

[0104] Example 1

[0105] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0106] Previous emotional education systems faced challenges, such as a lack of an easy-to-use interface for users and the lack of the ability to automatically generate individual stories and images based on emotional domains. Furthermore, selecting a device that provides a visual experience that is easy on children's eyes is also important as a means of displaying the generated content, but this was often not properly implemented. As a result, there were problems with the creation of teaching materials suitable for emotional education and the learning process using them not being effectively supported.

[0107] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0108] In this invention, the server includes a selection means for a user to select an emotional area, a story generation means including a generative AI model for generating a story based on the selected emotional area, an image generation means including an image generation AI model for generating image prompts based on the generated story and generating images based on the prompts, a transmission means for transmitting the generated story and images to the user's terminal, and a display means for displaying the transmitted story and images. This makes it possible to automatically generate individual stories and images based on emotions using an interface that is easy for users to operate, and to display the generated content on a device that is easy on children's eyes.

[0109] A "selection means" is a device or software that provides an interface for a user to select a particular emotional domain.

[0110] A "generative AI model" is an artificial intelligence model that automatically generates stories and sentences based on selected emotional domains.

[0111] The "story generation means" is a means for generating a story using a generative AI model with emotional domain data selected via a selection means as input.

[0112] A "prompt" is a textual instruction given to an image-generating AI model to generate a specific image.

[0113] An "image generation AI model" is an artificial intelligence model that automatically generates corresponding images based on given prompts.

[0114] The "image generation means" is a means for generating an image generation prompt based on the story generated by the story generation means, and passing the prompt to an image generation AI model to generate an image.

[0115] The "transmission means" is a means for transmitting the generated story and images from the server to the user's terminal.

[0116] "Display means" refers to a device that displays the received story and images on the user's terminal, allowing the user to check the content.

[0117] "Electronic paper" is a visually friendly display technology that can display text and images without the use of a backlight.

[0118] This system automatically generates picture books for children with the aim of emotional education, and supports the development of children's emotions and interpersonal relationships through reading aloud. This system consists of a terminal operated by the user, a server that processes data, and a display means that displays the generated stories and images.

[0119] First, the user launches the emotion education system application on a smartphone or tablet device (e.g., iOS or Android). The user selects an emotion area through the application interface. For example, the user selects "empathy."

[0120] Next, the device sends the data of the user's selected emotional domain to the server. The server uses a cloud server (e.g., AWS, GCP) to analyze the received data and input the emotional domain data into a generative AI model (e.g., OpenAI GPT-4). This generates a story about the selected emotional domain. As a specific example, a story based on the theme of "empathy" might be generated, such as "The Story of Little Rabbit Ravi and His Friend Nico."

[0121] The generated story is then used to generate image generation prompts. The server creates image generation prompts based on the content of the story and passes them to an image generation AI model (e.g., DALL-E, MidJourney). This automatically generates images corresponding to scenes in the story. For example, an image depicting "Ravi and Nico sitting together" is generated.

[0122] The generated story and images are sent from the server to the device, which then analyzes the data and displays it on an electronic paper device (e.g., Kindle, e-ink device). This display device does not use a backlight to provide a visual experience that is gentle on children's eyes.

[0123] Finally, the user reads a story to the child using the e-paper device. For example, a parent might read the story, "Ravi looks at his friend Nico and realizes he's sad..." In addition, the user can deepen the child's understanding of emotions by interacting with the child and discussing the content of the story. For example, emotional education is provided by asking, "How did Ravi realize Nico was sad?"

[0124] As a result, this system provides an easy-to-use interface for users, automatically generates individual stories and images based on emotional domains, and displays them on an electronic paper device that is easy on children's eyes. By utilizing a generative AI model and an image-generating AI model, the present invention is characterized by its ability to quickly provide high-quality content suitable for emotional education.

[0125] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0126] Step 1:

[0127] The user launches the application and selects an emotion area.

[0128] Input: The user launches the application on their smartphone or tablet and selects an emotion domain (e.g., "empathy").

[0129] How it works: The application displays an interface for the user to select an emotional area. When the user selects an emotional area, the information is temporarily stored on the device.

[0130] Output: Data from the selected emotion domains is generated and passed to the next processing step.

[0131] Step 2:

[0132] The device sends the emotion area data to the server.

[0133] Input: Emotion domain data selected by the user in Step 1 (e.g., "Empathy")

[0134] How it works: The device sends emotional domain data to the server using an API request over the internet. The data is securely sent to the server, where analysis begins.

[0135] Output: Emotional domain data is sent to the server.

[0136] Step 3:

[0137] The server analyzes the received emotional data and generates a story.

[0138] Input: Emotion domain data sent to the server in step 2 (e.g., "Empathy")

[0139] How it works: The server analyzes the received emotional domain data and inputs it into a generative AI model (e.g., OpenAI GPT-4). The generative AI model generates a story based on the emotional domain. For example, it generates "The Story of Little Rabbit Ravi and His Friend Nico."

[0140] Output: The generated story text (e.g., "One day, Ravi the little rabbit found his friend Niko sad...")

[0141] Step 4:

[0142] The server generates image prompts based on the story and generates images.

[0143] Input: The narrative text generated in step 3

[0144] How it works: The server generates an image prompt based on the content of the generated story (e.g., "A scene where Ravi and Nico are together"). It then passes the image prompt to an image generation AI model (e.g., DALL-E, MidJourney). The image generation AI model generates an image based on the prompt. For example, it generates an image of "A scene where Ravi and Nico are sitting together."

[0145] Output: Generated image file (e.g. JPEG, PNG format)

[0146] Step 5:

[0147] The server sends the generated story and images to the device.

[0148] Input: The story text generated in step 3 and the image files generated in step 4

[0149] How it works: The server compiles the generated story text and image files and sends them to the user's device via an API request over the internet.

[0150] Output: A data package containing the story and images

[0151] Step 6:

[0152] The terminal analyzes the received data and displays it on the e-paper device.

[0153] Input: The data package sent by the server in step 5 (story text and image files)

[0154] How it works: The device analyzes the received data and displays the story text and images on an electronic paper device (e.g., Kindle, e-ink device). Electronic paper devices display without a backlight, providing a visual experience that is gentle on children's eyes.

[0155] Output: The displayed story and images

[0156] Step 7:

[0157] A user reads a story to a child on an e-paper device

[0158] Input: The story and image displayed in step 6

[0159] How it works: A user reads a story to a child using an e-paper device. For example, a parent reads, "Ravi looks at his friend Niko and realizes he's sad..." The user then interacts with the child, asking questions about the story to deepen the child's understanding of their emotions.

[0160] Output: Responses of children who received emotional education

[0161] The above are the specific processing steps of the program of this system.

[0162] (Application example 1)

[0163] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0164] The present invention relates to a system that automatically generates picture books to support children's emotional education. With conventional systems, it is difficult to find the time and timing for reading aloud, making it difficult to provide effective emotional education, especially when parents are busy or children have to wait long periods of time. Furthermore, there is a lack of a way to effectively utilize spare time, such as while waiting for delivery. Therefore, an objective of the present invention is to provide a system that can provide emotional education by effectively utilizing waiting time for delivery.

[0165] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0166] In this invention, the server includes a generation means for generating a story for each emotion, an image generation means for generating an image based on the generated story, a display means for displaying the generated story and image, and a promotion means for generating a story and image for each emotion, displaying them when a user places an order, and promoting reading aloud while waiting for delivery. This makes it possible to effectively utilize the waiting time for delivery to provide emotional education.

[0167] "Emotional stories" are stories generated based on specific emotional domains, featuring situations and characters related to those emotions.

[0168] "Generation means" refers to a device or program that has the function of automatically generating stories and images based on the emotional domain selected by the user.

[0169] "Image generation means" refers to a device or program that has the function of automatically generating images of scenes and characters related to the generated story.

[0170] "Display means" refers to a device or function that displays the generated story or images so that the user can view them.

[0171] "Promotion means" refers to a device or function that displays the story and images generated when the user places an order and encourages reading aloud while waiting for delivery.

[0172] The "selection means" refers to a device or program that provides an interface or function for the user to select a desired emotional area.

[0173] "Transmission means" refers to a device or program for transmitting the generated story and images to the user's terminal.

[0174] "Delivery Order" means an online or offline order placed by a User for delivery of food or merchandise.

[0175] This invention is a system that automatically generates picture books for the purpose of emotional education and supports the development of children's emotions and interpersonal relationships through reading aloud. This system promotes emotional education by allowing parents to read picture books to their children, making effective use of waiting times for delivery services.

[0176] 1. System Configuration

[0177] The system includes the following major components:

[0178] Server: Includes a generative AI model and an image generation model that generates stories and images based on emotional domains.

[0179] Device: A device on which parents place delivery orders and view the generated stories and images. Specifically, this refers to a smartphone or tablet.

[0180] User: In this case, the parent who places the delivery order and reads to their child.

[0181] 2. Explanation of program processing

[0182] 1. User operations

[0183] A user launches a food delivery app, orders a meal, and selects the "Emotional Education" option, where the user selects a specific emotional domain (e.g., empathy).

[0184] 2. Story Generation Based on Emotional Domains

[0185] The user's selection data is sent to a server, which uses a generative AI model (e.g., OpenAI GPT-4) to generate a narrative based on the selected emotional domains.

[0186] Example: Generate a story about a little rabbit character named "Ravi" and his friend "Nico" that has an empathy theme. The prompt is:

[0187] Input prompt: "Generate a story about Ravi and Nico based on empathy."

[0188] 3. Narrative-based image generation

[0189] Based on the generated story, the server generates appropriate images using an image generation model (e.g., DALL-E 2).

[0190] Example: An image depicting "Ravi and Nico together" is generated.

[0191] 4. Data transmission and display

[0192] The server sends the generated story and images to the user's terminal, which displays the received story and images.

[0193] Parents can read stories to their children while they wait for their delivery.

[0194] 5. Reading aloud and emotional education

[0195] Emotional education is carried out by parents reading stories and discussing emotions with their children through dialogue.

[0196] For example, asking, "How did Ravi know Nico was sad?" provides an opportunity for children to think about empathy.

[0197] 3. Hardware and Software Use

[0198] The server uses a generative AI model (OpenAI GPT-4) for story generation and a generative image model (DALL-E 2) for image generation.

[0199] The user's device runs a delivery application (e.g., a generic food delivery service app) and has a display for displaying stories and images.

[0200] The server receives the user's selection data and inputs the data into the AI ​​model for processing.

[0201] This makes it possible to provide a system that makes effective use of the waiting time for delivery and allows parents and children to enjoy reading picture books suitable for emotional education.

[0202] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0203] Step 1:

[0204] A user launches a food delivery app and selects the "Emotional Education" option when ordering food. The user also selects an emotional domain, such as "Empathy," through an interface that identifies emotional domains. This generates food delivery order data and emotional domain data.

[0205] Step 2:

[0206] The terminal transmits the emotional domain data selected by the user to the server. The server receives the user's selection data and analyzes the emotional domains. This data includes the emotional domain selected by the user, such as "empathy."

[0207] Step 3:

[0208] The server inputs the emotional domain data into a generative AI model (e.g., OpenAI GPT-4) for generating stories based on emotional domains. The server uses the AI ​​model to generate stories based on specific emotional domains. The input is a prompt statement: "Generate a story about Ravi and Nico based on empathy." The output is a story text with an empathy theme.

[0209] Step 4:

[0210] The server analyzes the generated story text and generates image prompts based on the story content. Image prompts are text data containing specific scenes in the story (e.g., "A scene where Ravi and Nico are together"). These prompts are input into an image generation AI model (e.g., DALL-E 2). The result is an image depicting a scene from the story.

[0211] Step 5:

[0212] The server compiles the generated story and image data and sends them in digital format to the terminal, which then retrieves the received story data and image data and displays them to the user in an appropriate format, allowing the user to view the story and corresponding images while waiting for delivery.

[0213] Step 6:

[0214] The user reads a story displayed on the device to the child. The user then interacts with the child about the story's contents and provides emotional education. Specifically, by discussing the actions and emotions of the characters in the story, the child can deepen their understanding of emotions.

[0215] Through these steps, parents and children can effectively utilize the time they spend waiting for delivery by reading picture books suitable for emotional education.

[0216] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0217] This invention is a system that automatically generates picture books for children with the aim of emotional education, and supports the development of children's emotions and interpersonal relationships through reading aloud. In addition to selecting emotional domains based on user input, this system incorporates an emotion engine that recognizes the user's emotions to generate more personalized picture books.

[0218] System Overview

[0219] 1. User operations

[0220] The user launches the application and manually selects an emotion area, or the device automatically recognizes the user's emotion using an emotion engine.

[0221] Example: The user manually selects "empathy," or the emotion engine recognizes the user's facial expression and voice as "calm" and selects "empathy."

[0222] 2. Story generation based on emotional domains

[0223] The server analyzes the emotion area data received from the device.

[0224] Example: The server checks that the emotion domain is "empathy."

[0225] The emotional domain of "empathy" is input into the generative AI model on the server to generate a story.

[0226] Example: An AI model generates stories with empathy as the theme, using a rabbit character named "Rabi."

[0227] Some of the stories generated:

[0228] One day, Ravi, a little rabbit, finds his friend Niko sad, and he wants to know why Niko is sad.

[0229] 3. Narrative-based image generation

[0230] The server generates image prompts based on the generated story.

[0231] Example: Generate a prompt for "A scene with Ravi and Nico together."

[0232] The server passes the prompt to an image generation AI model, which generates an appropriate image.

[0233] Example: An image generation AI generates an image of a scene where "Ravi and Nico are sitting together."

[0234] 4. Data transmission and display

[0235] The server compiles the generated story and images into a single document and sends it to the device.

[0236] Example: Sending story text and image files to the device.

[0237] The terminal analyzes the received data and displays it on the electronic paper device.

[0238] Example: An e-paper device displays "a story and corresponding images that show the relatable nature of Ravi and Nico."

[0239] 5. Reading aloud

[0240] A user reads a story to a child using an e-paper device.

[0241] Example: A parent reads part of a story: "Ravi looks at his friend Niko and realizes he is sad..."

[0242] The user interacts with the child to discuss the content of the story and develop the child's emotions.

[0243] Example: "How did Ravi know Nico was sad?" helps children develop an understanding of emotions.

[0244] This system allows parents to easily create picture books suitable for emotional education and read them to their children. Using an emotion engine, it provides a more personalized experience according to the user's mood and state. Using an e-paper device, it provides a visual experience that is gentle on children's eyes while supporting the development of emotions and interpersonal relationships.

[0245] The processing flow will be explained below.

[0246] Step 1:

[0247] When a user launches the application, the emotion engine recognizes the user's emotion, and the device displays an emotion area based on the user's input.

[0248] Example: A user launches an application, and the emotion engine selects "empathy" based on the user's facial expressions and voice.

[0249] Step 2:

[0250] The terminal transmits the user's selection data to the server.

[0251] Example: Send data from the emotion field "empathy" to the server.

[0252] Step 3:

[0253] The server analyzes the emotion area data received from the device.

[0254] Example: The server checks that the emotion domain is "empathy."

[0255] Step 4:

[0256] The emotional domain of "empathy" is input into the generative AI model on the server to generate a story.

[0257] Example: An AI model generates stories with empathy as the theme, using a rabbit character named "Rabi."

[0258] Some of the stories generated:

[0259] One day, Ravi, a little rabbit, finds his friend Niko sad, and he wants to know why Niko is sad.

[0260] Step 5:

[0261] The server generates image prompts based on the generated story.

[0262] Example: Generate a prompt for "A scene with Ravi and Nico together."

[0263] Step 6:

[0264] The server passes the prompt to an image generation AI model, which generates an appropriate image.

[0265] Example: An image generation AI generates an image of a scene where "Ravi and Nico are sitting together."

[0266] Step 7:

[0267] The server compiles the generated story and images into a single document.

[0268] Example: Combining generated narrative text and images.

[0269] Examples of compiled data:

[0270] Story text:

[0271] One day, a little rabbit named Ravi found his friend Niko sad.

[0272] Story Image:

[0273] Image file of Ravi and Nico sitting together

[0274] Step 8:

[0275] The server sends the compiled data to the terminal.

[0276] Example: Sending a story and image file to a device.

[0277] Step 9:

[0278] The terminal analyzes the received data and displays it on the electronic paper device.

[0279] Example: An e-paper device displays "a story and corresponding images that show the relatable nature of Ravi and Nico."

[0280] Step 10:

[0281] A user reads a story to a child using an e-paper device.

[0282] Example: A parent reads part of a story: "Ravi looks at his friend Niko and realizes he is sad..."

[0283] Step 11:

[0284] The user interacts with the child to discuss the content of the story and develop the child's emotions.

[0285] Example: "How did Ravi know Nico was sad?" helps children develop an understanding of emotions.

[0286] Example 2

[0287] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0288] A major challenge is the lack of teaching materials and tools for effectively educating children about emotions. There are also limited ways for parents to easily create picture books suitable for emotional education and use them to help their children understand emotions. Furthermore, as digital devices are used, there is a need to support the development of emotions and interpersonal relationships while reducing visual burden.

[0289] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a recognition means for recognizing a user's emotion, a story generation means for generating a story based on the recognized emotion, an image generation means for generating images based on the generated story, and a display means for displaying the generated story and images. This makes it possible to automatically generate a personalized picture book according to the user's emotion and display it using an electronic paper device, thereby enabling effective emotional education for children while reducing visual burden.

[0290] The "recognition means" is a means for recognizing the user's emotions.

[0291] A "narrative generation means" is a means for generating a narrative based on recognized emotions.

[0292] The "image generation means" is a means for generating an image based on the generated story.

[0293] "Display means" is a means for displaying the generated story and images.

[0294] The "selection means" is a means for selecting an emotion area based on a user's input.

[0295] "Electronic paper" is a backlit display device used to reduce visual strain.

[0296] The "transmission means" is a means for transmitting the generated story to the terminal.

[0297] The "reading support tool" is a tool for reading the generated story to a child.

[0298] This invention is a system that automatically generates picture books for children with the aim of emotional education, and supports the development of children's emotions and interpersonal relationships through reading aloud. In addition to selecting emotional domains based on user input, this system incorporates an emotion engine to recognize the user's emotions and generate more personalized picture books.

[0299] First, the user launches the application and manually selects an emotion area, or the device automatically recognizes the user's emotion using an emotion engine. For example, the user manually selects "empathy," or the emotion engine recognizes the user's facial expression and voice as "calm" and selects "empathy."

[0300] Next, the device sends the selected emotional domain to the server. The server analyzes this data and confirms that the emotional domain is "empathy." The "empathy" emotional domain is input into the generative AI model, which then generates a story. As a concrete example, the generative AI model uses a rabbit character named "Rabi" to generate a story with an empathy theme.

[0301] Some examples of generated stories:

[0302] One day, Ravi, a little rabbit, finds his friend Niko sad, and he wants to know why Niko is sad.

[0303] The server then generates image prompts based on the generated story. For example, it generates a prompt for "a scene where Ravi and Nico are together." The server then passes the prompt to an image generation AI model, which generates an appropriate image. For example, the image generation AI generates an image of a scene where "Ravi and Nico are sitting together."

[0304] Finally, the server compiles the story and the generated images into a single document and sends it to the device. The device then analyzes the received data and displays it on the e-paper device. For example, the e-paper device displays "a story that shows the empathy between Ravi and Nico and the corresponding images."

[0305] The user uses the e-paper device to read a story to the child. For example, a parent might read part of the story, "Ravi looks at his friend Nico and realizes he's sad..." The user interacts with the child, discussing the content of the story and developing the child's emotions. For example, the parent might ask, "How did Ravi realize Nico was sad?", helping the child deepen their understanding of emotions.

[0306] This system allows parents to easily create picture books suitable for emotional education and read them to their children. Using an emotion engine, it provides a more personalized experience according to the user's mood and state. Using an e-paper device, it provides a visual experience that is gentle on children's eyes while supporting the development of emotions and interpersonal relationships.

[0307] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0308] Step 1:

[0309] The user launches the application and selects an emotion area, or the terminal uses an emotion engine to recognize the user's emotion.

[0310] Input: Emotion area selection screen display, user's facial expression data and voice data.

[0311] Data processing: Users can select emotion areas from a drop-down menu, or the emotion engine analyzes facial expressions and voice.

[0312] Output: Selected emotion domain data.

[0313] What happens: The user operates their smartphone to launch the app and taps the "Empathy" button. The device uses the camera to scan the user's facial expressions and performs voice recognition.

[0314] Step 2:

[0315] The terminal transmits the selected emotion area to the server.

[0316] Input: Emotion domain data.

[0317] Data processing: Packaging and sending emotional domain data.

[0318] Output: Emotional domain data sent to the server.

[0319] Specific operation: The device sends data on the emotional domain "empathy" to the server via the emotion engine API.

[0320] Step 3:

[0321] The server analyzes the received emotional domain data and generates a story for the picture book using a story generation means.

[0322] Input: Emotion domain data.

[0323] Data processing: Analysis of emotional domain data, input into generative AI models, and narrative generation.

[0324] Output: The generated narrative text.

[0325] Specific operation: The server inputs the emotional domain data "empathy" into the generative AI model and generates a story themed around empathy using a character named "Rabi."

[0326] Step 4:

[0327] The server generates image prompts based on the generated story and generates images using an image generation means.

[0328] Input: Narrative text.

[0329] Data processing: Generate image prompts from narrative text, input them into an image generation AI model, and generate images.

[0330] Output: The generated image data.

[0331] Specific operation: The server creates a prompt describing "a scene where Ravi and Nico are together" and inputs it into an image generation AI model to generate an image.

[0332] Step 5:

[0333] The server compiles the generated story text and images into a single document and sends it to the terminal.

[0334] Input: Narrative text and image data.

[0335] Data processing: Organizing data into document formats such as PDF and sending it to the device.

[0336] Output: The document data sent to the device.

[0337] Specific operation: The server compiles the story text and images into PDF format and sends it to the terminal.

[0338] Step 6:

[0339] The terminal analyzes the received data and displays it on the electronic paper device.

[0340] Input: The received document data.

[0341] Data processing: Analyzing document data and preparing it for display on e-paper devices.

[0342] Output: Story and images displayed on an e-paper device.

[0343] Specific operation: The terminal displays the received PDF on the e-paper device.

[0344] Step 7:

[0345] A user reads a story to a child using an e-paper device.

[0346] Input: A story and an image displayed on an e-paper device.

[0347] Data processing: Reading aloud was conducted.

[0348] Output: Providing emotional education to children.

[0349] Specific Action: The user reads the story "Ravi looks at his friend Niko and realizes he is sad..." and interacts with the child to deepen their understanding of emotions.

[0350] (Application example 2)

[0351] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0352] Previous story generation systems were limited in their ability to provide appropriate stories and images to support emotional education and interpersonal relationship development. They also struggled to recognize users' emotions in real time and provide personalized stories and images based on those emotions. Furthermore, the means to interactively experience the generated stories were limited, limiting the enrichment of children's learning experiences.

[0353] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a generation means for generating a story for each emotion, an image generation means for generating images based on the generated story, a display means for displaying the generated story and images, an emotion recognition means for automatically recognizing emotions, an audio playback means for reading aloud the generated story, and a cross-platform support means for supporting multiple display devices. This makes it possible to recognize the user's emotions in real time and provide personalized stories and images based on them, thereby providing a more interactive story-telling experience.

[0354] The "generation means for generating a story for each emotion" is a means for automatically generating a story corresponding to a specific emotion based on a user's selection and emotion recognition data.

[0355] The "image generating means for generating images based on the generated story" is a means for automatically generating images corresponding to the content of the generated story.

[0356] A "display means for displaying the generated story and images" is a device or interface for visually presenting the generated story and images to a user.

[0357] "Emotion recognition means for automatically recognizing emotions" is a means for automatically analyzing and recognizing emotions from the user's facial expressions and voice using sensors such as a camera and a microphone.

[0358] The "audio playback means for reading aloud" is a means for playing back the generated story aloud and providing it to the user audibly.

[0359] "Cross-platform support for multiple display devices" refers to the means by which the generated story and images are adjusted to be displayed appropriately on different types of devices (smartphones, tablets, smart glasses, head-mounted displays, etc.).

[0360] The "selection means for selecting an emotional area based on user input" is a means that allows a user to specify a particular emotional area based on their current feelings and intentions.

[0361] The "display means without backlight" refers to a means for displaying without using a backlight in order to reduce power consumption.

[0362] "Display means using electronic paper" refers to a display device that uses electronic ink technology and combines the ease of viewing as paper with power-saving performance.

[0363] This invention is a system that automatically generates stories and images for children using real-time emotion recognition and generation AI for the purpose of emotional education. The system automatically recognizes the user's emotions, generates a story based on those emotions, and then generates corresponding images. The generated stories and images are displayed on multiple display devices, providing an interactive reading experience through a text-to-speech function.

[0364] Hardware and software used

[0365] 1. Hardware

[0366] Smartphone

[0367] tablet

[0368] Smart glasses (e.g. Google Glass)

[0369] Head-mounted displays (e.g., Oculus Quest)

[0370] Electronic Paper Device

[0371] 2. Software

[0372] Azure Cognitive Services (for emotion recognition)

[0373] AWS Lambda (serverless computing)

[0374] S3 bucket (data storage)

[0375] OpenAI GPT-4 (a generative AI model for story generation)

[0376] DALL-E 2 (generative AI model for image generation)

[0377] React Native (for cross-platform compatibility)

[0378] Amazon Polly (for voice reading)

[0379] System Overview

[0380] 1. User Emotion Recognition

[0381] When a user launches an application, Azure Cognitive Services uses the device's camera and microphone to recognize the user's emotions. For example, if facial expression recognition technology determines that the user is "calm," this information is sent to the server.

[0382] 2. Narrative Generation Based on Emotional Data

[0383] Once emotion recognition data is sent to the server, OpenAI GPT-4 is invoked using AWS Lambda. Using emotion data (e.g., "empathy") as input, the generative AI model generates a related story. As a specific example, it generates an empathetic story using a rabbit character named "Rabi."

[0384] 3. Image Generation

[0385] Based on the content of the generated story, images that fit the story are generated using DALL-E 2. For example, a scene in the story where Ravi and Nico are sitting together is depicted.

[0386] 4. Data transmission and display

[0387] The generated stories and images are compiled into documents and sent using React Native to smartphones, tablets, smart glasses, head-mounted displays, and potentially even backlit e-paper devices.

[0388] 5. Reading aloud

[0389] The story is generated using Amazon Polly and read aloud, and an interface is provided for users to enjoy the story interactively.

[0390] Examples of concrete examples and prompts

[0391] For example, a user launches an app and Azure Cognitive Services determines that the user feels "calm." Emotional data "empathy" is sent to the server, OpenAI GPT-4 generates an "empathetic story about Ravi and Nico," and DALL-E 2 generates an image of "a scene where Ravi and Nico are sitting together." The generated story and image are sent to the device and read aloud.

[0392] An example prompt is:

[0393] "Generate stories with empathy as a theme based on emotional data"

[0394] "Based on the story you've created, generate an image of the rabbit character and his friends together."

[0395] The present invention generates stories and images that correspond to the user's real-time emotions, making it possible to support children's emotional and interpersonal development through an interactive storytelling experience.

[0396] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0397] Step 1:

[0398] The device launches an application and activates the device's camera and microphone to automatically recognize the user's emotions. Azure Cognitive Services is used to analyze data obtained from the camera and microphone and recognize the user's emotions. For example, a facial expression recognition algorithm determines that the user is "calm." The input is the user's facial expression and voice data, and the output is the recognized emotion data.

[0399] Step 2:

[0400] The device sends the recognized emotion data to the server. Specifically, the emotion data called "empathy" is sent in JSON format. The input is the emotion data output from Step 1, and the output is a notification of completion of data transmission to the server.

[0401] Step 3:

[0402] Based on the emotion data received by the server, AWS Lambda is triggered to generate a story. The Lambda function calls OpenAI GPT-4 and uses the emotion data as input to create a story generation prompt. The input is emotion data and the story generation prompt, and the output is the generated story text. Specifically, the prompt text "Generate a story with the theme of empathy" with the emotion "empathy" is sent to GPT-4, which then generates a story.

[0403] Step 4:

[0404] The server uses DALL-E 2 to generate images based on the generated story text. Specific scenes from the story are sent as prompts to DALL-E 2, which then generates corresponding images. The input is the story text and the image generation prompt, and the output is the generated image file. For example, an image is generated based on the prompt "A scene where Ravi and Nico are sitting together."

[0405] Step 5:

[0406] The server compiles the generated story text and images into a single document and sends it to the terminal. The input is the story text and image files, and the output is the compiled document file and a notification of completion of transmission. The data is compiled into a JSON format document file and sent to the terminal.

[0407] Step 6:

[0408] The device analyzes the received document file and provides it to the user through a display method. For example, using React Native, it can be displayed on a smartphone, tablet, smart glasses, or head-mounted display. The input is the document file, and the output is the story and images displayed on the device.

[0409] Step 7:

[0410] The device uses Amazon Polly to read aloud the story text generated, allowing users to enjoy the story not only visually but also aurally. The input is the story text, and the output is audio data, specifically the content of the story being read aloud.

[0411] Each step is processed sequentially to generate a personalized story and images based on the user's emotions, ultimately providing an interactive storytelling experience.

[0412] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0413] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0414] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0415] [Second embodiment]

[0416] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0417] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0418] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0419] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0420] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0421] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0422] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0423] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0424] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0425] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0426] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0427] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0428] This invention is a system that automatically generates picture books for children and supports the development of their emotions and interpersonal relationships through reading aloud, with the aim of emotional education. The system provides an interface that is easy for users to operate, and generates and displays stories and images based on emotional domains.

[0429] System Overview

[0430] 1. User operations

[0431] The user launches the application and selects an emotional area.

[0432] Example: A user selects "empathy" within an application.

[0433] The terminal transmits the user's selection data to the server.

[0434] 2. Story generation based on emotional domains

[0435] The server analyzes the emotion area data received from the device.

[0436] Example: The server checks that the emotion domain is "empathy."

[0437] Emotional domains are input into the generative AI model on the server to generate a story.

[0438] Example: An AI model generates a story about a little rabbit character named Ravi and his friend Nico, with a theme of empathy.

[0439] Some of the stories generated:

[0440] One day, Ravi, a little rabbit, finds his friend Niko sad, and he wants to know why Niko is sad.

[0441] 3. Narrative-based image generation

[0442] The server generates image prompts based on the generated story.

[0443] Example: Generate a prompt for "A scene with Ravi and Nico together."

[0444] The server passes the prompt to an image generation AI model, which generates an appropriate image.

[0445] Example: An image generation AI generates an image depicting a scene in which "Ravi and Nico are sitting together."

[0446] 4. Data transmission and display

[0447] The server compiles the generated story and images and sends them to the device.

[0448] Example: Sending story text and image files to a device.

[0449] The terminal analyzes the received data and displays it on the electronic paper device.

[0450] Example: An e-paper device displays "a story and corresponding images that show the relatable nature of Ravi and Nico."

[0451] 5. Reading aloud

[0452] A user reads a story to a child using an e-paper device.

[0453] Example: A parent reads part of a story: "Ravi looks at his friend Niko and realizes he is sad..."

[0454] The user interacts with the child to discuss the content of the story and develop the child's emotions.

[0455] Example: "How did Ravi know Nico was sad?" helps children develop an understanding of emotions.

[0456] This system allows parents to easily create picture books suitable for emotional education and read them to their children. By using an e-paper device, it provides a visual experience that is gentle on children's eyes while supporting the development of emotions and interpersonal relationships.

[0457] The processing flow will be explained below.

[0458] Step 1:

[0459] The user launches the application and selects an emotional area.

[0460] Example: A user selects "empathy" within an application.

[0461] Step 2:

[0462] The terminal transmits the user's selection data to the server.

[0463] Example: Send data from the emotion field "empathy" to the server.

[0464] Step 3:

[0465] The server analyzes the emotion area data received from the device.

[0466] Example: The server checks that the emotion domain is "empathy."

[0467] Step 4:

[0468] The emotional domain of "empathy" is input into the generative AI model on the server to generate a story.

[0469] Example: An AI model generates stories with empathy as the theme, using a rabbit character named "Rabi."

[0470] Some of the stories generated:

[0471] One day, Ravi, a little rabbit, finds his friend Niko sad, and he wants to know why Niko is sad.

[0472] Step 5:

[0473] The server generates image prompts based on the generated story.

[0474] Example: Generate a prompt for "A scene with Ravi and Nico together."

[0475] Step 6:

[0476] The server passes the prompt to an image generation AI model, which generates an appropriate image.

[0477] Example: An image generation AI generates an image of a scene where "Ravi and Nico are sitting together."

[0478] Step 7:

[0479] The server compiles the generated story and images into a single document.

[0480] Example: Combining generated narrative text and images.

[0481] Examples of compiled data:

[0482] Story text:

[0483] One day, a little rabbit named Ravi found his friend Niko sad.

[0484] Story Image:

[0485] Image file of Ravi and Nico sitting together

[0486] Step 8:

[0487] The server sends the compiled data to the terminal.

[0488] Example: Sending a story and image file to a device.

[0489] Step 9:

[0490] The terminal analyzes the received data and displays it on the electronic paper device.

[0491] Example: An e-paper device displays "a story and corresponding images that show the relatable nature of Ravi and Nico."

[0492] Step 10:

[0493] A user reads a story to a child using an e-paper device.

[0494] Example: A parent reads part of a story: "Ravi looks at his friend Niko and realizes he is sad..."

[0495] Step 11:

[0496] The user interacts with the child to discuss the content of the story and develop the child's emotions.

[0497] Example: "How did Ravi know Nico was sad?" helps children develop an understanding of emotions.

[0498] Example 1

[0499] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0500] Previous emotional education systems faced challenges, such as a lack of an easy-to-use interface for users and the lack of the ability to automatically generate individual stories and images based on emotional domains. Furthermore, selecting a device that provides a visual experience that is easy on children's eyes is also important as a means of displaying the generated content, but this was often not properly implemented. As a result, there were problems with the creation of teaching materials suitable for emotional education and the learning process using them not being effectively supported.

[0501] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0502] In this invention, the server includes a selection means for a user to select an emotional area, a story generation means including a generative AI model for generating a story based on the selected emotional area, an image generation means including an image generation AI model for generating image prompts based on the generated story and generating images based on the prompts, a transmission means for transmitting the generated story and images to the user's terminal, and a display means for displaying the transmitted story and images. This makes it possible to automatically generate individual stories and images based on emotions using an interface that is easy for users to operate, and to display the generated content on a device that is easy on children's eyes.

[0503] A "selection means" is a device or software that provides an interface for a user to select a particular emotional domain.

[0504] A "generative AI model" is an artificial intelligence model that automatically generates stories and sentences based on selected emotional domains.

[0505] The "story generation means" is a means for generating a story using a generative AI model with emotional domain data selected via a selection means as input.

[0506] A "prompt" is a textual instruction given to an image-generating AI model to generate a specific image.

[0507] An "image generation AI model" is an artificial intelligence model that automatically generates corresponding images based on given prompts.

[0508] The "image generation means" is a means for generating an image generation prompt based on the story generated by the story generation means, and passing the prompt to an image generation AI model to generate an image.

[0509] The "transmission means" is a means for transmitting the generated story and images from the server to the user's terminal.

[0510] "Display means" refers to a device that displays the received story and images on the user's terminal, allowing the user to check the content.

[0511] "Electronic paper" is a visually friendly display technology that can display text and images without the use of a backlight.

[0512] This system automatically generates picture books for children with the aim of emotional education, and supports the development of children's emotions and interpersonal relationships through reading aloud. This system consists of a terminal operated by the user, a server that processes data, and a display means that displays the generated stories and images.

[0513] First, the user launches the emotion education system application on a smartphone or tablet device (e.g., iOS or Android). The user selects an emotion area through the application interface. For example, the user selects "empathy."

[0514] Next, the device sends the data of the user's selected emotional domain to the server. The server uses a cloud server (e.g., AWS, GCP) to analyze the received data and input the emotional domain data into a generative AI model (e.g., OpenAI GPT-4). This generates a story about the selected emotional domain. As a specific example, a story based on the theme of "empathy" might be generated, such as "The Story of Little Rabbit Ravi and His Friend Nico."

[0515] The generated story is then used to generate image generation prompts. The server creates image generation prompts based on the content of the story and passes them to an image generation AI model (e.g., DALL-E, MidJourney). This automatically generates images corresponding to scenes in the story. For example, an image depicting "Ravi and Nico sitting together" is generated.

[0516] The generated story and images are sent from the server to the device, which then analyzes the data and displays it on an electronic paper device (e.g., Kindle, e-ink device). This display device does not use a backlight to provide a visual experience that is gentle on children's eyes.

[0517] Finally, the user reads a story to the child using the e-paper device. For example, a parent might read the story, "Ravi looks at his friend Nico and realizes he's sad..." In addition, the user can deepen the child's understanding of emotions by interacting with the child and discussing the content of the story. For example, emotional education is provided by asking, "How did Ravi realize Nico was sad?"

[0518] As a result, this system provides an easy-to-use interface for users, automatically generates individual stories and images based on emotional domains, and displays them on an electronic paper device that is easy on children's eyes. By utilizing a generative AI model and an image-generating AI model, the present invention is characterized by its ability to quickly provide high-quality content suitable for emotional education.

[0519] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0520] Step 1:

[0521] The user launches the application and selects an emotion area.

[0522] Input: The user launches the application on their smartphone or tablet and selects an emotion domain (e.g., "empathy").

[0523] How it works: The application displays an interface for the user to select an emotional area. When the user selects an emotional area, the information is temporarily stored on the device.

[0524] Output: Data from the selected emotion domains is generated and passed to the next processing step.

[0525] Step 2:

[0526] The device sends the emotion area data to the server.

[0527] Input: Emotion domain data selected by the user in Step 1 (e.g., "Empathy")

[0528] How it works: The device sends emotional domain data to the server using an API request over the internet. The data is securely sent to the server, where analysis begins.

[0529] Output: Emotional domain data is sent to the server.

[0530] Step 3:

[0531] The server analyzes the received emotional data and generates a story.

[0532] Input: Emotion domain data sent to the server in step 2 (e.g., "Empathy")

[0533] How it works: The server analyzes the received emotional domain data and inputs it into a generative AI model (e.g., OpenAI GPT-4). The generative AI model generates a story based on the emotional domain. For example, it generates "The Story of Little Rabbit Ravi and His Friend Nico."

[0534] Output: The generated story text (e.g., "One day, Ravi the little rabbit found his friend Niko sad...")

[0535] Step 4:

[0536] The server generates image prompts based on the story and generates images.

[0537] Input: The narrative text generated in step 3

[0538] How it works: The server generates an image prompt based on the content of the generated story (e.g., "A scene where Ravi and Nico are together"). It then passes the image prompt to an image generation AI model (e.g., DALL-E, MidJourney). The image generation AI model generates an image based on the prompt. For example, it generates an image of "A scene where Ravi and Nico are sitting together."

[0539] Output: Generated image file (e.g. JPEG, PNG format)

[0540] Step 5:

[0541] The server sends the generated story and images to the device.

[0542] Input: The story text generated in step 3 and the image files generated in step 4

[0543] How it works: The server compiles the generated story text and image files and sends them to the user's device via an API request over the internet.

[0544] Output: A data package containing the story and images

[0545] Step 6:

[0546] The terminal analyzes the received data and displays it on the e-paper device.

[0547] Input: The data package sent by the server in step 5 (story text and image files)

[0548] How it works: The device analyzes the received data and displays the story text and images on an electronic paper device (e.g., Kindle, e-ink device). Electronic paper devices display without a backlight, providing a visual experience that is gentle on children's eyes.

[0549] Output: The displayed story and images

[0550] Step 7:

[0551] A user reads a story to a child on an e-paper device

[0552] Input: The story and image displayed in step 6

[0553] How it works: A user reads a story to a child using an e-paper device. For example, a parent reads, "Ravi looks at his friend Niko and realizes he's sad..." The user then interacts with the child, asking questions about the story to deepen the child's understanding of their emotions.

[0554] Output: Responses of children who received emotional education

[0555] The above are the specific processing steps of the program of this system.

[0556] (Application example 1)

[0557] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0558] The present invention relates to a system that automatically generates picture books to support children's emotional education. With conventional systems, it is difficult to find the time and timing for reading aloud, making it difficult to provide effective emotional education, especially when parents are busy or children have to wait long periods of time. Furthermore, there is a lack of a way to effectively utilize spare time, such as while waiting for delivery. Therefore, an objective of the present invention is to provide a system that can provide emotional education by effectively utilizing waiting time for delivery.

[0559] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0560] In this invention, the server includes a generation means for generating a story for each emotion, an image generation means for generating an image based on the generated story, a display means for displaying the generated story and image, and a promotion means for generating a story and image for each emotion, displaying them when a user places an order, and promoting reading aloud while waiting for delivery. This makes it possible to effectively utilize the waiting time for delivery to provide emotional education.

[0561] "Emotional stories" are stories generated based on specific emotional domains, featuring situations and characters related to those emotions.

[0562] "Generation means" refers to a device or program that has the function of automatically generating stories and images based on the emotional domain selected by the user.

[0563] "Image generation means" refers to a device or program that has the function of automatically generating images of scenes and characters related to the generated story.

[0564] "Display means" refers to a device or function that displays the generated story or images so that the user can view them.

[0565] "Promotion means" refers to a device or function that displays the story and images generated when the user places an order and encourages reading aloud while waiting for delivery.

[0566] The "selection means" refers to a device or program that provides an interface or function for the user to select a desired emotional area.

[0567] "Transmission means" refers to a device or program for transmitting the generated story and images to the user's terminal.

[0568] "Delivery Order" means an online or offline order placed by a User for delivery of food or merchandise.

[0569] This invention is a system that automatically generates picture books for the purpose of emotional education and supports the development of children's emotions and interpersonal relationships through reading aloud. This system promotes emotional education by allowing parents to read picture books to their children, making effective use of waiting times for delivery services.

[0570] 1. System Configuration

[0571] The system includes the following major components:

[0572] Server: Includes a generative AI model and an image generation model that generates stories and images based on emotional domains.

[0573] Device: A device on which parents place delivery orders and view the generated stories and images. Specifically, this refers to a smartphone or tablet.

[0574] User: In this case, the parent who places the delivery order and reads to their child.

[0575] 2. Explanation of program processing

[0576] 1. User operations

[0577] A user launches a food delivery app, orders a meal, and selects the "Emotional Education" option, where the user selects a specific emotional domain (e.g., empathy).

[0578] 2. Story Generation Based on Emotional Domains

[0579] The user's selection data is sent to a server, which uses a generative AI model (e.g., OpenAI GPT-4) to generate a narrative based on the selected emotional domains.

[0580] Example: Generate a story about a little rabbit character named "Ravi" and his friend "Nico" that has an empathy theme. The prompt is:

[0581] Input prompt: "Generate a story about Ravi and Nico based on empathy."

[0582] 3. Narrative-based image generation

[0583] Based on the generated story, the server generates appropriate images using an image generation model (e.g., DALL-E 2).

[0584] Example: An image depicting "Ravi and Nico together" is generated.

[0585] 4. Data transmission and display

[0586] The server sends the generated story and images to the user's terminal, which displays the received story and images.

[0587] Parents can read stories to their children while they wait for their delivery.

[0588] 5. Reading aloud and emotional education

[0589] Emotional education is carried out by parents reading stories and discussing emotions with their children through dialogue.

[0590] For example, asking, "How did Ravi know Nico was sad?" provides an opportunity for children to think about empathy.

[0591] 3. Hardware and Software Use

[0592] The server uses a generative AI model (OpenAI GPT-4) for story generation and a generative image model (DALL-E 2) for image generation.

[0593] The user's device runs a delivery application (e.g., a generic food delivery service app) and has a display for displaying stories and images.

[0594] The server receives the user's selection data and inputs the data into the AI ​​model for processing.

[0595] This makes it possible to provide a system that makes effective use of the waiting time for delivery and allows parents and children to enjoy reading picture books suitable for emotional education.

[0596] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0597] Step 1:

[0598] A user launches a food delivery app and selects the "Emotional Education" option when ordering food. The user also selects an emotional domain, such as "Empathy," through an interface that identifies emotional domains. This generates food delivery order data and emotional domain data.

[0599] Step 2:

[0600] The terminal transmits the emotional domain data selected by the user to the server. The server receives the user's selection data and analyzes the emotional domains. This data includes the emotional domain selected by the user, such as "empathy."

[0601] Step 3:

[0602] The server inputs the emotional domain data into a generative AI model (e.g., OpenAI GPT-4) for generating stories based on emotional domains. The server uses the AI ​​model to generate stories based on specific emotional domains. The input is a prompt statement: "Generate a story about Ravi and Nico based on empathy." The output is a story text with an empathy theme.

[0603] Step 4:

[0604] The server analyzes the generated story text and generates image prompts based on the story content. Image prompts are text data containing specific scenes in the story (e.g., "A scene where Ravi and Nico are together"). These prompts are input into an image generation AI model (e.g., DALL-E 2). The result is an image depicting a scene from the story.

[0605] Step 5:

[0606] The server compiles the generated story and image data and sends them in digital format to the terminal, which then retrieves the received story data and image data and displays them to the user in an appropriate format, allowing the user to view the story and corresponding images while waiting for delivery.

[0607] Step 6:

[0608] The user reads a story displayed on the device to the child. The user then interacts with the child about the story's contents and provides emotional education. Specifically, by discussing the actions and emotions of the characters in the story, the child can deepen their understanding of emotions.

[0609] Through these steps, parents and children can effectively utilize the time they spend waiting for delivery by reading picture books suitable for emotional education.

[0610] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0611] This invention is a system that automatically generates picture books for children with the aim of emotional education, and supports the development of children's emotions and interpersonal relationships through reading aloud. In addition to selecting emotional domains based on user input, this system incorporates an emotion engine that recognizes the user's emotions to generate more personalized picture books.

[0612] System Overview

[0613] 1. User operations

[0614] The user launches the application and manually selects an emotion area, or the device automatically recognizes the user's emotion using an emotion engine.

[0615] Example: The user manually selects "empathy," or the emotion engine recognizes the user's facial expression and voice as "calm" and selects "empathy."

[0616] 2. Story generation based on emotional domains

[0617] The server analyzes the emotion area data received from the device.

[0618] Example: The server checks that the emotion domain is "empathy."

[0619] The emotional domain of "empathy" is input into the generative AI model on the server to generate a story.

[0620] Example: An AI model generates stories with empathy as the theme, using a rabbit character named "Rabi."

[0621] Some of the stories generated:

[0622] One day, Ravi, a little rabbit, finds his friend Niko sad, and he wants to know why Niko is sad.

[0623] 3. Narrative-based image generation

[0624] The server generates image prompts based on the generated story.

[0625] Example: Generate a prompt for "A scene with Ravi and Nico together."

[0626] The server passes the prompt to an image generation AI model, which generates an appropriate image.

[0627] Example: An image generation AI generates an image of a scene where "Ravi and Nico are sitting together."

[0628] 4. Data transmission and display

[0629] The server compiles the generated story and images into a single document and sends it to the device.

[0630] Example: Sending story text and image files to the device.

[0631] The terminal analyzes the received data and displays it on the electronic paper device.

[0632] Example: An e-paper device displays "a story and corresponding images that show the relatable nature of Ravi and Nico."

[0633] 5. Reading aloud

[0634] A user reads a story to a child using an e-paper device.

[0635] Example: A parent reads part of a story: "Ravi looks at his friend Niko and realizes he is sad..."

[0636] The user interacts with the child to discuss the content of the story and develop the child's emotions.

[0637] Example: "How did Ravi know Nico was sad?" helps children develop an understanding of emotions.

[0638] This system allows parents to easily create picture books suitable for emotional education and read them to their children. Using an emotion engine, it provides a more personalized experience according to the user's mood and state. Using an e-paper device, it provides a visual experience that is gentle on children's eyes while supporting the development of emotions and interpersonal relationships.

[0639] The processing flow will be explained below.

[0640] Step 1:

[0641] When a user launches the application, the emotion engine recognizes the user's emotion, and the device displays an emotion area based on the user's input.

[0642] Example: A user launches an application, and the emotion engine selects "empathy" based on the user's facial expressions and voice.

[0643] Step 2:

[0644] The terminal transmits the user's selection data to the server.

[0645] Example: Send data from the emotion field "empathy" to the server.

[0646] Step 3:

[0647] The server analyzes the emotion area data received from the device.

[0648] Example: The server checks that the emotion domain is "empathy."

[0649] Step 4:

[0650] The emotional domain of "empathy" is input into the generative AI model on the server to generate a story.

[0651] Example: An AI model generates stories with empathy as the theme, using a rabbit character named "Rabi."

[0652] Some of the stories generated:

[0653] One day, Ravi, a little rabbit, finds his friend Niko sad, and he wants to know why Niko is sad.

[0654] Step 5:

[0655] The server generates image prompts based on the generated story.

[0656] Example: Generate a prompt for "A scene with Ravi and Nico together."

[0657] Step 6:

[0658] The server passes the prompt to an image generation AI model, which generates an appropriate image.

[0659] Example: An image generation AI generates an image of a scene where "Ravi and Nico are sitting together."

[0660] Step 7:

[0661] The server compiles the generated story and images into a single document.

[0662] Example: Combining generated narrative text and images.

[0663] Examples of compiled data:

[0664] Story text:

[0665] One day, a little rabbit named Ravi found his friend Niko sad.

[0666] Story Image:

[0667] Image file of Ravi and Nico sitting together

[0668] Step 8:

[0669] The server sends the compiled data to the terminal.

[0670] Example: Sending a story and image file to a device.

[0671] Step 9:

[0672] The terminal analyzes the received data and displays it on the electronic paper device.

[0673] Example: An e-paper device displays "a story and corresponding images that show the relatable nature of Ravi and Nico."

[0674] Step 10:

[0675] A user reads a story to a child using an e-paper device.

[0676] Example: A parent reads part of a story: "Ravi looks at his friend Niko and realizes he is sad..."

[0677] Step 11:

[0678] The user interacts with the child to discuss the content of the story and develop the child's emotions.

[0679] Example: "How did Ravi know Nico was sad?" helps children develop an understanding of emotions.

[0680] Example 2

[0681] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0682] A major challenge is the lack of teaching materials and tools for effectively educating children about emotions. There are also limited ways for parents to easily create picture books suitable for emotional education and use them to help their children understand emotions. Furthermore, as digital devices are used, there is a need to support the development of emotions and interpersonal relationships while reducing visual burden.

[0683] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a recognition means for recognizing a user's emotion, a story generation means for generating a story based on the recognized emotion, an image generation means for generating images based on the generated story, and a display means for displaying the generated story and images. This makes it possible to automatically generate a personalized picture book according to the user's emotion and display it using an electronic paper device, thereby enabling effective emotional education for children while reducing visual burden.

[0684] The "recognition means" is a means for recognizing the user's emotions.

[0685] A "narrative generation means" is a means for generating a narrative based on recognized emotions.

[0686] The "image generation means" is a means for generating an image based on the generated story.

[0687] "Display means" is a means for displaying the generated story and images.

[0688] The "selection means" is a means for selecting an emotion area based on a user's input.

[0689] "Electronic paper" is a backlit display device used to reduce visual strain.

[0690] The "transmission means" is a means for transmitting the generated story to the terminal.

[0691] The "reading support tool" is a tool for reading the generated story to a child.

[0692] This invention is a system that automatically generates picture books for children with the aim of emotional education, and supports the development of children's emotions and interpersonal relationships through reading aloud. In addition to selecting emotional domains based on user input, this system incorporates an emotion engine to recognize the user's emotions and generate more personalized picture books.

[0693] First, the user launches the application and manually selects an emotion area, or the device automatically recognizes the user's emotion using an emotion engine. For example, the user manually selects "empathy," or the emotion engine recognizes the user's facial expression and voice as "calm" and selects "empathy."

[0694] Next, the device sends the selected emotional domain to the server. The server analyzes this data and confirms that the emotional domain is "empathy." The "empathy" emotional domain is input into the generative AI model, which then generates a story. As a concrete example, the generative AI model uses a rabbit character named "Rabi" to generate a story with an empathy theme.

[0695] Some examples of generated stories:

[0696] One day, Ravi, a little rabbit, finds his friend Niko sad, and he wants to know why Niko is sad.

[0697] The server then generates image prompts based on the generated story. For example, it generates a prompt for "a scene where Ravi and Nico are together." The server then passes the prompt to an image generation AI model, which generates an appropriate image. For example, the image generation AI generates an image of a scene where "Ravi and Nico are sitting together."

[0698] Finally, the server compiles the story and the generated images into a single document and sends it to the device. The device then analyzes the received data and displays it on the e-paper device. For example, the e-paper device displays "a story that shows the empathy between Ravi and Nico and the corresponding images."

[0699] The user uses the e-paper device to read a story to the child. For example, a parent might read part of the story, "Ravi looks at his friend Nico and realizes he's sad..." The user interacts with the child, discussing the content of the story and developing the child's emotions. For example, the parent might ask, "How did Ravi realize Nico was sad?", helping the child deepen their understanding of emotions.

[0700] This system allows parents to easily create picture books suitable for emotional education and read them to their children. Using an emotion engine, it provides a more personalized experience according to the user's mood and state. Using an e-paper device, it provides a visual experience that is gentle on children's eyes while supporting the development of emotions and interpersonal relationships.

[0701] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0702] Step 1:

[0703] The user launches the application and selects an emotion area, or the terminal uses an emotion engine to recognize the user's emotion.

[0704] Input: Emotion area selection screen display, user's facial expression data and voice data.

[0705] Data processing: Users can select emotion areas from a drop-down menu, or the emotion engine analyzes facial expressions and voice.

[0706] Output: Selected emotion domain data.

[0707] What happens: The user operates their smartphone to launch the app and taps the "Empathy" button. The device uses the camera to scan the user's facial expressions and performs voice recognition.

[0708] Step 2:

[0709] The terminal transmits the selected emotion area to the server.

[0710] Input: Emotion domain data.

[0711] Data processing: Packaging and sending emotional domain data.

[0712] Output: Emotional domain data sent to the server.

[0713] Specific operation: The device sends data on the emotional domain "empathy" to the server via the emotion engine API.

[0714] Step 3:

[0715] The server analyzes the received emotional domain data and generates a story for the picture book using a story generation means.

[0716] Input: Emotion domain data.

[0717] Data processing: Analysis of emotional domain data, input into generative AI models, and narrative generation.

[0718] Output: The generated narrative text.

[0719] Specific operation: The server inputs the emotional domain data "empathy" into the generative AI model and generates a story themed around empathy using a character named "Rabi."

[0720] Step 4:

[0721] The server generates image prompts based on the generated story and generates images using an image generation means.

[0722] Input: Narrative text.

[0723] Data processing: Generate image prompts from narrative text, input them into an image generation AI model, and generate images.

[0724] Output: The generated image data.

[0725] Specific operation: The server creates a prompt describing "a scene where Ravi and Nico are together" and inputs it into an image generation AI model to generate an image.

[0726] Step 5:

[0727] The server compiles the generated story text and images into a single document and sends it to the terminal.

[0728] Input: Narrative text and image data.

[0729] Data processing: Organizing data into document formats such as PDF and sending it to the device.

[0730] Output: The document data sent to the device.

[0731] Specific operation: The server compiles the story text and images into PDF format and sends it to the terminal.

[0732] Step 6:

[0733] The terminal analyzes the received data and displays it on the electronic paper device.

[0734] Input: The received document data.

[0735] Data processing: Analyzing document data and preparing it for display on e-paper devices.

[0736] Output: Story and images displayed on an e-paper device.

[0737] Specific operation: The terminal displays the received PDF on the e-paper device.

[0738] Step 7:

[0739] A user reads a story to a child using an e-paper device.

[0740] Input: A story and an image displayed on an e-paper device.

[0741] Data processing: Reading aloud was conducted.

[0742] Output: Providing emotional education to children.

[0743] Specific Action: The user reads the story "Ravi looks at his friend Niko and realizes he is sad..." and interacts with the child to deepen their understanding of emotions.

[0744] (Application example 2)

[0745] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0746] Previous story generation systems were limited in their ability to provide appropriate stories and images to support emotional education and interpersonal relationship development. They also struggled to recognize users' emotions in real time and provide personalized stories and images based on those emotions. Furthermore, the means to interactively experience the generated stories were limited, limiting the enrichment of children's learning experiences.

[0747] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a generation means for generating a story for each emotion, an image generation means for generating images based on the generated story, a display means for displaying the generated story and images, an emotion recognition means for automatically recognizing emotions, an audio playback means for reading aloud the generated story, and a cross-platform support means for supporting multiple display devices. This makes it possible to recognize the user's emotions in real time and provide personalized stories and images based on them, thereby providing a more interactive story-telling experience.

[0748] The "generation means for generating a story for each emotion" is a means for automatically generating a story corresponding to a specific emotion based on a user's selection and emotion recognition data.

[0749] The "image generating means for generating images based on the generated story" is a means for automatically generating images corresponding to the content of the generated story.

[0750] A "display means for displaying the generated story and images" is a device or interface for visually presenting the generated story and images to a user.

[0751] "Emotion recognition means for automatically recognizing emotions" is a means for automatically analyzing and recognizing emotions from the user's facial expressions and voice using sensors such as a camera and a microphone.

[0752] The "audio playback means for reading aloud" is a means for playing back the generated story aloud and providing it to the user audibly.

[0753] "Cross-platform support for multiple display devices" refers to the means by which the generated story and images are adjusted to be displayed appropriately on different types of devices (smartphones, tablets, smart glasses, head-mounted displays, etc.).

[0754] The "selection means for selecting an emotional area based on user input" is a means that allows a user to specify a particular emotional area based on their current feelings and intentions.

[0755] The "display means without backlight" refers to a means for displaying without using a backlight in order to reduce power consumption.

[0756] "Display means using electronic paper" refers to a display device that uses electronic ink technology and combines the ease of viewing as paper with power-saving performance.

[0757] This invention is a system that automatically generates stories and images for children using real-time emotion recognition and generation AI for the purpose of emotional education. The system automatically recognizes the user's emotions, generates a story based on those emotions, and then generates corresponding images. The generated stories and images are displayed on multiple display devices, providing an interactive reading experience through a text-to-speech function.

[0758] Hardware and software used

[0759] 1. Hardware

[0760] Smartphone

[0761] tablet

[0762] Smart glasses (e.g. Google Glass)

[0763] Head-mounted displays (e.g., Oculus Quest)

[0764] Electronic Paper Device

[0765] 2. Software

[0766] Azure Cognitive Services (for emotion recognition)

[0767] AWS Lambda (serverless computing)

[0768] S3 bucket (data storage)

[0769] OpenAI GPT-4 (a generative AI model for story generation)

[0770] DALL-E 2 (generative AI model for image generation)

[0771] React Native (for cross-platform compatibility)

[0772] Amazon Polly (for voice reading)

[0773] System Overview

[0774] 1. User Emotion Recognition

[0775] When a user launches an application, Azure Cognitive Services uses the device's camera and microphone to recognize the user's emotions. For example, if facial expression recognition technology determines that the user is "calm," this information is sent to the server.

[0776] 2. Narrative Generation Based on Emotional Data

[0777] Once emotion recognition data is sent to the server, OpenAI GPT-4 is invoked using AWS Lambda. Using emotion data (e.g., "empathy") as input, the generative AI model generates a related story. As a specific example, it generates an empathetic story using a rabbit character named "Rabi."

[0778] 3. Image Generation

[0779] Based on the content of the generated story, images that fit the story are generated using DALL-E 2. For example, a scene in the story where Ravi and Nico are sitting together is depicted.

[0780] 4. Data transmission and display

[0781] The generated stories and images are compiled into documents and sent using React Native to smartphones, tablets, smart glasses, head-mounted displays, and potentially even backlit e-paper devices.

[0782] 5. Reading aloud

[0783] The story is generated using Amazon Polly and read aloud, and an interface is provided for users to enjoy the story interactively.

[0784] Examples of concrete examples and prompts

[0785] For example, a user launches an app and Azure Cognitive Services determines that the user feels "calm." Emotional data "empathy" is sent to the server, OpenAI GPT-4 generates an "empathetic story about Ravi and Nico," and DALL-E 2 generates an image of "a scene where Ravi and Nico are sitting together." The generated story and image are sent to the device and read aloud.

[0786] An example prompt is:

[0787] "Generate stories with empathy as a theme based on emotional data"

[0788] "Based on the story you've created, generate an image of the rabbit character and his friends together."

[0789] The present invention generates stories and images that correspond to the user's real-time emotions, making it possible to support children's emotional and interpersonal development through an interactive storytelling experience.

[0790] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0791] Step 1:

[0792] The device launches an application and activates the device's camera and microphone to automatically recognize the user's emotions. Azure Cognitive Services is used to analyze data obtained from the camera and microphone and recognize the user's emotions. For example, a facial expression recognition algorithm determines that the user is "calm." The input is the user's facial expression and voice data, and the output is the recognized emotion data.

[0793] Step 2:

[0794] The device sends the recognized emotion data to the server. Specifically, the emotion data called "empathy" is sent in JSON format. The input is the emotion data output from Step 1, and the output is a notification of completion of data transmission to the server.

[0795] Step 3:

[0796] Based on the emotion data received by the server, AWS Lambda is triggered to generate a story. The Lambda function calls OpenAI GPT-4 and uses the emotion data as input to create a story generation prompt. The input is emotion data and the story generation prompt, and the output is the generated story text. Specifically, the prompt text "Generate a story with the theme of empathy" with the emotion "empathy" is sent to GPT-4, which then generates a story.

[0797] Step 4:

[0798] The server uses DALL-E 2 to generate images based on the generated story text. Specific scenes from the story are sent as prompts to DALL-E 2, which then generates corresponding images. The input is the story text and the image generation prompt, and the output is the generated image file. For example, an image is generated based on the prompt "A scene where Ravi and Nico are sitting together."

[0799] Step 5:

[0800] The server compiles the generated story text and images into a single document and sends it to the terminal. The input is the story text and image files, and the output is the compiled document file and a notification of completion of transmission. The data is compiled into a JSON format document file and sent to the terminal.

[0801] Step 6:

[0802] The device analyzes the received document file and provides it to the user through a display method. For example, using React Native, it can be displayed on a smartphone, tablet, smart glasses, or head-mounted display. The input is the document file, and the output is the story and images displayed on the device.

[0803] Step 7:

[0804] The device uses Amazon Polly to read aloud the story text generated, allowing users to enjoy the story not only visually but also aurally. The input is the story text, and the output is audio data, specifically the content of the story being read aloud.

[0805] Each step is processed sequentially to generate a personalized story and images based on the user's emotions, ultimately providing an interactive storytelling experience.

[0806] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0807] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0808] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0809] [Third embodiment]

[0810] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0811] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0812] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0813] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0814] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0815] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0816] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0817] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0818] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0819] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0820] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0821] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0822] This invention is a system that automatically generates picture books for children and supports the development of their emotions and interpersonal relationships through reading aloud, with the aim of emotional education. The system provides an interface that is easy for users to operate, and generates and displays stories and images based on emotional domains.

[0823] System Overview

[0824] 1. User operations

[0825] The user launches the application and selects an emotional area.

[0826] Example: A user selects "empathy" within an application.

[0827] The terminal transmits the user's selection data to the server.

[0828] 2. Story generation based on emotional domains

[0829] The server analyzes the emotion area data received from the device.

[0830] Example: The server checks that the emotion domain is "empathy."

[0831] Emotional domains are input into the generative AI model on the server to generate a story.

[0832] Example: An AI model generates a story about a little rabbit character named Ravi and his friend Nico, with a theme of empathy.

[0833] Some of the stories generated:

[0834] One day, Ravi, a little rabbit, finds his friend Niko sad, and he wants to know why Niko is sad.

[0835] 3. Narrative-based image generation

[0836] The server generates image prompts based on the generated story.

[0837] Example: Generate a prompt for "A scene with Ravi and Nico together."

[0838] The server passes the prompt to an image generation AI model, which generates an appropriate image.

[0839] Example: An image generation AI generates an image depicting a scene in which "Ravi and Nico are sitting together."

[0840] 4. Data transmission and display

[0841] The server compiles the generated story and images and sends them to the device.

[0842] Example: Sending story text and image files to a device.

[0843] The terminal analyzes the received data and displays it on the electronic paper device.

[0844] Example: An e-paper device displays "a story and corresponding images that show the relatable nature of Ravi and Nico."

[0845] 5. Reading aloud

[0846] A user reads a story to a child using an e-paper device.

[0847] Example: A parent reads part of a story: "Ravi looks at his friend Niko and realizes he is sad..."

[0848] The user interacts with the child to discuss the content of the story and develop the child's emotions.

[0849] Example: "How did Ravi know Nico was sad?" helps children develop an understanding of emotions.

[0850] This system allows parents to easily create picture books suitable for emotional education and read them to their children. By using an e-paper device, it provides a visual experience that is gentle on children's eyes while supporting the development of emotions and interpersonal relationships.

[0851] The processing flow will be explained below.

[0852] Step 1:

[0853] The user launches the application and selects an emotional area.

[0854] Example: A user selects "empathy" within an application.

[0855] Step 2:

[0856] The terminal transmits the user's selection data to the server.

[0857] Example: Send data from the emotion field "empathy" to the server.

[0858] Step 3:

[0859] The server analyzes the emotion area data received from the device.

[0860] Example: The server checks that the emotion domain is "empathy."

[0861] Step 4:

[0862] The emotional domain of "empathy" is input into the generative AI model on the server to generate a story.

[0863] Example: An AI model generates stories with empathy as the theme, using a rabbit character named "Rabi."

[0864] Some of the stories generated:

[0865] One day, Ravi, a little rabbit, finds his friend Niko sad, and he wants to know why Niko is sad.

[0866] Step 5:

[0867] The server generates image prompts based on the generated story.

[0868] Example: Generate a prompt for "A scene with Ravi and Nico together."

[0869] Step 6:

[0870] The server passes the prompt to an image generation AI model, which generates an appropriate image.

[0871] Example: An image generation AI generates an image of a scene where "Ravi and Nico are sitting together."

[0872] Step 7:

[0873] The server compiles the generated story and images into a single document.

[0874] Example: Combining generated narrative text and images.

[0875] Examples of compiled data:

[0876] Story text:

[0877] One day, a little rabbit named Ravi found his friend Niko sad.

[0878] Story Image:

[0879] Image file of Ravi and Nico sitting together

[0880] Step 8:

[0881] The server sends the compiled data to the terminal.

[0882] Example: Sending a story and image file to a device.

[0883] Step 9:

[0884] The terminal analyzes the received data and displays it on the electronic paper device.

[0885] Example: An e-paper device displays "a story and corresponding images that show the relatable nature of Ravi and Nico."

[0886] Step 10:

[0887] A user reads a story to a child using an e-paper device.

[0888] Example: A parent reads part of a story: "Ravi looks at his friend Niko and realizes he is sad..."

[0889] Step 11:

[0890] The user interacts with the child to discuss the content of the story and develop the child's emotions.

[0891] Example: "How did Ravi know Nico was sad?" helps children develop an understanding of emotions.

[0892] Example 1

[0893] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0894] Previous emotional education systems faced challenges, such as a lack of an easy-to-use interface for users and the lack of the ability to automatically generate individual stories and images based on emotional domains. Furthermore, selecting a device that provides a visual experience that is easy on children's eyes is also important as a means of displaying the generated content, but this was often not properly implemented. As a result, there were problems with the creation of teaching materials suitable for emotional education and the learning process using them not being effectively supported.

[0895] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0896] In this invention, the server includes a selection means for a user to select an emotional area, a story generation means including a generative AI model for generating a story based on the selected emotional area, an image generation means including an image generation AI model for generating image prompts based on the generated story and generating images based on the prompts, a transmission means for transmitting the generated story and images to the user's terminal, and a display means for displaying the transmitted story and images. This makes it possible to automatically generate individual stories and images based on emotions using an interface that is easy for users to operate, and to display the generated content on a device that is easy on children's eyes.

[0897] A "selection means" is a device or software that provides an interface for a user to select a particular emotional domain.

[0898] A "generative AI model" is an artificial intelligence model that automatically generates stories and sentences based on selected emotional domains.

[0899] The "story generation means" is a means for generating a story using a generative AI model with emotional domain data selected via a selection means as input.

[0900] A "prompt" is a textual instruction given to an image-generating AI model to generate a specific image.

[0901] An "image generation AI model" is an artificial intelligence model that automatically generates corresponding images based on given prompts.

[0902] The "image generation means" is a means for generating an image generation prompt based on the story generated by the story generation means, and passing the prompt to an image generation AI model to generate an image.

[0903] The "transmission means" is a means for transmitting the generated story and images from the server to the user's terminal.

[0904] "Display means" refers to a device that displays the received story and images on the user's terminal, allowing the user to check the content.

[0905] "Electronic paper" is a visually friendly display technology that can display text and images without the use of a backlight.

[0906] This system automatically generates picture books for children with the aim of emotional education, and supports the development of children's emotions and interpersonal relationships through reading aloud. This system consists of a terminal operated by the user, a server that processes data, and a display means that displays the generated stories and images.

[0907] First, the user launches the emotion education system application on a smartphone or tablet device (e.g., iOS or Android). The user selects an emotion area through the application interface. For example, the user selects "empathy."

[0908] Next, the device sends the data of the user's selected emotional domain to the server. The server uses a cloud server (e.g., AWS, GCP) to analyze the received data and input the emotional domain data into a generative AI model (e.g., OpenAI GPT-4). This generates a story about the selected emotional domain. As a specific example, a story based on the theme of "empathy" might be generated, such as "The Story of Little Rabbit Ravi and His Friend Nico."

[0909] The generated story is then used to generate image generation prompts. The server creates image generation prompts based on the content of the story and passes them to an image generation AI model (e.g., DALL-E, MidJourney). This automatically generates images corresponding to scenes in the story. For example, an image depicting "Ravi and Nico sitting together" is generated.

[0910] The generated story and images are sent from the server to the device, which then analyzes the data and displays it on an electronic paper device (e.g., Kindle, e-ink device). This display device does not use a backlight to provide a visual experience that is gentle on children's eyes.

[0911] Finally, the user reads a story to the child using the e-paper device. For example, a parent might read the story, "Ravi looks at his friend Nico and realizes he's sad..." In addition, the user can deepen the child's understanding of emotions by interacting with the child and discussing the content of the story. For example, emotional education is provided by asking, "How did Ravi realize Nico was sad?"

[0912] As a result, this system provides an easy-to-use interface for users, automatically generates individual stories and images based on emotional domains, and displays them on an electronic paper device that is easy on children's eyes. By utilizing a generative AI model and an image-generating AI model, the present invention is characterized by its ability to quickly provide high-quality content suitable for emotional education.

[0913] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0914] Step 1:

[0915] The user launches the application and selects an emotion area.

[0916] Input: The user launches the application on their smartphone or tablet and selects an emotion domain (e.g., "empathy").

[0917] How it works: The application displays an interface for the user to select an emotional area. When the user selects an emotional area, the information is temporarily stored on the device.

[0918] Output: Data from the selected emotion domains is generated and passed to the next processing step.

[0919] Step 2:

[0920] The device sends the emotion area data to the server.

[0921] Input: Emotion domain data selected by the user in Step 1 (e.g., "Empathy")

[0922] How it works: The device sends emotional domain data to the server using an API request over the internet. The data is securely sent to the server, where analysis begins.

[0923] Output: Emotional domain data is sent to the server.

[0924] Step 3:

[0925] The server analyzes the received emotional data and generates a story.

[0926] Input: Emotion domain data sent to the server in step 2 (e.g., "Empathy")

[0927] How it works: The server analyzes the received emotional domain data and inputs it into a generative AI model (e.g., OpenAI GPT-4). The generative AI model generates a story based on the emotional domain. For example, it generates "The Story of Little Rabbit Ravi and His Friend Nico."

[0928] Output: The generated story text (e.g., "One day, Ravi the little rabbit found his friend Niko sad...")

[0929] Step 4:

[0930] The server generates image prompts based on the story and generates images.

[0931] Input: The narrative text generated in step 3

[0932] How it works: The server generates an image prompt based on the content of the generated story (e.g., "A scene where Ravi and Nico are together"). It then passes the image prompt to an image generation AI model (e.g., DALL-E, MidJourney). The image generation AI model generates an image based on the prompt. For example, it generates an image of "A scene where Ravi and Nico are sitting together."

[0933] Output: Generated image file (e.g. JPEG, PNG format)

[0934] Step 5:

[0935] The server sends the generated story and images to the device.

[0936] Input: The story text generated in step 3 and the image files generated in step 4

[0937] How it works: The server compiles the generated story text and image files and sends them to the user's device via an API request over the internet.

[0938] Output: A data package containing the story and images

[0939] Step 6:

[0940] The terminal analyzes the received data and displays it on the e-paper device.

[0941] Input: The data package sent by the server in step 5 (story text and image files)

[0942] How it works: The device analyzes the received data and displays the story text and images on an electronic paper device (e.g., Kindle, e-ink device). Electronic paper devices display without a backlight, providing a visual experience that is gentle on children's eyes.

[0943] Output: The displayed story and images

[0944] Step 7:

[0945] A user reads a story to a child on an e-paper device

[0946] Input: The story and image displayed in step 6

[0947] How it works: A user reads a story to a child using an e-paper device. For example, a parent reads, "Ravi looks at his friend Niko and realizes he's sad..." The user then interacts with the child, asking questions about the story to deepen the child's understanding of their emotions.

[0948] Output: Responses of children who received emotional education

[0949] The above are the specific processing steps of the program of this system.

[0950] (Application example 1)

[0951] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0952] The present invention relates to a system that automatically generates picture books to support children's emotional education. With conventional systems, it is difficult to find the time and timing for reading aloud, making it difficult to provide effective emotional education, especially when parents are busy or children have to wait long periods of time. Furthermore, there is a lack of a way to effectively utilize spare time, such as while waiting for delivery. Therefore, an objective of the present invention is to provide a system that can provide emotional education by effectively utilizing waiting time for delivery.

[0953] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0954] In this invention, the server includes a generation means for generating a story for each emotion, an image generation means for generating an image based on the generated story, a display means for displaying the generated story and image, and a promotion means for generating a story and image for each emotion, displaying them when a user places an order, and promoting reading aloud while waiting for delivery. This makes it possible to effectively utilize the waiting time for delivery to provide emotional education.

[0955] "Emotional stories" are stories generated based on specific emotional domains, featuring situations and characters related to those emotions.

[0956] "Generation means" refers to a device or program that has the function of automatically generating stories and images based on the emotional domain selected by the user.

[0957] "Image generation means" refers to a device or program that has the function of automatically generating images of scenes and characters related to the generated story.

[0958] "Display means" refers to a device or function that displays the generated story or images so that the user can view them.

[0959] "Promotion means" refers to a device or function that displays the story and images generated when the user places an order and encourages reading aloud while waiting for delivery.

[0960] The "selection means" refers to a device or program that provides an interface or function for the user to select a desired emotional area.

[0961] "Transmission means" refers to a device or program for transmitting the generated story and images to the user's terminal.

[0962] "Delivery Order" means an online or offline order placed by a User for delivery of food or merchandise.

[0963] This invention is a system that automatically generates picture books for the purpose of emotional education and supports the development of children's emotions and interpersonal relationships through reading aloud. This system promotes emotional education by allowing parents to read picture books to their children, making effective use of waiting times for delivery services.

[0964] 1. System Configuration

[0965] The system includes the following major components:

[0966] Server: Includes a generative AI model and an image generation model that generates stories and images based on emotional domains.

[0967] Device: A device on which parents place delivery orders and view the generated stories and images. Specifically, this refers to a smartphone or tablet.

[0968] User: In this case, the parent who places the delivery order and reads to their child.

[0969] 2. Explanation of program processing

[0970] 1. User operations

[0971] A user launches a food delivery app, orders a meal, and selects the "Emotional Education" option, where the user selects a specific emotional domain (e.g., empathy).

[0972] 2. Story Generation Based on Emotional Domains

[0973] The user's selection data is sent to a server, which uses a generative AI model (e.g., OpenAI GPT-4) to generate a narrative based on the selected emotional domains.

[0974] Example: Generate a story about a little rabbit character named "Ravi" and his friend "Nico" that has an empathy theme. The prompt is:

[0975] Input prompt: "Generate a story about Ravi and Nico based on empathy."

[0976] 3. Narrative-based image generation

[0977] Based on the generated story, the server generates appropriate images using an image generation model (e.g., DALL-E 2).

[0978] Example: An image depicting "Ravi and Nico together" is generated.

[0979] 4. Data transmission and display

[0980] The server sends the generated story and images to the user's terminal, which displays the received story and images.

[0981] Parents can read stories to their children while they wait for their delivery.

[0982] 5. Reading aloud and emotional education

[0983] Emotional education is carried out by parents reading stories and discussing emotions with their children through dialogue.

[0984] For example, asking, "How did Ravi know Nico was sad?" provides an opportunity for children to think about empathy.

[0985] 3. Hardware and Software Use

[0986] The server uses a generative AI model (OpenAI GPT-4) for story generation and a generative image model (DALL-E 2) for image generation.

[0987] The user's device runs a delivery application (e.g., a generic food delivery service app) and has a display for displaying stories and images.

[0988] The server receives the user's selection data and inputs the data into the AI ​​model for processing.

[0989] This makes it possible to provide a system that makes effective use of the waiting time for delivery and allows parents and children to enjoy reading picture books suitable for emotional education.

[0990] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0991] Step 1:

[0992] A user launches a food delivery app and selects the "Emotional Education" option when ordering food. The user also selects an emotional domain, such as "Empathy," through an interface that identifies emotional domains. This generates food delivery order data and emotional domain data.

[0993] Step 2:

[0994] The terminal transmits the emotional domain data selected by the user to the server. The server receives the user's selection data and analyzes the emotional domains. This data includes the emotional domain selected by the user, such as "empathy."

[0995] Step 3:

[0996] The server inputs the emotional domain data into a generative AI model (e.g., OpenAI GPT-4) for generating stories based on emotional domains. The server uses the AI ​​model to generate stories based on specific emotional domains. The input is a prompt statement: "Generate a story about Ravi and Nico based on empathy." The output is a story text with an empathy theme.

[0997] Step 4:

[0998] The server analyzes the generated story text and generates image prompts based on the story content. Image prompts are text data containing specific scenes in the story (e.g., "A scene where Ravi and Nico are together"). These prompts are input into an image generation AI model (e.g., DALL-E 2). The result is an image depicting a scene from the story.

[0999] Step 5:

[1000] The server compiles the generated story and image data and sends them in digital format to the terminal, which then retrieves the received story data and image data and displays them to the user in an appropriate format, allowing the user to view the story and corresponding images while waiting for delivery.

[1001] Step 6:

[1002] The user reads a story displayed on the device to the child. The user then interacts with the child about the story's contents and provides emotional education. Specifically, by discussing the actions and emotions of the characters in the story, the child can deepen their understanding of emotions.

[1003] Through these steps, parents and children can effectively utilize the time they spend waiting for delivery by reading picture books suitable for emotional education.

[1004] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1005] This invention is a system that automatically generates picture books for children with the aim of emotional education, and supports the development of children's emotions and interpersonal relationships through reading aloud. In addition to selecting emotional domains based on user input, this system incorporates an emotion engine that recognizes the user's emotions to generate more personalized picture books.

[1006] System Overview

[1007] 1. User operations

[1008] The user launches the application and manually selects an emotion area, or the device automatically recognizes the user's emotion using an emotion engine.

[1009] Example: The user manually selects "empathy," or the emotion engine recognizes the user's facial expression and voice as "calm" and selects "empathy."

[1010] 2. Story generation based on emotional domains

[1011] The server analyzes the emotion area data received from the device.

[1012] Example: The server checks that the emotion domain is "empathy."

[1013] The emotional domain of "empathy" is input into the generative AI model on the server to generate a story.

[1014] Example: An AI model generates stories with empathy as the theme, using a rabbit character named "Rabi."

[1015] Some of the stories generated:

[1016] One day, Ravi, a little rabbit, finds his friend Niko sad, and he wants to know why Niko is sad.

[1017] 3. Narrative-based image generation

[1018] The server generates image prompts based on the generated story.

[1019] Example: Generate a prompt for "A scene with Ravi and Nico together."

[1020] The server passes the prompt to an image generation AI model, which generates an appropriate image.

[1021] Example: An image generation AI generates an image of a scene where "Ravi and Nico are sitting together."

[1022] 4. Data transmission and display

[1023] The server compiles the generated story and images into a single document and sends it to the device.

[1024] Example: Sending story text and image files to the device.

[1025] The terminal analyzes the received data and displays it on the electronic paper device.

[1026] Example: An e-paper device displays "a story and corresponding images that show the relatable nature of Ravi and Nico."

[1027] 5. Reading aloud

[1028] A user reads a story to a child using an e-paper device.

[1029] Example: A parent reads part of a story: "Ravi looks at his friend Niko and realizes he is sad..."

[1030] The user interacts with the child to discuss the content of the story and develop the child's emotions.

[1031] Example: "How did Ravi know Nico was sad?" helps children develop an understanding of emotions.

[1032] This system allows parents to easily create picture books suitable for emotional education and read them to their children. Using an emotion engine, it provides a more personalized experience according to the user's mood and state. Using an e-paper device, it provides a visual experience that is gentle on children's eyes while supporting the development of emotions and interpersonal relationships.

[1033] The processing flow will be explained below.

[1034] Step 1:

[1035] When a user launches the application, the emotion engine recognizes the user's emotion, and the device displays an emotion area based on the user's input.

[1036] Example: A user launches an application, and the emotion engine selects "empathy" based on the user's facial expressions and voice.

[1037] Step 2:

[1038] The terminal transmits the user's selection data to the server.

[1039] Example: Send data from the emotion field "empathy" to the server.

[1040] Step 3:

[1041] The server analyzes the emotion area data received from the device.

[1042] Example: The server checks that the emotion domain is "empathy."

[1043] Step 4:

[1044] The emotional domain of "empathy" is input into the generative AI model on the server to generate a story.

[1045] Example: An AI model generates stories with empathy as the theme, using a rabbit character named "Rabi."

[1046] Some of the stories generated:

[1047] One day, Ravi, a little rabbit, finds his friend Niko sad, and he wants to know why Niko is sad.

[1048] Step 5:

[1049] The server generates image prompts based on the generated story.

[1050] Example: Generate a prompt for "A scene with Ravi and Nico together."

[1051] Step 6:

[1052] The server passes the prompt to an image generation AI model, which generates an appropriate image.

[1053] Example: An image generation AI generates an image of a scene where "Ravi and Nico are sitting together."

[1054] Step 7:

[1055] The server compiles the generated story and images into a single document.

[1056] Example: Combining generated narrative text and images.

[1057] Examples of compiled data:

[1058] Story text:

[1059] One day, a little rabbit named Ravi found his friend Niko sad.

[1060] Story Image:

[1061] Image file of Ravi and Nico sitting together

[1062] Step 8:

[1063] The server sends the compiled data to the terminal.

[1064] Example: Sending a story and image file to a device.

[1065] Step 9:

[1066] The terminal analyzes the received data and displays it on the electronic paper device.

[1067] Example: An e-paper device displays "a story and corresponding images that show the relatable nature of Ravi and Nico."

[1068] Step 10:

[1069] A user reads a story to a child using an e-paper device.

[1070] Example: A parent reads part of a story: "Ravi looks at his friend Niko and realizes he is sad..."

[1071] Step 11:

[1072] The user interacts with the child to discuss the content of the story and develop the child's emotions.

[1073] Example: "How did Ravi know Nico was sad?" helps children develop an understanding of emotions.

[1074] Example 2

[1075] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1076] A major challenge is the lack of teaching materials and tools for effectively educating children about emotions. There are also limited ways for parents to easily create picture books suitable for emotional education and use them to help their children understand emotions. Furthermore, as digital devices are used, there is a need to support the development of emotions and interpersonal relationships while reducing visual burden.

[1077] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a recognition means for recognizing a user's emotion, a story generation means for generating a story based on the recognized emotion, an image generation means for generating images based on the generated story, and a display means for displaying the generated story and images. This makes it possible to automatically generate a personalized picture book according to the user's emotion and display it using an electronic paper device, thereby enabling effective emotional education for children while reducing visual burden.

[1078] The "recognition means" is a means for recognizing the user's emotions.

[1079] A "narrative generation means" is a means for generating a narrative based on recognized emotions.

[1080] The "image generation means" is a means for generating an image based on the generated story.

[1081] "Display means" is a means for displaying the generated story and images.

[1082] The "selection means" is a means for selecting an emotion area based on a user's input.

[1083] "Electronic paper" is a backlit display device used to reduce visual strain.

[1084] The "transmission means" is a means for transmitting the generated story to the terminal.

[1085] The "reading support tool" is a tool for reading the generated story to a child.

[1086] This invention is a system that automatically generates picture books for children with the aim of emotional education, and supports the development of children's emotions and interpersonal relationships through reading aloud. In addition to selecting emotional domains based on user input, this system incorporates an emotion engine to recognize the user's emotions and generate more personalized picture books.

[1087] First, the user launches the application and manually selects an emotion area, or the device automatically recognizes the user's emotion using an emotion engine. For example, the user manually selects "empathy," or the emotion engine recognizes the user's facial expression and voice as "calm" and selects "empathy."

[1088] Next, the device sends the selected emotional domain to the server. The server analyzes this data and confirms that the emotional domain is "empathy." The "empathy" emotional domain is input into the generative AI model, which then generates a story. As a concrete example, the generative AI model uses a rabbit character named "Rabi" to generate a story with an empathy theme.

[1089] Some examples of generated stories:

[1090] One day, Ravi, a little rabbit, finds his friend Niko sad, and he wants to know why Niko is sad.

[1091] The server then generates image prompts based on the generated story. For example, it generates a prompt for "a scene where Ravi and Nico are together." The server then passes the prompt to an image generation AI model, which generates an appropriate image. For example, the image generation AI generates an image of a scene where "Ravi and Nico are sitting together."

[1092] Finally, the server compiles the story and the generated images into a single document and sends it to the device. The device then analyzes the received data and displays it on the e-paper device. For example, the e-paper device displays "a story that shows the empathy between Ravi and Nico and the corresponding images."

[1093] The user uses the e-paper device to read a story to the child. For example, a parent might read part of the story, "Ravi looks at his friend Nico and realizes he's sad..." The user interacts with the child, discussing the content of the story and developing the child's emotions. For example, the parent might ask, "How did Ravi realize Nico was sad?", helping the child deepen their understanding of emotions.

[1094] This system allows parents to easily create picture books suitable for emotional education and read them to their children. Using an emotion engine, it provides a more personalized experience according to the user's mood and state. Using an e-paper device, it provides a visual experience that is gentle on children's eyes while supporting the development of emotions and interpersonal relationships.

[1095] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1096] Step 1:

[1097] The user launches the application and selects an emotion area, or the terminal uses an emotion engine to recognize the user's emotion.

[1098] Input: Emotion area selection screen display, user's facial expression data and voice data.

[1099] Data processing: Users can select emotion areas from a drop-down menu, or the emotion engine analyzes facial expressions and voice.

[1100] Output: Selected emotion domain data.

[1101] What happens: The user operates their smartphone to launch the app and taps the "Empathy" button. The device uses the camera to scan the user's facial expressions and performs voice recognition.

[1102] Step 2:

[1103] The terminal transmits the selected emotion area to the server.

[1104] Input: Emotion domain data.

[1105] Data processing: Packaging and sending emotional domain data.

[1106] Output: Emotional domain data sent to the server.

[1107] Specific operation: The device sends data on the emotional domain "empathy" to the server via the emotion engine API.

[1108] Step 3:

[1109] The server analyzes the received emotional domain data and generates a story for the picture book using a story generation means.

[1110] Input: Emotion domain data.

[1111] Data processing: Analysis of emotional domain data, input into generative AI models, and narrative generation.

[1112] Output: The generated narrative text.

[1113] Specific operation: The server inputs the emotional domain data "empathy" into the generative AI model and generates a story themed around empathy using a character named "Rabi."

[1114] Step 4:

[1115] The server generates image prompts based on the generated story and generates images using an image generation means.

[1116] Input: Narrative text.

[1117] Data processing: Generate image prompts from narrative text, input them into an image generation AI model, and generate images.

[1118] Output: The generated image data.

[1119] Specific operation: The server creates a prompt describing "a scene where Ravi and Nico are together" and inputs it into an image generation AI model to generate an image.

[1120] Step 5:

[1121] The server compiles the generated story text and images into a single document and sends it to the terminal.

[1122] Input: Narrative text and image data.

[1123] Data processing: Organizing data into document formats such as PDF and sending it to the device.

[1124] Output: The document data sent to the device.

[1125] Specific operation: The server compiles the story text and images into PDF format and sends it to the terminal.

[1126] Step 6:

[1127] The terminal analyzes the received data and displays it on the electronic paper device.

[1128] Input: The received document data.

[1129] Data processing: Analyzing document data and preparing it for display on e-paper devices.

[1130] Output: Story and images displayed on an e-paper device.

[1131] Specific operation: The terminal displays the received PDF on the e-paper device.

[1132] Step 7:

[1133] A user reads a story to a child using an e-paper device.

[1134] Input: A story and an image displayed on an e-paper device.

[1135] Data processing: Reading aloud was conducted.

[1136] Output: Providing emotional education to children.

[1137] Specific Action: The user reads the story "Ravi looks at his friend Niko and realizes he is sad..." and interacts with the child to deepen their understanding of emotions.

[1138] (Application example 2)

[1139] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1140] Previous story generation systems were limited in their ability to provide appropriate stories and images to support emotional education and interpersonal relationship development. They also struggled to recognize users' emotions in real time and provide personalized stories and images based on those emotions. Furthermore, the means to interactively experience the generated stories were limited, limiting the enrichment of children's learning experiences.

[1141] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a generation means for generating a story for each emotion, an image generation means for generating images based on the generated story, a display means for displaying the generated story and images, an emotion recognition means for automatically recognizing emotions, an audio playback means for reading aloud the generated story, and a cross-platform support means for supporting multiple display devices. This makes it possible to recognize the user's emotions in real time and provide personalized stories and images based on them, thereby providing a more interactive story-telling experience.

[1142] The "generation means for generating a story for each emotion" is a means for automatically generating a story corresponding to a specific emotion based on a user's selection and emotion recognition data.

[1143] The "image generating means for generating images based on the generated story" is a means for automatically generating images corresponding to the content of the generated story.

[1144] A "display means for displaying the generated story and images" is a device or interface for visually presenting the generated story and images to a user.

[1145] "Emotion recognition means for automatically recognizing emotions" is a means for automatically analyzing and recognizing emotions from the user's facial expressions and voice using sensors such as a camera and a microphone.

[1146] The "audio playback means for reading aloud" is a means for playing back the generated story aloud and providing it to the user audibly.

[1147] "Cross-platform support for multiple display devices" refers to the means by which the generated story and images are adjusted to be displayed appropriately on different types of devices (smartphones, tablets, smart glasses, head-mounted displays, etc.).

[1148] The "selection means for selecting an emotional area based on user input" is a means that allows a user to specify a particular emotional area based on their current feelings and intentions.

[1149] The "display means without backlight" refers to a means for displaying without using a backlight in order to reduce power consumption.

[1150] "Display means using electronic paper" refers to a display device that uses electronic ink technology and combines the ease of viewing as paper with power-saving performance.

[1151] This invention is a system that automatically generates stories and images for children using real-time emotion recognition and generation AI for the purpose of emotional education. The system automatically recognizes the user's emotions, generates a story based on those emotions, and then generates corresponding images. The generated stories and images are displayed on multiple display devices, providing an interactive reading experience through a text-to-speech function.

[1152] Hardware and software used

[1153] 1. Hardware

[1154] Smartphone

[1155] tablet

[1156] Smart glasses (e.g. Google Glass)

[1157] Head-mounted displays (e.g., Oculus Quest)

[1158] Electronic Paper Device

[1159] 2. Software

[1160] Azure Cognitive Services (for emotion recognition)

[1161] AWS Lambda (serverless computing)

[1162] S3 bucket (data storage)

[1163] OpenAI GPT-4 (a generative AI model for story generation)

[1164] DALL-E 2 (generative AI model for image generation)

[1165] React Native (for cross-platform compatibility)

[1166] Amazon Polly (for voice reading)

[1167] System Overview

[1168] 1. User Emotion Recognition

[1169] When a user launches an application, Azure Cognitive Services uses the device's camera and microphone to recognize the user's emotions. For example, if facial expression recognition technology determines that the user is "calm," this information is sent to the server.

[1170] 2. Narrative Generation Based on Emotional Data

[1171] Once emotion recognition data is sent to the server, OpenAI GPT-4 is invoked using AWS Lambda. Using emotion data (e.g., "empathy") as input, the generative AI model generates a related story. As a specific example, it generates an empathetic story using a rabbit character named "Rabi."

[1172] 3. Image Generation

[1173] Based on the content of the generated story, images that fit the story are generated using DALL-E 2. For example, a scene in the story where Ravi and Nico are sitting together is depicted.

[1174] 4. Data transmission and display

[1175] The generated stories and images are compiled into documents and sent using React Native to smartphones, tablets, smart glasses, head-mounted displays, and potentially even backlit e-paper devices.

[1176] 5. Reading aloud

[1177] The story is generated using Amazon Polly and read aloud, and an interface is provided for users to enjoy the story interactively.

[1178] Examples of concrete examples and prompts

[1179] For example, a user launches an app and Azure Cognitive Services determines that the user feels "calm." Emotional data "empathy" is sent to the server, OpenAI GPT-4 generates an "empathetic story about Ravi and Nico," and DALL-E 2 generates an image of "a scene where Ravi and Nico are sitting together." The generated story and image are sent to the device and read aloud.

[1180] An example prompt is:

[1181] "Generate stories with empathy as a theme based on emotional data"

[1182] "Based on the story you've created, generate an image of the rabbit character and his friends together."

[1183] The present invention generates stories and images that correspond to the user's real-time emotions, making it possible to support children's emotional and interpersonal development through an interactive storytelling experience.

[1184] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1185] Step 1:

[1186] The device launches an application and activates the device's camera and microphone to automatically recognize the user's emotions. Azure Cognitive Services is used to analyze data obtained from the camera and microphone and recognize the user's emotions. For example, a facial expression recognition algorithm determines that the user is "calm." The input is the user's facial expression and voice data, and the output is the recognized emotion data.

[1187] Step 2:

[1188] The device sends the recognized emotion data to the server. Specifically, the emotion data called "empathy" is sent in JSON format. The input is the emotion data output from Step 1, and the output is a notification of completion of data transmission to the server.

[1189] Step 3:

[1190] Based on the emotion data received by the server, AWS Lambda is triggered to generate a story. The Lambda function calls OpenAI GPT-4 and uses the emotion data as input to create a story generation prompt. The input is emotion data and the story generation prompt, and the output is the generated story text. Specifically, the prompt text "Generate a story with the theme of empathy" with the emotion "empathy" is sent to GPT-4, which then generates a story.

[1191] Step 4:

[1192] The server uses DALL-E 2 to generate images based on the generated story text. Specific scenes from the story are sent as prompts to DALL-E 2, which then generates corresponding images. The input is the story text and the image generation prompt, and the output is the generated image file. For example, an image is generated based on the prompt "A scene where Ravi and Nico are sitting together."

[1193] Step 5:

[1194] The server compiles the generated story text and images into a single document and sends it to the terminal. The input is the story text and image files, and the output is the compiled document file and a notification of completion of transmission. The data is compiled into a JSON format document file and sent to the terminal.

[1195] Step 6:

[1196] The device analyzes the received document file and provides it to the user through a display method. For example, using React Native, it can be displayed on a smartphone, tablet, smart glasses, or head-mounted display. The input is the document file, and the output is the story and images displayed on the device.

[1197] Step 7:

[1198] The device uses Amazon Polly to read aloud the story text generated, allowing users to enjoy the story not only visually but also aurally. The input is the story text, and the output is audio data, specifically the content of the story being read aloud.

[1199] Each step is processed sequentially to generate a personalized story and images based on the user's emotions, ultimately providing an interactive storytelling experience.

[1200] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1201] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1202] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1203] [Fourth embodiment]

[1204] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1205] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1206] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1207] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1208] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1209] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1210] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1211] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1212] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1213] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1214] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1215] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1216] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1217] This invention is a system that automatically generates picture books for children and supports the development of their emotions and interpersonal relationships through reading aloud, with the aim of emotional education. The system provides an interface that is easy for users to operate, and generates and displays stories and images based on emotional domains.

[1218] System Overview

[1219] 1. User operations

[1220] The user launches the application and selects an emotional area.

[1221] Example: A user selects "empathy" within an application.

[1222] The terminal transmits the user's selection data to the server.

[1223] 2. Story generation based on emotional domains

[1224] The server analyzes the emotion area data received from the device.

[1225] Example: The server checks that the emotion domain is "empathy."

[1226] Emotional domains are input into the generative AI model on the server to generate a story.

[1227] Example: An AI model generates a story about a little rabbit character named Ravi and his friend Nico, with a theme of empathy.

[1228] Some of the stories generated:

[1229] One day, Ravi, a little rabbit, finds his friend Niko sad, and he wants to know why Niko is sad.

[1230] 3. Narrative-based image generation

[1231] The server generates image prompts based on the generated story.

[1232] Example: Generate a prompt for "A scene with Ravi and Nico together."

[1233] The server passes the prompt to an image generation AI model, which generates an appropriate image.

[1234] Example: An image generation AI generates an image depicting a scene in which "Ravi and Nico are sitting together."

[1235] 4. Data transmission and display

[1236] The server compiles the generated story and images and sends them to the device.

[1237] Example: Sending story text and image files to a device.

[1238] The terminal analyzes the received data and displays it on the electronic paper device.

[1239] Example: An e-paper device displays "a story and corresponding images that show the relatable nature of Ravi and Nico."

[1240] 5. Reading aloud

[1241] A user reads a story to a child using an e-paper device.

[1242] Example: A parent reads part of a story: "Ravi looks at his friend Niko and realizes he is sad..."

[1243] The user interacts with the child to discuss the content of the story and develop the child's emotions.

[1244] Example: "How did Ravi know Nico was sad?" helps children develop an understanding of emotions.

[1245] This system allows parents to easily create picture books suitable for emotional education and read them to their children. By using an e-paper device, it provides a visual experience that is gentle on children's eyes while supporting the development of emotions and interpersonal relationships.

[1246] The processing flow will be explained below.

[1247] Step 1:

[1248] The user launches the application and selects an emotional area.

[1249] Example: A user selects "empathy" within an application.

[1250] Step 2:

[1251] The terminal transmits the user's selection data to the server.

[1252] Example: Send data from the emotion field "empathy" to the server.

[1253] Step 3:

[1254] The server analyzes the emotion area data received from the device.

[1255] Example: The server checks that the emotion domain is "empathy."

[1256] Step 4:

[1257] The emotional domain of "empathy" is input into the generative AI model on the server to generate a story.

[1258] Example: An AI model generates stories with empathy as the theme, using a rabbit character named "Rabi."

[1259] Some of the stories generated:

[1260] One day, Ravi, a little rabbit, finds his friend Niko sad, and he wants to know why Niko is sad.

[1261] Step 5:

[1262] The server generates image prompts based on the generated story.

[1263] Example: Generate a prompt for "A scene with Ravi and Nico together."

[1264] Step 6:

[1265] The server passes the prompt to an image generation AI model, which generates an appropriate image.

[1266] Example: An image generation AI generates an image of a scene where "Ravi and Nico are sitting together."

[1267] Step 7:

[1268] The server compiles the generated story and images into a single document.

[1269] Example: Combining generated narrative text and images.

[1270] Examples of compiled data:

[1271] Story text:

[1272] One day, a little rabbit named Ravi found his friend Niko sad.

[1273] Story Image:

[1274] Image file of Ravi and Nico sitting together

[1275] Step 8:

[1276] The server sends the compiled data to the terminal.

[1277] Example: Sending a story and image file to a device.

[1278] Step 9:

[1279] The terminal analyzes the received data and displays it on the electronic paper device.

[1280] Example: An e-paper device displays "a story and corresponding images that show the relatable nature of Ravi and Nico."

[1281] Step 10:

[1282] A user reads a story to a child using an e-paper device.

[1283] Example: A parent reads part of a story: "Ravi looks at his friend Niko and realizes he is sad..."

[1284] Step 11:

[1285] The user interacts with the child to discuss the content of the story and develop the child's emotions.

[1286] Example: "How did Ravi know Nico was sad?" helps children develop an understanding of emotions.

[1287] Example 1

[1288] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1289] Previous emotional education systems faced challenges, such as a lack of an easy-to-use interface for users and the lack of the ability to automatically generate individual stories and images based on emotional domains. Furthermore, selecting a device that provides a visual experience that is easy on children's eyes is also important as a means of displaying the generated content, but this was often not properly implemented. As a result, there were problems with the creation of teaching materials suitable for emotional education and the learning process using them not being effectively supported.

[1290] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1291] In this invention, the server includes a selection means for a user to select an emotional area, a story generation means including a generative AI model for generating a story based on the selected emotional area, an image generation means including an image generation AI model for generating image prompts based on the generated story and generating images based on the prompts, a transmission means for transmitting the generated story and images to the user's terminal, and a display means for displaying the transmitted story and images. This makes it possible to automatically generate individual stories and images based on emotions using an interface that is easy for users to operate, and to display the generated content on a device that is easy on children's eyes.

[1292] A "selection means" is a device or software that provides an interface for a user to select a particular emotional domain.

[1293] A "generative AI model" is an artificial intelligence model that automatically generates stories and sentences based on selected emotional domains.

[1294] The "story generation means" is a means for generating a story using a generative AI model with emotional domain data selected via a selection means as input.

[1295] A "prompt" is a textual instruction given to an image-generating AI model to generate a specific image.

[1296] An "image generation AI model" is an artificial intelligence model that automatically generates corresponding images based on given prompts.

[1297] The "image generation means" is a means for generating an image generation prompt based on the story generated by the story generation means, and passing the prompt to an image generation AI model to generate an image.

[1298] The "transmission means" is a means for transmitting the generated story and images from the server to the user's terminal.

[1299] "Display means" refers to a device that displays the received story and images on the user's terminal, allowing the user to check the content.

[1300] "Electronic paper" is a visually friendly display technology that can display text and images without the use of a backlight.

[1301] This system automatically generates picture books for children with the aim of emotional education, and supports the development of children's emotions and interpersonal relationships through reading aloud. This system consists of a terminal operated by the user, a server that processes data, and a display means that displays the generated stories and images.

[1302] First, the user launches the emotion education system application on a smartphone or tablet device (e.g., iOS or Android). The user selects an emotion area through the application interface. For example, the user selects "empathy."

[1303] Next, the device sends the data of the user's selected emotional domain to the server. The server uses a cloud server (e.g., AWS, GCP) to analyze the received data and input the emotional domain data into a generative AI model (e.g., OpenAI GPT-4). This generates a story about the selected emotional domain. As a specific example, a story based on the theme of "empathy" might be generated, such as "The Story of Little Rabbit Ravi and His Friend Nico."

[1304] The generated story is then used to generate image generation prompts. The server creates image generation prompts based on the content of the story and passes them to an image generation AI model (e.g., DALL-E, MidJourney). This automatically generates images corresponding to scenes in the story. For example, an image depicting "Ravi and Nico sitting together" is generated.

[1305] The generated story and images are sent from the server to the device, which then analyzes the data and displays it on an electronic paper device (e.g., Kindle, e-ink device). This display device does not use a backlight to provide a visual experience that is gentle on children's eyes.

[1306] Finally, the user reads a story to the child using the e-paper device. For example, a parent might read the story, "Ravi looks at his friend Nico and realizes he's sad..." In addition, the user can deepen the child's understanding of emotions by interacting with the child and discussing the content of the story. For example, emotional education is provided by asking, "How did Ravi realize Nico was sad?"

[1307] As a result, this system provides an easy-to-use interface for users, automatically generates individual stories and images based on emotional domains, and displays them on an electronic paper device that is easy on children's eyes. By utilizing a generative AI model and an image-generating AI model, the present invention is characterized by its ability to quickly provide high-quality content suitable for emotional education.

[1308] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1309] Step 1:

[1310] The user launches the application and selects an emotion area.

[1311] Input: The user launches the application on their smartphone or tablet and selects an emotion domain (e.g., "empathy").

[1312] How it works: The application displays an interface for the user to select an emotional area. When the user selects an emotional area, the information is temporarily stored on the device.

[1313] Output: Data from the selected emotion domains is generated and passed to the next processing step.

[1314] Step 2:

[1315] The device sends the emotion area data to the server.

[1316] Input: Emotion domain data selected by the user in Step 1 (e.g., "Empathy")

[1317] How it works: The device sends emotional domain data to the server using an API request over the internet. The data is securely sent to the server, where analysis begins.

[1318] Output: Emotional domain data is sent to the server.

[1319] Step 3:

[1320] The server analyzes the received emotional data and generates a story.

[1321] Input: Emotion domain data sent to the server in step 2 (e.g., "Empathy")

[1322] How it works: The server analyzes the received emotional domain data and inputs it into a generative AI model (e.g., OpenAI GPT-4). The generative AI model generates a story based on the emotional domain. For example, it generates "The Story of Little Rabbit Ravi and His Friend Nico."

[1323] Output: The generated story text (e.g., "One day, Ravi the little rabbit found his friend Niko sad...")

[1324] Step 4:

[1325] The server generates image prompts based on the story and generates images.

[1326] Input: The narrative text generated in step 3

[1327] How it works: The server generates an image prompt based on the content of the generated story (e.g., "A scene where Ravi and Nico are together"). It then passes the image prompt to an image generation AI model (e.g., DALL-E, MidJourney). The image generation AI model generates an image based on the prompt. For example, it generates an image of "A scene where Ravi and Nico are sitting together."

[1328] Output: Generated image file (e.g. JPEG, PNG format)

[1329] Step 5:

[1330] The server sends the generated story and images to the device.

[1331] Input: The story text generated in step 3 and the image files generated in step 4

[1332] How it works: The server compiles the generated story text and image files and sends them to the user's device via an API request over the internet.

[1333] Output: A data package containing the story and images

[1334] Step 6:

[1335] The terminal analyzes the received data and displays it on the e-paper device.

[1336] Input: The data package sent by the server in step 5 (story text and image files)

[1337] How it works: The device analyzes the received data and displays the story text and images on an electronic paper device (e.g., Kindle, e-ink device). Electronic paper devices display without a backlight, providing a visual experience that is gentle on children's eyes.

[1338] Output: The displayed story and images

[1339] Step 7:

[1340] A user reads a story to a child on an e-paper device

[1341] Input: The story and image displayed in step 6

[1342] How it works: A user reads a story to a child using an e-paper device. For example, a parent reads, "Ravi looks at his friend Niko and realizes he's sad..." The user then interacts with the child, asking questions about the story to deepen the child's understanding of their emotions.

[1343] Output: Responses of children who received emotional education

[1344] The above are the specific processing steps of the program of this system.

[1345] (Application example 1)

[1346] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1347] The present invention relates to a system that automatically generates picture books to support children's emotional education. With conventional systems, it is difficult to find the time and timing for reading aloud, making it difficult to provide effective emotional education, especially when parents are busy or children have to wait long periods of time. Furthermore, there is a lack of a way to effectively utilize spare time, such as while waiting for delivery. Therefore, an objective of the present invention is to provide a system that can provide emotional education by effectively utilizing waiting time for delivery.

[1348] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1349] In this invention, the server includes a generation means for generating a story for each emotion, an image generation means for generating an image based on the generated story, a display means for displaying the generated story and image, and a promotion means for generating a story and image for each emotion, displaying them when a user places an order, and promoting reading aloud while waiting for delivery. This makes it possible to effectively utilize the waiting time for delivery to provide emotional education.

[1350] "Emotional stories" are stories generated based on specific emotional domains, featuring situations and characters related to those emotions.

[1351] "Generation means" refers to a device or program that has the function of automatically generating stories and images based on the emotional domain selected by the user.

[1352] "Image generation means" refers to a device or program that has the function of automatically generating images of scenes and characters related to the generated story.

[1353] "Display means" refers to a device or function that displays the generated story or images so that the user can view them.

[1354] "Promotion means" refers to a device or function that displays the story and images generated when the user places an order and encourages reading aloud while waiting for delivery.

[1355] The "selection means" refers to a device or program that provides an interface or function for the user to select a desired emotional area.

[1356] "Transmission means" refers to a device or program for transmitting the generated story and images to the user's terminal.

[1357] "Delivery Order" means an online or offline order placed by a User for delivery of food or merchandise.

[1358] This invention is a system that automatically generates picture books for the purpose of emotional education and supports the development of children's emotions and interpersonal relationships through reading aloud. This system promotes emotional education by allowing parents to read picture books to their children, making effective use of waiting times for delivery services.

[1359] 1. System Configuration

[1360] The system includes the following major components:

[1361] Server: Includes a generative AI model and an image generation model that generates stories and images based on emotional domains.

[1362] Device: A device on which parents place delivery orders and view the generated stories and images. Specifically, this refers to a smartphone or tablet.

[1363] User: In this case, the parent who places the delivery order and reads to their child.

[1364] 2. Explanation of program processing

[1365] 1. User operations

[1366] A user launches a food delivery app, orders a meal, and selects the "Emotional Education" option, where the user selects a specific emotional domain (e.g., empathy).

[1367] 2. Story Generation Based on Emotional Domains

[1368] The user's selection data is sent to a server, which uses a generative AI model (e.g., OpenAI GPT-4) to generate a narrative based on the selected emotional domains.

[1369] Example: Generate a story about a little rabbit character named "Ravi" and his friend "Nico" that has an empathy theme. The prompt is:

[1370] Input prompt: "Generate a story about Ravi and Nico based on empathy."

[1371] 3. Narrative-based image generation

[1372] Based on the generated story, the server generates appropriate images using an image generation model (e.g., DALL-E 2).

[1373] Example: An image depicting "Ravi and Nico together" is generated.

[1374] 4. Data transmission and display

[1375] The server sends the generated story and images to the user's terminal, which displays the received story and images.

[1376] Parents can read stories to their children while they wait for their delivery.

[1377] 5. Reading aloud and emotional education

[1378] Emotional education is carried out by parents reading stories and discussing emotions with their children through dialogue.

[1379] For example, asking, "How did Ravi know Nico was sad?" provides an opportunity for children to think about empathy.

[1380] 3. Hardware and Software Use

[1381] The server uses a generative AI model (OpenAI GPT-4) for story generation and a generative image model (DALL-E 2) for image generation.

[1382] The user's device runs a delivery application (e.g., a generic food delivery service app) and has a display for displaying stories and images.

[1383] The server receives the user's selection data and inputs the data into the AI ​​model for processing.

[1384] This makes it possible to provide a system that makes effective use of the waiting time for delivery and allows parents and children to enjoy reading picture books suitable for emotional education.

[1385] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1386] Step 1:

[1387] A user launches a food delivery app and selects the "Emotional Education" option when ordering food. The user also selects an emotional domain, such as "Empathy," through an interface that identifies emotional domains. This generates food delivery order data and emotional domain data.

[1388] Step 2:

[1389] The terminal transmits the emotional domain data selected by the user to the server. The server receives the user's selection data and analyzes the emotional domains. This data includes the emotional domain selected by the user, such as "empathy."

[1390] Step 3:

[1391] The server inputs the emotional domain data into a generative AI model (e.g., OpenAI GPT-4) for generating stories based on emotional domains. The server uses the AI ​​model to generate stories based on specific emotional domains. The input is a prompt statement: "Generate a story about Ravi and Nico based on empathy." The output is a story text with an empathy theme.

[1392] Step 4:

[1393] The server analyzes the generated story text and generates image prompts based on the story content. Image prompts are text data containing specific scenes in the story (e.g., "A scene where Ravi and Nico are together"). These prompts are input into an image generation AI model (e.g., DALL-E 2). The result is an image depicting a scene from the story.

[1394] Step 5:

[1395] The server compiles the generated story and image data and sends them in digital format to the terminal, which then retrieves the received story data and image data and displays them to the user in an appropriate format, allowing the user to view the story and corresponding images while waiting for delivery.

[1396] Step 6:

[1397] The user reads a story displayed on the device to the child. The user then interacts with the child about the story's contents and provides emotional education. Specifically, by discussing the actions and emotions of the characters in the story, the child can deepen their understanding of emotions.

[1398] Through these steps, parents and children can effectively utilize the time they spend waiting for delivery by reading picture books suitable for emotional education.

[1399] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1400] This invention is a system that automatically generates picture books for children with the aim of emotional education, and supports the development of children's emotions and interpersonal relationships through reading aloud. In addition to selecting emotional domains based on user input, this system incorporates an emotion engine that recognizes the user's emotions to generate more personalized picture books.

[1401] System Overview

[1402] 1. User operations

[1403] The user launches the application and manually selects an emotion area, or the device automatically recognizes the user's emotion using an emotion engine.

[1404] Example: The user manually selects "empathy," or the emotion engine recognizes the user's facial expression and voice as "calm" and selects "empathy."

[1405] 2. Story generation based on emotional domains

[1406] The server analyzes the emotion area data received from the device.

[1407] Example: The server checks that the emotion domain is "empathy."

[1408] The emotional domain of "empathy" is input into the generative AI model on the server to generate a story.

[1409] Example: An AI model generates stories with empathy as the theme, using a rabbit character named "Rabi."

[1410] Some of the stories generated:

[1411] One day, Ravi, a little rabbit, finds his friend Niko sad, and he wants to know why Niko is sad.

[1412] 3. Narrative-based image generation

[1413] The server generates image prompts based on the generated story.

[1414] Example: Generate a prompt for "A scene with Ravi and Nico together."

[1415] The server passes the prompt to an image generation AI model, which generates an appropriate image.

[1416] Example: An image generation AI generates an image of a scene where "Ravi and Nico are sitting together."

[1417] 4. Data transmission and display

[1418] The server compiles the generated story and images into a single document and sends it to the device.

[1419] Example: Sending story text and image files to the device.

[1420] The terminal analyzes the received data and displays it on the electronic paper device.

[1421] Example: An e-paper device displays "a story and corresponding images that show the relatable nature of Ravi and Nico."

[1422] 5. Reading aloud

[1423] A user reads a story to a child using an e-paper device.

[1424] Example: A parent reads part of a story: "Ravi looks at his friend Niko and realizes he is sad..."

[1425] The user interacts with the child to discuss the content of the story and develop the child's emotions.

[1426] Example: "How did Ravi know Nico was sad?" helps children develop an understanding of emotions.

[1427] This system allows parents to easily create picture books suitable for emotional education and read them to their children. Using an emotion engine, it provides a more personalized experience according to the user's mood and state. Using an e-paper device, it provides a visual experience that is gentle on children's eyes while supporting the development of emotions and interpersonal relationships.

[1428] The processing flow will be explained below.

[1429] Step 1:

[1430] When a user launches the application, the emotion engine recognizes the user's emotion, and the device displays an emotion area based on the user's input.

[1431] Example: A user launches an application, and the emotion engine selects "empathy" based on the user's facial expressions and voice.

[1432] Step 2:

[1433] The terminal transmits the user's selection data to the server.

[1434] Example: Send data from the emotion field "empathy" to the server.

[1435] Step 3:

[1436] The server analyzes the emotion area data received from the device.

[1437] Example: The server checks that the emotion domain is "empathy."

[1438] Step 4:

[1439] The emotional domain of "empathy" is input into the generative AI model on the server to generate a story.

[1440] Example: An AI model generates stories with empathy as the theme, using a rabbit character named "Rabi."

[1441] Some of the stories generated:

[1442] One day, Ravi, a little rabbit, finds his friend Niko sad, and he wants to know why Niko is sad.

[1443] Step 5:

[1444] The server generates image prompts based on the generated story.

[1445] Example: Generate a prompt for "A scene with Ravi and Nico together."

[1446] Step 6:

[1447] The server passes the prompt to an image generation AI model, which generates an appropriate image.

[1448] Example: An image generation AI generates an image of a scene where "Ravi and Nico are sitting together."

[1449] Step 7:

[1450] The server compiles the generated story and images into a single document.

[1451] Example: Combining generated narrative text and images.

[1452] Examples of compiled data:

[1453] Story text:

[1454] One day, a little rabbit named Ravi found his friend Niko sad.

[1455] Story Image:

[1456] Image file of Ravi and Nico sitting together

[1457] Step 8:

[1458] The server sends the compiled data to the terminal.

[1459] Example: Sending a story and image file to a device.

[1460] Step 9:

[1461] The terminal analyzes the received data and displays it on the electronic paper device.

[1462] Example: An e-paper device displays "a story and corresponding images that show the relatable nature of Ravi and Nico."

[1463] Step 10:

[1464] A user reads a story to a child using an e-paper device.

[1465] Example: A parent reads part of a story: "Ravi looks at his friend Niko and realizes he is sad..."

[1466] Step 11:

[1467] The user interacts with the child to discuss the content of the story and develop the child's emotions.

[1468] Example: "How did Ravi know Nico was sad?" helps children develop an understanding of emotions.

[1469] Example 2

[1470] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1471] A major challenge is the lack of teaching materials and tools for effectively educating children about emotions. There are also limited ways for parents to easily create picture books suitable for emotional education and use them to help their children understand emotions. Furthermore, as digital devices are used, there is a need to support the development of emotions and interpersonal relationships while reducing visual burden.

[1472] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a recognition means for recognizing a user's emotion, a story generation means for generating a story based on the recognized emotion, an image generation means for generating images based on the generated story, and a display means for displaying the generated story and images. This makes it possible to automatically generate a personalized picture book according to the user's emotion and display it using an electronic paper device, thereby enabling effective emotional education for children while reducing visual burden.

[1473] The "recognition means" is a means for recognizing the user's emotions.

[1474] A "narrative generation means" is a means for generating a narrative based on recognized emotions.

[1475] The "image generation means" is a means for generating an image based on the generated story.

[1476] "Display means" is a means for displaying the generated story and images.

[1477] The "selection means" is a means for selecting an emotion area based on a user's input.

[1478] "Electronic paper" is a backlit display device used to reduce visual strain.

[1479] The "transmission means" is a means for transmitting the generated story to the terminal.

[1480] The "reading support tool" is a tool for reading the generated story to a child.

[1481] This invention is a system that automatically generates picture books for children with the aim of emotional education, and supports the development of children's emotions and interpersonal relationships through reading aloud. In addition to selecting emotional domains based on user input, this system incorporates an emotion engine to recognize the user's emotions and generate more personalized picture books.

[1482] First, the user launches the application and manually selects an emotion area, or the device automatically recognizes the user's emotion using an emotion engine. For example, the user manually selects "empathy," or the emotion engine recognizes the user's facial expression and voice as "calm" and selects "empathy."

[1483] Next, the device sends the selected emotional domain to the server. The server analyzes this data and confirms that the emotional domain is "empathy." The "empathy" emotional domain is input into the generative AI model, which then generates a story. As a concrete example, the generative AI model uses a rabbit character named "Rabi" to generate a story with an empathy theme.

[1484] Some examples of generated stories:

[1485] One day, Ravi, a little rabbit, finds his friend Niko sad, and he wants to know why Niko is sad.

[1486] The server then generates image prompts based on the generated story. For example, it generates a prompt for "a scene where Ravi and Nico are together." The server then passes the prompt to an image generation AI model, which generates an appropriate image. For example, the image generation AI generates an image of a scene where "Ravi and Nico are sitting together."

[1487] Finally, the server compiles the story and the generated images into a single document and sends it to the device. The device then analyzes the received data and displays it on the e-paper device. For example, the e-paper device displays "a story that shows the empathy between Ravi and Nico and the corresponding images."

[1488] The user uses the e-paper device to read a story to the child. For example, a parent might read part of the story, "Ravi looks at his friend Nico and realizes he's sad..." The user interacts with the child, discussing the content of the story and developing the child's emotions. For example, the parent might ask, "How did Ravi realize Nico was sad?", helping the child deepen their understanding of emotions.

[1489] This system allows parents to easily create picture books suitable for emotional education and read them to their children. Using an emotion engine, it provides a more personalized experience according to the user's mood and state. Using an e-paper device, it provides a visual experience that is gentle on children's eyes while supporting the development of emotions and interpersonal relationships.

[1490] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1491] Step 1:

[1492] The user launches the application and selects an emotion area, or the terminal uses an emotion engine to recognize the user's emotion.

[1493] Input: Emotion area selection screen display, user's facial expression data and voice data.

[1494] Data processing: Users can select emotion areas from a drop-down menu, or the emotion engine analyzes facial expressions and voice.

[1495] Output: Selected emotion domain data.

[1496] What happens: The user operates their smartphone to launch the app and taps the "Empathy" button. The device uses the camera to scan the user's facial expressions and performs voice recognition.

[1497] Step 2:

[1498] The terminal transmits the selected emotion area to the server.

[1499] Input: Emotion domain data.

[1500] Data processing: Packaging and sending emotional domain data.

[1501] Output: Emotional domain data sent to the server.

[1502] Specific operation: The device sends data on the emotional domain "empathy" to the server via the emotion engine API.

[1503] Step 3:

[1504] The server analyzes the received emotional domain data and generates a story for the picture book using a story generation means.

[1505] Input: Emotion domain data.

[1506] Data processing: Analysis of emotional domain data, input into generative AI models, and narrative generation.

[1507] Output: The generated narrative text.

[1508] Specific operation: The server inputs the emotional domain data "empathy" into the generative AI model and generates a story themed around empathy using a character named "Rabi."

[1509] Step 4:

[1510] The server generates image prompts based on the generated story and generates images using an image generation means.

[1511] Input: Narrative text.

[1512] Data processing: Generate image prompts from narrative text, input them into an image generation AI model, and generate images.

[1513] Output: The generated image data.

[1514] Specific operation: The server creates a prompt describing "a scene where Ravi and Nico are together" and inputs it into an image generation AI model to generate an image.

[1515] Step 5:

[1516] The server compiles the generated story text and images into a single document and sends it to the terminal.

[1517] Input: Narrative text and image data.

[1518] Data processing: Organizing data into document formats such as PDF and sending it to the device.

[1519] Output: The document data sent to the device.

[1520] Specific operation: The server compiles the story text and images into PDF format and sends it to the terminal.

[1521] Step 6:

[1522] The terminal analyzes the received data and displays it on the electronic paper device.

[1523] Input: The received document data.

[1524] Data processing: Analyzing document data and preparing it for display on e-paper devices.

[1525] Output: Story and images displayed on an e-paper device.

[1526] Specific operation: The terminal displays the received PDF on the e-paper device.

[1527] Step 7:

[1528] A user reads a story to a child using an e-paper device.

[1529] Input: A story and an image displayed on an e-paper device.

[1530] Data processing: Reading aloud was conducted.

[1531] Output: Providing emotional education to children.

[1532] Specific Action: The user reads the story "Ravi looks at his friend Niko and realizes he is sad..." and interacts with the child to deepen their understanding of emotions.

[1533] (Application example 2)

[1534] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1535] Previous story generation systems were limited in their ability to provide appropriate stories and images to support emotional education and interpersonal relationship development. They also struggled to recognize users' emotions in real time and provide personalized stories and images based on those emotions. Furthermore, the means to interactively experience the generated stories were limited, limiting the enrichment of children's learning experiences.

[1536] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a generation means for generating a story for each emotion, an image generation means for generating images based on the generated story, a display means for displaying the generated story and images, an emotion recognition means for automatically recognizing emotions, an audio playback means for reading aloud the generated story, and a cross-platform support means for supporting multiple display devices. This makes it possible to recognize the user's emotions in real time and provide personalized stories and images based on them, thereby providing a more interactive story-telling experience.

[1537] The "generation means for generating a story for each emotion" is a means for automatically generating a story corresponding to a specific emotion based on a user's selection and emotion recognition data.

[1538] The "image generating means for generating images based on the generated story" is a means for automatically generating images corresponding to the content of the generated story.

[1539] A "display means for displaying the generated story and images" is a device or interface for visually presenting the generated story and images to a user.

[1540] "Emotion recognition means for automatically recognizing emotions" is a means for automatically analyzing and recognizing emotions from the user's facial expressions and voice using sensors such as a camera and a microphone.

[1541] The "audio playback means for reading aloud" is a means for playing back the generated story aloud and providing it to the user audibly.

[1542] "Cross-platform support for multiple display devices" refers to the means by which the generated story and images are adjusted to be displayed appropriately on different types of devices (smartphones, tablets, smart glasses, head-mounted displays, etc.).

[1543] The "selection means for selecting an emotional area based on user input" is a means that allows a user to specify a particular emotional area based on their current feelings and intentions.

[1544] The "display means without backlight" refers to a means for displaying without using a backlight in order to reduce power consumption.

[1545] "Display means using electronic paper" refers to a display device that uses electronic ink technology and combines the ease of viewing as paper with power-saving performance.

[1546] This invention is a system that automatically generates stories and images for children using real-time emotion recognition and generation AI for the purpose of emotional education. The system automatically recognizes the user's emotions, generates a story based on those emotions, and then generates corresponding images. The generated stories and images are displayed on multiple display devices, providing an interactive reading experience through a text-to-speech function.

[1547] Hardware and software used

[1548] 1. Hardware

[1549] Smartphone

[1550] tablet

[1551] Smart glasses (e.g. Google Glass)

[1552] Head-mounted displays (e.g., Oculus Quest)

[1553] Electronic Paper Device

[1554] 2. Software

[1555] Azure Cognitive Services (for emotion recognition)

[1556] AWS Lambda (serverless computing)

[1557] S3 bucket (data storage)

[1558] OpenAI GPT-4 (a generative AI model for story generation)

[1559] DALL-E 2 (generative AI model for image generation)

[1560] React Native (for cross-platform compatibility)

[1561] Amazon Polly (for voice reading)

[1562] System Overview

[1563] 1. User Emotion Recognition

[1564] When a user launches an application, Azure Cognitive Services uses the device's camera and microphone to recognize the user's emotions. For example, if facial expression recognition technology determines that the user is "calm," this information is sent to the server.

[1565] 2. Narrative Generation Based on Emotional Data

[1566] Once emotion recognition data is sent to the server, OpenAI GPT-4 is invoked using AWS Lambda. Using emotion data (e.g., "empathy") as input, the generative AI model generates a related story. As a specific example, it generates an empathetic story using a rabbit character named "Rabi."

[1567] 3. Image Generation

[1568] Based on the content of the generated story, images that fit the story are generated using DALL-E 2. For example, a scene in the story where Ravi and Nico are sitting together is depicted.

[1569] 4. Data transmission and display

[1570] The generated stories and images are compiled into documents and sent using React Native to smartphones, tablets, smart glasses, head-mounted displays, and potentially even backlit e-paper devices.

[1571] 5. Reading aloud

[1572] The story is generated using Amazon Polly and read aloud, and an interface is provided for users to enjoy the story interactively.

[1573] Examples of concrete examples and prompts

[1574] For example, a user launches an app and Azure Cognitive Services determines that the user feels "calm." Emotional data "empathy" is sent to the server, OpenAI GPT-4 generates an "empathetic story about Ravi and Nico," and DALL-E 2 generates an image of "a scene where Ravi and Nico are sitting together." The generated story and image are sent to the device and read aloud.

[1575] An example prompt is:

[1576] "Generate stories with empathy as a theme based on emotional data"

[1577] "Based on the story you've created, generate an image of the rabbit character and his friends together."

[1578] The present invention generates stories and images that correspond to the user's real-time emotions, making it possible to support children's emotional and interpersonal development through an interactive storytelling experience.

[1579] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1580] Step 1:

[1581] The device launches an application and activates the device's camera and microphone to automatically recognize the user's emotions. Azure Cognitive Services is used to analyze data obtained from the camera and microphone and recognize the user's emotions. For example, a facial expression recognition algorithm determines that the user is "calm." The input is the user's facial expression and voice data, and the output is the recognized emotion data.

[1582] Step 2:

[1583] The device sends the recognized emotion data to the server. Specifically, the emotion data called "empathy" is sent in JSON format. The input is the emotion data output from Step 1, and the output is a notification of completion of data transmission to the server.

[1584] Step 3:

[1585] Based on the emotion data received by the server, AWS Lambda is triggered to generate a story. The Lambda function calls OpenAI GPT-4 and uses the emotion data as input to create a story generation prompt. The input is emotion data and the story generation prompt, and the output is the generated story text. Specifically, the prompt text "Generate a story with the theme of empathy" with the emotion "empathy" is sent to GPT-4, which then generates a story.

[1586] Step 4:

[1587] The server uses DALL-E 2 to generate images based on the generated story text. Specific scenes from the story are sent as prompts to DALL-E 2, which then generates corresponding images. The input is the story text and the image generation prompt, and the output is the generated image file. For example, an image is generated based on the prompt "A scene where Ravi and Nico are sitting together."

[1588] Step 5:

[1589] The server compiles the generated story text and images into a single document and sends it to the terminal. The input is the story text and image files, and the output is the compiled document file and a notification of completion of transmission. The data is compiled into a JSON format document file and sent to the terminal.

[1590] Step 6:

[1591] The device analyzes the received document file and provides it to the user through a display method. For example, using React Native, it can be displayed on a smartphone, tablet, smart glasses, or head-mounted display. The input is the document file, and the output is the story and images displayed on the device.

[1592] Step 7:

[1593] The device uses Amazon Polly to read aloud the story text generated, allowing users to enjoy the story not only visually but also aurally. The input is the story text, and the output is audio data, specifically the content of the story being read aloud.

[1594] Each step is processed sequentially to generate a personalized story and images based on the user's emotions, ultimately providing an interactive storytelling experience.

[1595] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1596] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1597] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1598] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1599] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1600] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1601] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1602] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1603] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1604] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1605] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1606] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1607] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1608] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1609] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1610] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1611] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1612] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1613] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1614] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1615] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1616] The following is further disclosed regarding the above embodiment.

[1617] (Claim 1)

[1618] A generating means for generating a story for each emotion;

[1619] image generation means for generating images based on the generated story;

[1620] display means for displaying the generated story and images;

[1621] A system including:

[1622] (Claim 2)

[1623] selection means for selecting an emotional domain based on user input;

[1624] a story generation means for generating a story based on the selected emotional domain;

[1625] The system of claim 1 further comprising:

[1626] (Claim 3)

[1627] 10. The system of claim 1, including a display means without a backlight, the display means using electronic paper.

[1628] "Example 1"

[1629] (Claim 1)

[1630] a selection means for a user to select an emotion area;

[1631] a story generation means including a generative AI model for generating a story based on the selected emotional domains;

[1632] image generation means for generating image prompts based on the generated narrative and for generating images based on the prompts, the image generation means including an image generation AI model;

[1633] a transmitting means for transmitting the generated story and images to a user's terminal;

[1634] display means for displaying the transmitted story and images;

[1635] A system including:

[1636] (Claim 2)

[1637] selection means for selecting an emotional domain based on user input;

[1638] 10. The system of claim 1, further comprising a storytelling means for a user to read a story to a child.

[1639] (Claim 3)

[1640] 10. The system of claim 1, including a display means without a backlight, the display means using electronic paper.

[1641] "Application Example 1"

[1642] (Claim 1)

[1643] A generating means for generating a story for each emotion;

[1644] image generation means for generating images based on the generated story;

[1645] display means for displaying the generated story and images;

[1646] Generate stories and images for each emotion, display them when users place orders, and encourage them to read aloud while waiting for delivery.

[1647] A system including:

[1648] (Claim 2)

[1649] selection means for selecting an emotional domain based on user input;

[1650] a story generation means for generating a story based on the selected emotional domain;

[1651] a transmission means for transmitting the generated story and images to a user's terminal when ordering delivery, and reading the story and images to the user's terminal;

[1652] The system of claim 1 further comprising:

[1653] (Claim 3)

[1654] 10. The system of claim 1, including a display means without a backlight, the display means using electronic paper.

[1655] "Example 2: Combining Emotion Engines"

[1656] (Claim 1)

[1657] recognition means for recognizing an emotion of a user;

[1658] a story generation means for generating a story based on the recognized emotions;

[1659] image generation means for generating images based on the generated story;

[1660] display means for displaying the generated story and images;

[1661] A system including:

[1662] (Claim 2)

[1663] selection means for selecting an emotional domain based on user input;

[1664] a story generation means for generating a story based on the selected emotional domain;

[1665] The system of claim 1 further comprising:

[1666] (Claim 3)

[1667] Display means using electronic paper,

[1668] a transmitting means for transmitting the generated story to a terminal;

[1669] a reading support means for reading the generated story to a child;

[1670] The system of claim 1 further comprising:

[1671] "Application example 2 when combining emotion engines"

[1672] (Claim 1)

[1673] A generating means for generating a story for each emotion;

[1674] image generation means for generating images based on the generated story;

[1675] display means for displaying the generated story and images;

[1676] an emotion recognition means for automatically recognizing emotions;

[1677] an audio playback means for reading the generated story aloud;

[1678] Cross-platform support for multiple display devices;

[1679] A system including:

[1680] (Claim 2)

[1681] selection means for selecting an emotional domain based on user input;

[1682] a story generation means for generating a story based on the selected emotional domain;

[1683] 10. The system of claim 1.

[1684] (Claim 3)

[1685] 10. The system of claim 1, including a display means without a backlight, the display means using electronic paper. [Explanation of symbols]

[1686] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A generating means for generating a story for each emotion; image generation means for generating images based on the generated story; display means for displaying the generated story and images; A system including:

2. selection means for selecting an emotional domain based on user input; a story generation means for generating a story based on the selected emotional domain; The system of claim 1 further comprising:

3. 10. The system of claim 1, further comprising a display means without a backlight, the display means using electronic paper.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A