system

A customizable picture book system using AI generation and feedback mechanisms addresses the challenge of busy guardians by providing tailored, engaging books that adapt to children's growth and interests.

JP2026073479APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Busy guardians lack time to read picture books to children, and there is a need for a method to efficiently provide picture books that individually consider the vocabulary and interests necessary for a child's growth, with a mechanism that supports continuous growth by reflecting the child's reaction to the generated picture book in the creation of the next book.

Method used

A system that allows users to customize picture books through input on story content, image style, number of pages, and message, using AI to generate illustrations and text, with a feedback function to incorporate children's impressions into the next book creation, supporting individual growth.

Benefits of technology

Enables easy creation of customized picture books that enhance learning experiences for children, allowing for continuous improvement based on their feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026073479000001_ABST
    Figure 2026073479000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of receiving input from users regarding the story content, image style, number of pages, and the message they want to convey. A means for creating instructions necessary for story generation based on the above input information, A means for generating images and text based on the above instructions, A means of creating a manuscript for an electronic picture book by combining the generated images and text, A means to enable users to download the above-mentioned digital picture book, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] There is a problem that busy guardians lack time to read picture books to children. Therefore, there is a need for a method to efficiently provide picture books that individually consider the vocabulary and interests necessary for a child's growth. In addition, there is a need for a means to easily customize the content of picture books and provide a more familiar experience for children. Furthermore, a mechanism that continuously supports a child's growth by reflecting the child's reaction to the generated picture book in the creation of the next picture book is also essential.

Means for Solving the Problems

[0005] This invention enables the creation of individually customized picture books by providing a means for receiving input from users regarding story content, image style, number of pages, and the message they wish to convey. Furthermore, by providing a means for creating instructions necessary for story generation based on the input information and generating illustrations and text, the manuscript for an electronic picture book can be created quickly and efficiently. In addition, the generated electronic picture book can be easily used by the user by downloading it, and the option of printing and delivery is also provided. Moreover, by providing a feedback function that obtains children's impressions of the generated electronic picture book and incorporates them into the creation of the next picture book, a system that supports intellectual development tailored to the individual growth of each child is realized.

[0006] A "user" is an individual or group that operates the system and generates picture books for themselves or their children.

[0007] "The content of the story" refers to the specific details of the story or theme that the picture book aims to convey.

[0008] "Image style" is a concept that refers to the visual form and design principles of illustrations in picture books.

[0009] "Page count" refers to the total number of pages in the picture book that will be produced.

[0010] A "message" is the main theme or educational intention that the picture book aims to convey to the reader.

[0011] A "command" is a specific instruction given to the generated AI based on the input information.

[0012] "Picture" refers to an image containing visual information created by a generative AI.

[0013] "Text" refers to character information, including sentences created by generative AI.

[0014] An "electronic picture book" is a picture book in digital format that can be viewed on a display or similar device.

[0015] The "manuscript" is a pre-completed document combining pictures and text, which will become the design drawing of the final picture book.

[0016] "Downloadable" means that data can be stored on a terminal via the Internet.

[0017] The "feedback function" is a mechanism that collects impressions and evaluations from users or readers and reflects them in the system.

Brief Description of Drawings

[0018] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.

Mode for Carrying Out the Invention

[0019] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be described.

[0021] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0022] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0023] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0024] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0026] [First Embodiment]

[0027] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0028] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0031] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0034] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0038] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0039] The system of the present invention automatically generates picture books via a web interface. Specific embodiments are described below.

[0040] First, the device displays an input form for the user to customize the content of the picture book. This form can accept input such as the story content, image style, number of pages, message to be conveyed, and, if necessary, the names of the characters.

[0041] Users enter this information based on their preferences. Once the input is complete, this information is sent to the server in JSON format.

[0042] The server analyzes the received information and creates the necessary instructions for generating images and text. These instructions are then passed to the image generation AI and text generation AI, which carry out the specific generation processes.

[0043] The server invokes an image generation AI based on prompts and generates an image in the specified style. For example, an illustration of a starry sky in a watercolor style is generated. Additionally, a text generation AI is used to generate narration that takes vocabulary level into consideration. For example, a sentence like, "That night, under countless stars, the adventurers..." is generated.

[0044] The generated images and text are combined by the server to create a picture book manuscript in HTML or PDF format. This harmonizes the images and text, resulting in a digitally completed, original picture book based on the user's specified content.

[0045] Ultimately, the server makes this digital picture book available for download on devices, providing users with free access to it. Users can view this picture book on digital devices such as tablets and read it aloud to their children, and they can also optionally order a printed version.

[0046] Furthermore, feedback and reactions to the generated picture books are sent back to the server using a dedicated input device or stuffed animal. Based on this feedback, it becomes possible to provide content tailored to the child's interests and development when creating the next picture book.

[0047] This system allows parents to easily create customized picture books, providing their children with an enjoyable learning experience.

[0048] The following describes the processing flow.

[0049] Step 1:

[0050] The device provides the user with a web interface, displaying a form to enter the picture book's content, image style, number of pages, message to be conveyed, and, if necessary, the names of the characters.

[0051] Step 2:

[0052] The user enters the required information into the input form mentioned above and submits the completed data. The input content is sent to the server in JSON format.

[0053] Step 3:

[0054] The server parses the received JSON data and generates an AI prompt based on the input. This prompt contains the necessary instructions for both the image generation AI and the text generation AI.

[0055] Step 4:

[0056] The server invokes an image generation AI and generates an image in the specified style according to the prompts. The generated image is temporarily stored in cloud storage.

[0057] Step 5:

[0058] The server invokes a text generation AI to generate sentences suitable for the story based on prompts. The generated text is checked to ensure it is easy to understand and matches the child's vocabulary level.

[0059] Step 6:

[0060] The server combines the generated images and text to create a digital picture book manuscript in HTML or PDF format. It then arranges the layout page by page to complete the final manuscript.

[0061] Step 7:

[0062] The server generates a download link for the completed digital picture book and provides it to the user's device. The user can then download the picture book from this link and view it on a tablet or other device.

[0063] Step 8:

[0064] If the user selects the bookbinding option, they will enter the additional information required for bookbinding and delivery on their terminal and send that information to the server. The server will then arrange the bookbinding process and delivery.

[0065] Step 9:

[0066] The system records children's thoughts and reactions to the generated picture books via devices or stuffed animals, and sends this feedback to a server. This information is then considered when creating future picture books, and used to provide content that is appropriate for the child's development.

[0067] (Example 1)

[0068] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0069] Traditional picture book creation systems had problems such as insufficient customization options and inadequate incorporation of feedback on the generated products. Furthermore, the options for providing the generated digital books in physical form were limited. As a result, it was difficult for users to create picture books that met their individual needs and to make subsequent improvements.

[0070] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0071] In this invention, the server includes means for receiving input from the user regarding the story content, image style, number of pages, and the message to be conveyed; means for converting the input information into a data format and analyzing it; and means for creating instructions for an image generation AI model and a text generation AI model based on the data analysis. This enables the creation of customizable picture books that meet the diverse needs of users and allows for continuous quality improvement through feedback on the generated content.

[0072] "User" refers to an individual or user who uses the system to input customized information for stories and images and receives the generated content.

[0073] A "server" is an information processing device that receives and analyzes input data, sends commands to the generating AI model, and stores and provides the generated content.

[0074] "Data format" refers to a format that structures information and converts it into a form that can be easily parsed, and generally includes formats such as JSON and XML.

[0075] "Data analysis" is the process of processing input information and extracting useful commands and patterns.

[0076] An "image generation AI model" refers to an algorithm that generates images in an artistic or requested style based on given instructions.

[0077] A "text generation AI model" refers to an algorithm that generates natural language stories and explanatory texts based on given instructions.

[0078] A "command" refers to a command or prompt that instructs an AI model to generate content in a specific form or with specific content.

[0079] A "digital book" refers to a book in electronic format that is accessible on a computer or electronic device, and is presented as multimedia content including text and images.

[0080] "Feedback" refers to evaluations and opinions provided by users, and is information used to improve the system and for future updates.

[0081] A "physical book" refers to a product in book form created by printing digitally generated content and using paper or other physical materials.

[0082] This invention is a system that automatically generates picture books customized by users via a web interface. The system mainly consists of a server, a terminal, and a generation AI model.

[0083] The device displays an input form to the user, which is necessary for customizing the story. This form is used to input the story's content, image style, number of pages, message to convey, character names, and other information. After the user enters this information, the device uses JavaScript® and HTML to convert the data into JSON format and send it to the server.

[0084] The server analyzes the received data and creates instructions to pass to the image generation AI model and the text generation AI model. A Python script is used for this analysis, and the instructions are formed as specific prompts. For example, the image generation AI is given a prompt such as "Generate a watercolor-style adventure scene with a starry sky background," and the text generation AI is given a prompt such as "Create a story about children adventuring under a starry sky."

[0085] The server then uses this prompt to call an image generation AI model to generate an image in the specified style, and a text generation AI model to generate narrative text that takes vocabulary and expression into consideration. The specific processing utilizes APIs for the generation AI. For example, a Generative Adversarial Network (GAN) is used for image generation, and a natural language processing model such as GPT-3 (registered trademark) is used for text generation.

[0086] The generated images and text are combined by the server in HTML or PDF format and saved as a digital book. The Python PIL library and ReportLab are used in this process.

[0087] Finally, the server provides the terminal with a download link for the finished product. Users can use the download link to view the picture book on their digital device and, if necessary, send feedback back to the server through a dedicated form. This feedback will be considered in the next generation process and used to improve the system.

[0088] This system allows users to easily customize and create picture books tailored to their individual needs, providing an experience that enhances learning effectiveness, especially for children.

[0089] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0090] Step 1:

[0091] The device displays a customization input form for the story to the user. Here, the user can enter details such as the story's content, image style, number of pages, message, and character names. This form is implemented using HTML and JavaScript. Upon completion, the device converts the input data into JSON format and sends it to the server.

[0092] Step 2:

[0093] The server parses the JSON data received from the terminal. This parsing process uses Python, and based on the provided data, it creates instructions for the image generation AI model and the text generation AI model. Specifically, the data processing involves parsing user input to generate prompt sentences, such as "Generate a watercolor-style adventure scene with a starry sky background."

[0094] Step 3:

[0095] The server calls an image generation AI model based on the generated prompt text and generates an image in the specified style. In this step, the prompt text is passed to the image generation AI via API, the AI ​​generates an appropriate image, and the image data is returned to the server. For example, a "scene of an adventurer with stars shining in the night sky" generated using GAN is one such example.

[0096] Step 4:

[0097] The server invokes a text generation AI model according to a prompt and generates narrative text using natural language processing. A prompt such as "Create a story about children adventuring under the starry sky" is used, and the AI ​​generates narration accordingly, outputting text such as "That night, under the shining starry sky, the adventure began..."

[0098] Step 5:

[0099] The server combines the generated images and text and records them as a digital book. This step uses Python's PIL library or ReportLab to arrange the images and text page by page and save them in HTML or PDF format. The output is an ebook file, ready for user viewing.

[0100] Step 6:

[0101] The server provides the terminal with a download link for the completed digital book. Users can download the picture book via this link and view it on their device. Furthermore, users can provide feedback to the server after use, which can be used in the next generation process. This feedback is accumulated to improve the system.

[0102] (Application Example 1)

[0103] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0104] In recent years, with the widespread adoption of digital devices, the demand for digital content that parents and children can enjoy together has increased. However, existing picture book-related content has difficulty reflecting the individual preferences of users, and there are limited means of providing interactive experiences such as creating stories together as a family. This invention aims to solve these problems by allowing users to customize the content of stories, supporting collaborative creation between parents and children, and making it instantly viewable on digital devices.

[0105] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0106] In this invention, the server includes means for receiving input from the user regarding the content of the document, the style of the images, the number of pages, and the message to be conveyed; means for making the generated electronic document viewable on a mobile information terminal; and means for presenting options for a digital version and a printed version of the generated document. This makes it possible for parents and children to collaboratively create original stories and instantly share and enjoy them on digital devices.

[0107] "Users" refers to individuals who use this system to customize documents and images and generate personalized digital content.

[0108] "Document content" refers to the text information of the story and explanations included in the generated digital picture book.

[0109] "Image style" refers to the artistic expression method of the generated visual content, and includes styles such as watercolor.

[0110] A "generative AI model" refers to an artificial intelligence system that automatically generates specific images and text based on user input.

[0111] A "prompt statement" refers to an input statement used to instruct a generative AI model on specific outputs.

[0112] "Personal information terminals" refer to computing devices that users can carry with them, such as smartphones and tablets.

[0113] The "feedback function" refers to a feature that collects user opinions and reactions to generated documents and uses them to improve the next generation process.

[0114] "Created jointly by parent and child" refers to the process where a parent and child work together to conceive the content and create a single digital piece of content.

[0115] This invention provides a specific embodiment of an interactive picture book generation system that allows parents and children to instantly enjoy stories they have collaboratively edited on a digital device.

[0116] The server displays a form on the terminal via a web interface to receive input information from the user, such as the content of the document, image format, number of pages, and the message to be conveyed. This information is sent to the server in JSON format. The server uses a generative AI model to dynamically generate images and text based on the received information. Specifically, prompt statements are used to instruct the AI ​​on what kind of images and text it should generate. For example, a prompt statement such as "Draw a watercolor-style picture of a beach and adventurers" might be used.

[0117] The terminal makes the generated electronic document viewable on a mobile device. This allows users to instantly enjoy the original story they created together as a family. Furthermore, the server presents users with the option of choosing between a digital version and a printed version of the generated document. Users can also obtain a physical printed picture book upon request.

[0118] The device also has a feedback function that sends comments and reactions to the generated document to the server. This feedback is reflected in the next story generation, allowing for content tailored to the child's interests and development. For example, when creating a picture book with a summer vacation adventure theme, the server generates images and text suitable for the document content, such as "adventurers visiting a seaside town."

[0119] In this way, the system of this invention makes it possible for parents and children to jointly generate customized digital content and provide enjoyable learning and experiences through it.

[0120] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0121] Step 1:

[0122] The terminal displays a web interface form for the user to input information such as story content, image style, number of pages, and the message to be conveyed. The entered information is organized as structured data and prepared for later transmission to the server. In this step, the input is information entered by the user, and the output is the organized data.

[0123] Step 2:

[0124] When a user presses the submit button on an input form, the terminal converts the prepared data into JSON format and sends it to the server. In this step, the input is prepared data, and the output is data encoded in JSON format.

[0125] Step 3:

[0126] The server receives JSON data sent from the terminal and analyzes it. Based on the received information, it creates the prompts necessary for the image generation AI and text generation AI. The input for this step is JSON data, and the output is prompts.

[0127] Step 4:

[0128] The server invokes a generative AI model, which generates an image in the specified style and constructs text according to the created prompt. The input for this step is the prompt, and the output is the generated image data and text data. Specifically, the image generation AI model receives instructions such as "Please draw a watercolor-style picture of a beach and adventurers" and generates an image.

[0129] Step 5:

[0130] The server combines the generated images and text to create the manuscript for the digital picture book. This manuscript is then formatted into a digital format such as HTML or PDF, making it viewable by users. The input for this step is image data and text data, and the output is the manuscript for the digital picture book.

[0131] Step 6:

[0132] The server generates and presents a link on the terminal that allows the user to download the completed digital picture book manuscript. The input for this step is the digital picture book manuscript, and the output is the download link.

[0133] Step 7:

[0134] The user uses a download link to save the e-book to their device and begins viewing it on their mobile device. In this step, the digital content, the e-book, is utilized at the user level.

[0135] Step 8:

[0136] After the user experiences the digital picture book, the device displays a feedback form, allowing them to enter their thoughts and reactions to the book. The entered feedback is sent to the server to be used in generating future stories. The input in this step is user feedback, and the output is feedback data.

[0137] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0138] This invention is a system that recognizes the user's emotions using an emotion engine and adjusts the content of the digital picture book based on that data to provide the most suitable picture book experience for each individual user. A specific embodiment of this system is described below.

[0139] First, the device provides the user with an input interface to customize the story content, image style, and number of pages. The user enters this information, and the data is sent to the server in JSON format.

[0140] The server analyzes the received data and generates AI prompts based on the input information. These generated prompts are passed to image and text generation AIs, which then generate materials based on the specified style and content. The generated illustrations and text are assembled into a story and saved in the form of a digital picture book.

[0141] Next, when a picture book is read aloud, the device activates an emotion engine. The emotion engine analyzes the user's, especially the child's, facial expressions and voice through the camera and microphone, and evaluates their current emotional state in real time. For example, if a child shows a surprised expression, the reaction can be visualized, and the next development of the story can be adjusted to increase engagement.

[0142] For example, if a child feels anxious during a tense scene in a story, the emotion engine can detect this and make adjustments such as inserting a softer, calmer scene on the next page.

[0143] Furthermore, the data accumulated by the emotion engine is stored on a server, and user preferences and reaction patterns are analyzed. This information is then fed back into future story generation, resulting in the creation of more personalized picture books.

[0144] Furthermore, user-customized digital picture books are provided via a downloadable link from the server, and if printed copies are desired as an option, production and delivery procedures are arranged.

[0145] This invention allows parents to read aloud in a way that resonates with their child's emotions, and allows children to enjoy stories tailored to their own needs.

[0146] The following describes the processing flow.

[0147] Step 1:

[0148] The device provides a web interface and displays a form where the user can input the story content, image style, number of pages, and the message they want to convey.

[0149] Step 2:

[0150] The user fills in the required information in the input form and completes the input. The data is then sent to the server in JSON format.

[0151] Step 3:

[0152] The server analyzes the received input data and creates prompts necessary for story generation. These prompts contain the information needed to generate both images and text.

[0153] Step 4:

[0154] The server invokes an image generation AI to generate an image in the specified style based on prompts. The generated image is temporarily stored in cloud storage.

[0155] Step 5:

[0156] The server invokes a text generation AI to generate sentences appropriate to the story from the prompts. This includes content tailored to a child's vocabulary level.

[0157] Step 6:

[0158] The server combines the generated images and text to create a digital picture book manuscript in HTML or PDF format. This manuscript includes the overall story structure and page layout.

[0159] Step 7:

[0160] The device activates an emotion engine and uses a camera and microphone to recognize the user's emotions in real time while reading picture books aloud.

[0161] Step 8:

[0162] The emotion engine detects emotional responses from the user's facial expressions and voice, and if, for example, a child is surprised, it receives data to adjust the content placed on the next page.

[0163] Step 9:

[0164] The server analyzes emotional data and adjusts the content of the picture book in real time based on the results. This provides a story experience tailored to the child's emotions.

[0165] Step 10:

[0166] The emotional data obtained by the emotion engine will be stored on a server and used to analyze the preferences of individual users in future picture book creation.

[0167] Step 11:

[0168] The server generates a link for the user to download the e-book and displays it on the device. If the user selects the printed book option, the server then handles information to arrange the shipping process.

[0169] (Example 2)

[0170] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0171] Existing digital picture books make it difficult to customize stories to suit the individual emotions and preferences of users. As a result, they cannot consistently provide users with an engaging and personalized storytelling experience, and there is a particular problem in that they cannot adequately deliver emotionally impactful read-aloud sessions, especially to children.

[0172] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0173] This invention includes a server that analyzes user input information to generate commands for a generation AI model and uses those commands to produce comprehensive story material; a server that analyzes the user's expressions and voice in real time and evaluates their emotions; and a server that dynamically adjusts the progression of the story according to the emotional evaluation to provide an adaptive electronic picture book experience. This makes it possible to dynamically adjust the story based on the user's individual emotions and preferences and provide an optimized picture book experience at all times.

[0174] "Users" refers to individual people who use the system to customize stories and receive an optimized picture book experience.

[0175] "The content of the story" refers to all elements related to the development of the story, such as the storyline, theme, and characters.

[0176] "Image style" refers to the artistic or design direction of the illustrations and visual materials used in a story.

[0177] "Page count" refers to the total number of pages that make up an electronic picture book, and is an element that indicates the length or volume of the story.

[0178] A "generative AI model" refers to artificial intelligence technology that automatically generates images and text based on input prompts.

[0179] A "prompt statement" is a text containing instructions given to a generative AI model, specifically a statement that directs the model to generate concrete content.

[0180] "Emotional assessment" refers to the act of analyzing and evaluating a user's current emotional state in real time based on their facial expressions and voice.

[0181] An "adaptive digital picture book experience" refers to a function that dynamically adjusts the story's progression and content according to the user's emotions and preferences, providing a personalized read-aloud experience.

[0182] This invention is a system that allows users to customize digital picture books based on their own emotions and preferences. First, the terminal provides the user with an intuitive interface where they can input the story content, image style, number of pages, etc. This allows the user to select specific keywords and create a customized story outline based on them.

[0183] The server receives the input data and analyzes it using programming languages ​​such as Python or JavaScript. From this analysis, it creates prompts to supply to the generating AI model. These prompts are provided in text format and might include something like, "Generate a space adventure story for children, focusing on friendship with colorful illustrations." Based on these prompts, AI technologies such as Stable Diffusion for image generation and the GPT series for text generation are used to specifically generate the story and illustrations.

[0184] The server integrates the generated illustrations and text to construct the manuscript for the digital picture book. This results in a digital work with narrative coherence and visual appeal.

[0185] Furthermore, during story time, the device uses its camera and microphone to transmit the user's facial expressions and voice to the emotion engine in real time. The emotion engine analyzes this data and evaluates the user's emotional state, dynamically adjusting the story's progression. For example, if a child shows a surprised expression, the story's development shifts towards a more reassuring direction to increase engagement.

[0186] The generated digital picture books are provided as downloadable links via a server, and printed and delivered versions are available upon request. In this way, users can enjoy a unique storytelling experience tailored to their individual emotions.

[0187] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0188] Step 1:

[0189] The terminal provides the user with an input interface. Here, the user can select or input information such as the story content, image style, and number of pages. This input information is collected as data in JSON format. Specifically, data is acquired using touchscreen or voice recognition technology.

[0190] Step 2:

[0191] The server receives JSON data sent from the terminal and performs analysis. A Python script is used for data analysis to extract necessary information and generate prompts required for the AI ​​model. At this stage, the data is structured, and specific keywords such as "space exploration" are extracted.

[0192] Step 3:

[0193] The server creates instructions for text and image generation AI based on the generated prompt text. These instructions are then passed to the generation AI model (e.g., Stable Diffusion or GPT series) to generate illustrations and narrative text based on the specified style and theme. Using the AI ​​prompt as input, the system obtains theme-appropriate images and text as output.

[0194] Step 4:

[0195] The server integrates the generated illustrations and text to create an electronic picture book. This creation utilizes programmatic digital layout, carefully considering the story's order and page count. As a result, a visually engaging and consistent work is produced.

[0196] Step 5:

[0197] The device uses a camera and microphone to capture the user's facial expressions and voice while a picture book is being read aloud, and transmits this data to the emotion engine. Real-time video and audio data are used as input for emotion analysis. Specifically, it performs facial recognition and voice analysis.

[0198] Step 6:

[0199] The server receives the results of emotion analysis and adjusts the story's progression accordingly. For example, if the user feels anxious, a calming scene is automatically inserted on the next page. This dynamic adjustment provides an optimal story experience that resonates with the user's emotions.

[0200] Step 7:

[0201] The server then provides the user with a customized digital picture book and generates a downloadable link. If the user wishes to have the book printed or delivered, the server handles those procedures. The final output is the individual picture book data, which is shared with the user.

[0202] (Application Example 2)

[0203] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0204] There is a challenge in providing individual users, especially children, with an appropriate and enjoyable story experience tailored to their emotions at any given time when enjoying ebooks. Furthermore, static content makes it difficult to fully personalize the user experience, resulting in a lack of increased user engagement.

[0205] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0206] In this invention, the server includes means for receiving input from the user regarding the story content, image style, number of pages, and message to be conveyed; means for creating an ebook manuscript by combining the generated images and text; and means for sensing the user's emotional state and dynamically adjusting the story content accordingly. This makes it possible to provide a personalized story experience that responds to the user's real-time emotions.

[0207] "User" refers to an individual or organization that uses the system to generate and customize stories.

[0208] A "story" refers to a series of texts and images combined to create an experience for the user, such as a narrative or episodes.

[0209] "Images" refer to the visual elements included in an e-book, specifically illustrations and diagrams generated to visually represent the story's content and style.

[0210] "Style" refers to the elements that characterize the appearance and design of a story or image, and it is the criterion that users use to select a particular theme or atmosphere.

[0211] "Page count" refers to the total number of pages in which the story unfolds within an ebook, and is an indicator used by users to determine the length of the story.

[0212] A "message" refers to the intention or meaning conveyed to the user through a story, and may include specific values ​​or lessons.

[0213] "Instructions" refer to data containing specific commands and settings necessary for generating a story, and are generated by the system.

[0214] "E-books" refer to a form of book that is stored in digital format and can be viewed by users through electronic devices.

[0215] "Emotional state" refers to data that indicates a user's psychological or emotional response, and includes information collected from the user's facial expressions and voice while they are browsing.

[0216] "Customization" refers to the process of adjusting the content and format of a story to suit the specific needs and preferences of a particular user.

[0217] This invention is an e-book system that dynamically adjusts the content of a story according to the user's emotional state. When a user inputs the story content, image style, number of pages, and message they want to convey into their device, this information is sent to a server. Based on this input information, the server creates the necessary instructions for story generation and generates prompt sentences using a generation AI model. These prompt sentences trigger image and text generation, creating the material for the story.

[0218] The device uses a camera and microphone to collect the user's facial expressions and voice in real time, and an emotion engine recognizes their emotional state. For example, OpenCV captures camera footage, and the microphone library acquires audio data. The server takes this emotional state into consideration and dynamically adjusts the content of the generated story. Depending on the emotional change, the next scene can be switched to content that provides a sense of security.

[0219] For example, if a user feels anxious during a tense scene in a story, a friendly character who alleviates the tension can be introduced in the next scene, creating a reassuring development. This allows users to receive a personalized experience tailored to their individual circumstances.

[0220] An example of a prompt to input into a generation AI model is, "The user is currently feeling anxious. Please take this emotion into consideration and generate a storyline in the next scene that will provide a sense of reassurance." By providing stories that resonate with the user in this way, the enjoyment of ebooks can be maximized.

[0221] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0222] Step 1:

[0223] The user inputs the story content, image style, number of pages, and message via their device. The entered data is temporarily stored on the device and converted to JSON format. This is the stage where the user's intentions and requests are concretized.

[0224] Step 2:

[0225] The terminal sends the converted JSON data to the server. The server parses the received data and creates the instructions necessary for story generation. Through the parsing process, the server extracts each element and reformats it into the appropriate format. The output is a prompt sentence to be input to the generating AI model.

[0226] Step 3:

[0227] The server inputs prompt text into the generative AI model, which then generates image and text materials. The generative AI model performs data calculations based on the instructions and constructs illustrations and text to match the settings. The generated materials are obtained as output.

[0228] Step 4:

[0229] The generated materials are combined on the server to create the manuscript for the ebook. The server integrates the materials to construct the overall story and compiles it into the final digital format. The output is the data of the completed ebook.

[0230] Step 5:

[0231] The device collects the user's facial expressions and voice using a camera and microphone, and analyzes them with an emotion engine. It captures video using OpenCV and records audio using the microphone library. The input consists of audio and video data, and the analysis yields the user's emotional state.

[0232] Step 6:

[0233] The server adjusts the content of the ebook based on the emotional state it receives. Using the results of the emotion engine, it replaces or rearranges text and images. As a result, the output is a story that corresponds to the user's emotions.

[0234] Step 7:

[0235] The adjusted ebook is returned to the user's device and becomes available for download. The final ebook data is sent to the device, and the user can use it for viewing or reading aloud. The output is a usable ebook.

[0236] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0237] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0238] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0239] [Second Embodiment]

[0240] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0241] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0242] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0243] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0244] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0245] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0246] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0247] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0248] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0249] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0250] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0251] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0252] The system of the present invention automatically generates picture books via a web interface. Specific embodiments are described below.

[0253] First, the device displays an input form for the user to customize the content of the picture book. This form can accept input such as the story content, image style, number of pages, message to be conveyed, and, if necessary, the names of the characters.

[0254] Users enter this information based on their preferences. Once the input is complete, this information is sent to the server in JSON format.

[0255] The server analyzes the received information and creates the necessary instructions for generating images and text. These instructions are then passed to the image generation AI and text generation AI, which carry out the specific generation processes.

[0256] The server invokes an image generation AI based on prompts and generates an image in the specified style. For example, an illustration of a starry sky in a watercolor style is generated. Additionally, a text generation AI is used to generate narration that takes vocabulary level into consideration. For example, a sentence like, "That night, under countless stars, the adventurers..." is generated.

[0257] The generated images and text are combined by the server to create a picture book manuscript in HTML or PDF format. This harmonizes the images and text, resulting in a digitally completed, original picture book based on the user's specified content.

[0258] Ultimately, the server makes this digital picture book available for download on devices, providing users with free access to it. Users can view this picture book on digital devices such as tablets and read it aloud to their children, and they can also optionally order a printed version.

[0259] Furthermore, feedback and reactions to the generated picture books are sent back to the server using a dedicated input device or stuffed animal. Based on this feedback, it becomes possible to provide content tailored to the child's interests and development when creating the next picture book.

[0260] This system allows parents to easily create customized picture books, providing their children with an enjoyable learning experience.

[0261] The following describes the processing flow.

[0262] Step 1:

[0263] The device provides the user with a web interface, displaying a form to enter the picture book's content, image style, number of pages, message to be conveyed, and, if necessary, the names of the characters.

[0264] Step 2:

[0265] The user enters the required information into the input form mentioned above and submits the completed data. The input content is sent to the server in JSON format.

[0266] Step 3:

[0267] The server parses the received JSON data and generates an AI prompt based on the input. This prompt contains the necessary instructions for both the image generation AI and the text generation AI.

[0268] Step 4:

[0269] The server invokes an image generation AI and generates an image in the specified style according to the prompts. The generated image is temporarily stored in cloud storage.

[0270] Step 5:

[0271] The server invokes a text generation AI to generate sentences suitable for the story based on prompts. The generated text is checked to ensure it is easy to understand and matches the child's vocabulary level.

[0272] Step 6:

[0273] The server combines the generated images and text to create a digital picture book manuscript in HTML or PDF format. It then arranges the layout page by page to complete the final manuscript.

[0274] Step 7:

[0275] The server generates a download link for the completed digital picture book and provides it to the user's device. The user can then download the picture book from this link and view it on a tablet or other device.

[0276] Step 8:

[0277] If the user selects the bookbinding option, they will enter the additional information required for bookbinding and delivery on their terminal and send that information to the server. The server will then arrange the bookbinding process and delivery.

[0278] Step 9:

[0279] The system records children's thoughts and reactions to the generated picture books via devices or stuffed animals, and sends this feedback to a server. This information is then considered when creating future picture books, and used to provide content that is appropriate for the child's development.

[0280] (Example 1)

[0281] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0282] Traditional picture book creation systems had problems such as insufficient customization options and inadequate incorporation of feedback on the generated products. Furthermore, the options for providing the generated digital books in physical form were limited. As a result, it was difficult for users to create picture books that met their individual needs and to make subsequent improvements.

[0283] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0284] In this invention, the server includes means for receiving input from the user regarding the content of the story, the style of the image, the number of pages, and the message to be conveyed, means for converting the input information into a data format for analysis, and means for creating instructions for an image generation AI model and a text generation AI model based on the data analysis. This enables the generation of customizable picture books according to the diverse needs of the user and continuous quality improvement through feedback on the generated content.

[0285] The "user" refers to an individual or user who inputs customization information for stories and images using the system and receives the generated content.

[0286] The "server" is an information processing device that receives and analyzes the input data, and is a device that plays a role in sending instructions to the generation AI model and storing and providing the generated content.

[0287] The "data format" refers to a format in which information is structured and converted into an easily analyzable form, and generally includes JSON, XML, etc.

[0288] "Data analysis" is a process of processing the input information and extracting useful instructions and patterns.

[0289] The "image generation AI model" refers to an algorithm that generates an artistic or required style image based on the given instructions.

[0290] The "text generation AI model" refers to an algorithm that generates a natural language story or explanatory text based on the given instructions.

[0291] The "instruction" refers to an instruction sentence or prompt that instructs the AI model to generate content in a specific form and content.

[0292] A "digital book" refers to a book in electronic format that is accessible on a computer or electronic device, and is presented as multimedia content including text and images.

[0293] "Feedback" refers to evaluations and opinions provided by users, and is information used to improve the system and for future updates.

[0294] A "physical book" refers to a product in book form created by printing digitally generated content and using paper or other physical materials.

[0295] This invention is a system that automatically generates picture books customized by users via a web interface. The system mainly consists of a server, a terminal, and a generation AI model.

[0296] The device displays an input form to the user, which is necessary for customizing the story. This form is used to input the story's content, image style, number of pages, message to convey, character names, and other information. After the user enters this information, the device uses JavaScript and HTML to convert the data into JSON format and send it to the server.

[0297] The server analyzes the received data and creates instructions to pass to the image generation AI model and the text generation AI model. A Python script is used for this analysis, and the instructions are formed as specific prompts. For example, the image generation AI is given a prompt such as "Generate a watercolor-style adventure scene with a starry sky background," and the text generation AI is given a prompt such as "Create a story about children adventuring under a starry sky."

[0298] The server then uses this prompt to call an image generation AI model to generate an image in the specified style, and a text generation AI model to generate narrative text that takes vocabulary and expression into consideration. The specific processing utilizes APIs for the generation AI. For example, a Generative Adversarial Network (GAN) is used for image generation, and a natural language processing model such as GPT-3 is used for text generation.

[0299] The generated images and text are combined by the server in HTML or PDF format and saved as a digital book. The Python PIL library and ReportLab are used in this process.

[0300] Finally, the server provides the terminal with a download link for the finished product. Users can use the download link to view the picture book on their digital device and, if necessary, send feedback back to the server through a dedicated form. This feedback will be considered in the next generation process and used to improve the system.

[0301] This system allows users to easily customize and create picture books tailored to their individual needs, providing an experience that enhances learning effectiveness, especially for children.

[0302] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0303] Step 1:

[0304] The device displays a customization input form for the story to the user. Here, the user can enter details such as the story's content, image style, number of pages, message, and character names. This form is implemented using HTML and JavaScript. Upon completion, the device converts the input data into JSON format and sends it to the server.

[0305] Step 2:

[0306] The server analyzes the JSON-formatted data received from the terminal. Python is used for this analysis process, and commands are created for the image generation AI model and the text generation AI model based on the provided data. As specific data processing, it analyzes the user input to generate a prompt sentence, and creates a prompt such as "Generate an adventure scene in the style of watercolor painting with a starry sky as the background".

[0307] Step 3:

[0308] Based on the generated prompt sentence, the server calls the image generation AI model to generate an image in the specified style. In this step, the prompt sentence is passed to the image generation AI via the API, and the AI generates an appropriate image, and the image data is returned to the server. For example, "A scene of adventurers with stars shining in the night sky" generated using GAN is an example of this.

[0309] Step 4:

[0310] The server calls the text generation AI model according to the prompt sentence and uses natural language processing to generate a story text. Prompts such as "Create a story about children adventuring under the starry sky" are used, and the AI generates a narration along these lines, and text such as "That night, under the shining starry sky, the adventure began..." is output.

[0311] Step 5:

[0312] The server combines the generated image and text and records them as a digital book. In this step, the Python PIL library and ReportLab are used to arrange the image and text page by page and save them in HTML or PDF format. The output is an electronic book in file format and is in a state where the user can view it.

[0313] Step 6:

[0314] The server provides the terminal with a download link for the completed digital book. Users can download the picture book via this link and view it on their device. Furthermore, users can provide feedback to the server after use, which can be used in the next generation process. This feedback is accumulated to improve the system.

[0315] (Application Example 1)

[0316] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0317] In recent years, with the widespread adoption of digital devices, the demand for digital content that parents and children can enjoy together has increased. However, existing picture book-related content has difficulty reflecting the individual preferences of users, and there are limited means of providing interactive experiences such as creating stories together as a family. This invention aims to solve these problems by allowing users to customize the content of stories, supporting collaborative creation between parents and children, and making it instantly viewable on digital devices.

[0318] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0319] In this invention, the server includes means for receiving input from the user regarding the content of the document, the style of the images, the number of pages, and the message to be conveyed; means for making the generated electronic document viewable on a mobile information terminal; and means for presenting options for a digital version and a printed version of the generated document. This makes it possible for parents and children to collaboratively create original stories and instantly share and enjoy them on digital devices.

[0320] "Users" refers to individuals who use this system to customize documents and images and generate personalized digital content.

[0321] "Document content" refers to the text information of the story and explanations included in the generated digital picture book.

[0322] "Image style" refers to the artistic expression method of the generated visual content, and includes styles such as watercolor.

[0323] A "generative AI model" refers to an artificial intelligence system that automatically generates specific images and text based on user input.

[0324] A "prompt statement" refers to an input statement used to instruct a generative AI model on specific outputs.

[0325] "Personal information terminals" refer to computing devices that users can carry with them, such as smartphones and tablets.

[0326] The "feedback function" refers to a feature that collects user opinions and reactions to generated documents and uses them to improve the next generation process.

[0327] "Created jointly by parent and child" refers to the process where a parent and child work together to conceive the content and create a single digital piece of content.

[0328] This invention provides a specific embodiment of an interactive picture book generation system that allows parents and children to instantly enjoy stories they have collaboratively edited on a digital device.

[0329] The server displays a form on the terminal via a web interface to receive input information from the user, such as the content of the document, image format, number of pages, and the message to be conveyed. This information is sent to the server in JSON format. The server uses a generative AI model to dynamically generate images and text based on the received information. Specifically, prompt statements are used to instruct the AI ​​on what kind of images and text it should generate. For example, a prompt statement such as "Draw a watercolor-style picture of a beach and adventurers" might be used.

[0330] The terminal makes the generated electronic document viewable on a mobile device. This allows users to instantly enjoy the original story they created together as a family. Furthermore, the server presents users with the option of choosing between a digital version and a printed version of the generated document. Users can also obtain a physical printed picture book upon request.

[0331] The device also has a feedback function that sends comments and reactions to the generated document to the server. This feedback is reflected in the next story generation, allowing for content tailored to the child's interests and development. For example, when creating a picture book with a summer vacation adventure theme, the server generates images and text suitable for the document content, such as "adventurers visiting a seaside town."

[0332] In this way, the system of this invention makes it possible for parents and children to jointly generate customized digital content and provide enjoyable learning and experiences through it.

[0333] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0334] Step 1:

[0335] The terminal displays a web interface form for the user to input information such as story content, image style, number of pages, and the message to be conveyed. The entered information is organized as structured data and prepared for later transmission to the server. In this step, the input is information entered by the user, and the output is the organized data.

[0336] Step 2:

[0337] When a user presses the submit button on an input form, the terminal converts the prepared data into JSON format and sends it to the server. In this step, the input is prepared data, and the output is data encoded in JSON format.

[0338] Step 3:

[0339] The server receives JSON data sent from the terminal and analyzes it. Based on the received information, it creates the prompts necessary for the image generation AI and text generation AI. The input for this step is JSON data, and the output is prompts.

[0340] Step 4:

[0341] The server invokes a generative AI model, which generates an image in the specified style and constructs text according to the created prompt. The input for this step is the prompt, and the output is the generated image data and text data. Specifically, the image generation AI model receives instructions such as "Please draw a watercolor-style picture of a beach and adventurers" and generates an image.

[0342] Step 5:

[0343] The server combines the generated images and text to create the manuscript for the digital picture book. This manuscript is then formatted into a digital format such as HTML or PDF, making it viewable by users. The input for this step is image data and text data, and the output is the manuscript for the digital picture book.

[0344] Step 6:

[0345] The server generates and presents a link on the terminal that allows the user to download the completed digital picture book manuscript. The input for this step is the digital picture book manuscript, and the output is the download link.

[0346] Step 7:

[0347] The user uses a download link to save the e-book to their device and begins viewing it on their mobile device. In this step, the digital content, the e-book, is utilized at the user level.

[0348] Step 8:

[0349] After the user experiences the digital picture book, the device displays a feedback form, allowing them to enter their thoughts and reactions to the book. The entered feedback is sent to the server to be used in generating future stories. The input in this step is user feedback, and the output is feedback data.

[0350] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0351] This invention is a system that recognizes the user's emotions using an emotion engine and adjusts the content of the digital picture book based on that data to provide the most suitable picture book experience for each individual user. A specific embodiment of this system is described below.

[0352] First, the device provides the user with an input interface to customize the story content, image style, and number of pages. The user enters this information, and the data is sent to the server in JSON format.

[0353] The server analyzes the received data and generates AI prompts based on the input information. These generated prompts are passed to image and text generation AIs, which then generate materials based on the specified style and content. The generated illustrations and text are assembled into a story and saved in the form of a digital picture book.

[0354] Next, when a picture book is read aloud, the device activates an emotion engine. The emotion engine analyzes the user's, especially the child's, facial expressions and voice through the camera and microphone, and evaluates their current emotional state in real time. For example, if a child shows a surprised expression, the reaction can be visualized, and the next development of the story can be adjusted to increase engagement.

[0355] For example, if a child feels anxious during a tense scene in a story, the emotion engine can detect this and make adjustments such as inserting a softer, calmer scene on the next page.

[0356] Furthermore, the data accumulated by the emotion engine is stored on a server, and user preferences and reaction patterns are analyzed. This information is then fed back into future story generation, resulting in the creation of more personalized picture books.

[0357] Furthermore, user-customized digital picture books are provided via a downloadable link from the server, and if printed copies are desired as an option, production and delivery procedures are arranged.

[0358] This invention allows parents to read aloud in a way that resonates with their child's emotions, and allows children to enjoy stories tailored to their own needs.

[0359] The following describes the processing flow.

[0360] Step 1:

[0361] The device provides a web interface and displays a form where the user can input the story content, image style, number of pages, and the message they want to convey.

[0362] Step 2:

[0363] The user fills in the required information in the input form and completes the input. The data is then sent to the server in JSON format.

[0364] Step 3:

[0365] The server analyzes the received input data and creates prompts necessary for story generation. These prompts contain the information needed to generate both images and text.

[0366] Step 4:

[0367] The server invokes an image generation AI to generate an image in the specified style based on prompts. The generated image is temporarily stored in cloud storage.

[0368] Step 5:

[0369] The server invokes a text generation AI to generate sentences appropriate to the story from the prompts. This includes content tailored to a child's vocabulary level.

[0370] Step 6:

[0371] The server combines the generated images and text to create a digital picture book manuscript in HTML or PDF format. This manuscript includes the overall story structure and page layout.

[0372] Step 7:

[0373] The device activates an emotion engine and uses a camera and microphone to recognize the user's emotions in real time while reading picture books aloud.

[0374] Step 8:

[0375] The emotion engine detects emotional responses from the user's facial expressions and voice, and if, for example, a child is surprised, it receives data to adjust the content placed on the next page.

[0376] Step 9:

[0377] The server analyzes emotional data and adjusts the content of the picture book in real time based on the results. This provides a story experience tailored to the child's emotions.

[0378] Step 10:

[0379] The emotional data obtained by the emotion engine will be stored on a server and used to analyze the preferences of individual users in future picture book creation.

[0380] Step 11:

[0381] The server generates a link for the user to download the e-book and displays it on the device. If the user selects the printed book option, the server then handles information to arrange the shipping process.

[0382] (Example 2)

[0383] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0384] Existing digital picture books make it difficult to customize stories to suit the individual emotions and preferences of users. As a result, they cannot consistently provide users with an engaging and personalized storytelling experience, and there is a particular problem in that they cannot adequately deliver emotionally impactful read-aloud sessions, especially to children.

[0385] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0386] This invention includes a server that analyzes user input information to generate commands for a generation AI model and uses those commands to produce comprehensive story material; a server that analyzes the user's expressions and voice in real time and evaluates their emotions; and a server that dynamically adjusts the progression of the story according to the emotional evaluation to provide an adaptive electronic picture book experience. This makes it possible to dynamically adjust the story based on the user's individual emotions and preferences and provide an optimized picture book experience at all times.

[0387] "Users" refers to individual people who use the system to customize stories and receive an optimized picture book experience.

[0388] "The content of the story" refers to all elements related to the development of the story, such as the storyline, theme, and characters.

[0389] "Image style" refers to the artistic or design direction of the illustrations and visual materials used in a story.

[0390] "Page count" refers to the total number of pages that make up an electronic picture book, and is an element that indicates the length or volume of the story.

[0391] A "generative AI model" refers to artificial intelligence technology that automatically generates images and text based on input prompts.

[0392] A "prompt statement" is a text containing instructions given to a generative AI model, specifically a statement that directs the model to generate concrete content.

[0393] "Emotional assessment" refers to the act of analyzing and evaluating a user's current emotional state in real time based on their facial expressions and voice.

[0394] An "adaptive digital picture book experience" refers to a function that dynamically adjusts the story's progression and content according to the user's emotions and preferences, providing a personalized read-aloud experience.

[0395] This invention is a system that allows users to customize digital picture books based on their own emotions and preferences. First, the terminal provides the user with an intuitive interface where they can input the story content, image style, number of pages, etc. This allows the user to select specific keywords and create a customized story outline based on them.

[0396] The server receives the input data and analyzes it using programming languages ​​such as Python or JavaScript. From this analysis, it creates prompts to supply to the generating AI model. These prompts are provided in text format and might include something like, "Generate a space adventure story for children, focusing on friendship with colorful illustrations." Based on these prompts, AI technologies such as Stable Diffusion for image generation and the GPT series for text generation are used to specifically generate the story and illustrations.

[0397] The server integrates the generated illustrations and text to construct the manuscript for the digital picture book. This results in a digital work with narrative coherence and visual appeal.

[0398] Furthermore, during story time, the device uses its camera and microphone to transmit the user's facial expressions and voice to the emotion engine in real time. The emotion engine analyzes this data and evaluates the user's emotional state, dynamically adjusting the story's progression. For example, if a child shows a surprised expression, the story's development shifts towards a more reassuring direction to increase engagement.

[0399] The generated digital picture books are provided as downloadable links via a server, and printed and delivered versions are available upon request. In this way, users can enjoy a unique storytelling experience tailored to their individual emotions.

[0400] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0401] Step 1:

[0402] The terminal provides the user with an input interface. Here, the user can select or input information such as the story content, image style, and number of pages. This input information is collected as data in JSON format. Specifically, data is acquired using touchscreen or voice recognition technology.

[0403] Step 2:

[0404] The server receives JSON data sent from the terminal and performs analysis. A Python script is used for data analysis to extract necessary information and generate prompts required for the AI ​​model. At this stage, the data is structured, and specific keywords such as "space exploration" are extracted.

[0405] Step 3:

[0406] The server creates instructions for text and image generation AI based on the generated prompt text. These instructions are then passed to the generation AI model (e.g., Stable Diffusion or GPT series) to generate illustrations and narrative text based on the specified style and theme. Using the AI ​​prompt as input, the system obtains theme-appropriate images and text as output.

[0407] Step 4:

[0408] The server integrates the generated illustrations and text to create an electronic picture book. This creation utilizes programmatic digital layout, carefully considering the story's order and page count. As a result, a visually engaging and consistent work is produced.

[0409] Step 5:

[0410] The device uses a camera and microphone to capture the user's facial expressions and voice while a picture book is being read aloud, and transmits this data to the emotion engine. Real-time video and audio data are used as input for emotion analysis. Specifically, it performs facial recognition and voice analysis.

[0411] Step 6:

[0412] The server receives the results of emotion analysis and adjusts the story's progression accordingly. For example, if the user feels anxious, a calming scene is automatically inserted on the next page. This dynamic adjustment provides an optimal story experience that resonates with the user's emotions.

[0413] Step 7:

[0414] The server then provides the user with a customized digital picture book and generates a downloadable link. If the user wishes to have the book printed or delivered, the server handles those procedures. The final output is the individual picture book data, which is shared with the user.

[0415] (Application Example 2)

[0416] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0417] There is a challenge in providing individual users, especially children, with an appropriate and enjoyable story experience tailored to their emotions at any given time when enjoying ebooks. Furthermore, static content makes it difficult to fully personalize the user experience, resulting in a lack of increased user engagement.

[0418] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0419] In this invention, the server includes means for receiving input from the user regarding the story content, image style, number of pages, and message to be conveyed; means for creating an ebook manuscript by combining the generated images and text; and means for sensing the user's emotional state and dynamically adjusting the story content accordingly. This makes it possible to provide a personalized story experience that responds to the user's real-time emotions.

[0420] "User" refers to an individual or organization that uses the system to generate and customize stories.

[0421] A "story" refers to a series of texts and images combined to create an experience for the user, such as a narrative or episodes.

[0422] "Images" refer to the visual elements included in an e-book, specifically illustrations and diagrams generated to visually represent the story's content and style.

[0423] "Style" refers to the elements that characterize the appearance and design of a story or image, and it is the criterion that users use to select a particular theme or atmosphere.

[0424] "Page count" refers to the total number of pages in which the story unfolds within an ebook, and is an indicator used by users to determine the length of the story.

[0425] A "message" refers to the intention or meaning conveyed to the user through a story, and may include specific values ​​or lessons.

[0426] "Instructions" refer to data containing specific commands and settings necessary for generating a story, and are generated by the system.

[0427] "E-books" refer to a form of book that is stored in digital format and can be viewed by users through electronic devices.

[0428] "Emotional state" refers to data that indicates a user's psychological or emotional response, and includes information collected from the user's facial expressions and voice while they are browsing.

[0429] "Customization" refers to the process of adjusting the content and format of a story to suit the specific needs and preferences of a particular user.

[0430] This invention is an e-book system that dynamically adjusts the content of a story according to the user's emotional state. When a user inputs the story content, image style, number of pages, and message they want to convey into their device, this information is sent to a server. Based on this input information, the server creates the necessary instructions for story generation and generates prompt sentences using a generation AI model. These prompt sentences trigger image and text generation, creating the material for the story.

[0431] The device uses a camera and microphone to collect the user's facial expressions and voice in real time, and an emotion engine recognizes their emotional state. For example, OpenCV captures camera footage, and the microphone library acquires audio data. The server takes this emotional state into consideration and dynamically adjusts the content of the generated story. Depending on the emotional change, the next scene can be switched to content that provides a sense of security.

[0432] For example, if a user feels anxious during a tense scene in a story, a friendly character who alleviates the tension can be introduced in the next scene, creating a reassuring development. This allows users to receive a personalized experience tailored to their individual circumstances.

[0433] An example of a prompt to input into a generation AI model is, "The user is currently feeling anxious. Please take this emotion into consideration and generate a storyline in the next scene that will provide a sense of reassurance." By providing stories that resonate with the user in this way, the enjoyment of ebooks can be maximized.

[0434] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0435] Step 1:

[0436] The user inputs the story content, image style, number of pages, and message via their device. The entered data is temporarily stored on the device and converted to JSON format. This is the stage where the user's intentions and requests are concretized.

[0437] Step 2:

[0438] The terminal sends the converted JSON data to the server. The server parses the received data and creates the instructions necessary for story generation. Through the parsing process, the server extracts each element and reformats it into the appropriate format. The output is a prompt sentence to be input to the generating AI model.

[0439] Step 3:

[0440] The server inputs prompt text into the generative AI model, which then generates image and text materials. The generative AI model performs data calculations based on the instructions and constructs illustrations and text to match the settings. The generated materials are obtained as output.

[0441] Step 4:

[0442] The generated materials are combined on the server to create the manuscript for the ebook. The server integrates the materials to construct the overall story and compiles it into the final digital format. The output is the data of the completed ebook.

[0443] Step 5:

[0444] The device collects the user's facial expressions and voice using a camera and microphone, and analyzes them with an emotion engine. It captures video using OpenCV and records audio using the microphone library. The input consists of audio and video data, and the analysis yields the user's emotional state.

[0445] Step 6:

[0446] The server adjusts the content of the ebook based on the emotional state it receives. Using the results of the emotion engine, it replaces or rearranges text and images. As a result, the output is a story that corresponds to the user's emotions.

[0447] Step 7:

[0448] The adjusted ebook is returned to the user's device and becomes available for download. The final ebook data is sent to the device, and the user can use it for viewing or reading aloud. The output is a usable ebook.

[0449] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0450] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0451] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0452] [Third Embodiment]

[0453] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0454] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0455] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0456] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0457] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0458] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0459] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0460] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0461] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0462] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0463] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0464] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0465] The system of the present invention automatically generates picture books via a web interface. Specific embodiments are described below.

[0466] First, the device displays an input form for the user to customize the content of the picture book. This form can accept input such as the story content, image style, number of pages, message to be conveyed, and, if necessary, the names of the characters.

[0467] Users enter this information based on their preferences. Once the input is complete, this information is sent to the server in JSON format.

[0468] The server analyzes the received information and creates the necessary instructions for generating images and text. These instructions are then passed to the image generation AI and text generation AI, which carry out the specific generation processes.

[0469] The server invokes an image generation AI based on prompts and generates an image in the specified style. For example, an illustration of a starry sky in a watercolor style is generated. Additionally, a text generation AI is used to generate narration that takes vocabulary level into consideration. For example, a sentence like, "That night, under countless stars, the adventurers..." is generated.

[0470] The generated images and text are combined by the server to create a picture book manuscript in HTML or PDF format. This harmonizes the images and text, resulting in a digitally completed, original picture book based on the user's specified content.

[0471] Ultimately, the server makes this digital picture book available for download on devices, providing users with free access to it. Users can view this picture book on digital devices such as tablets and read it aloud to their children, and they can also optionally order a printed version.

[0472] Furthermore, feedback and reactions to the generated picture books are sent back to the server using a dedicated input device or stuffed animal. Based on this feedback, it becomes possible to provide content tailored to the child's interests and development when creating the next picture book.

[0473] This system allows parents to easily create customized picture books, providing their children with an enjoyable learning experience.

[0474] The following describes the processing flow.

[0475] Step 1:

[0476] The device provides the user with a web interface, displaying a form to enter the picture book's content, image style, number of pages, message to be conveyed, and, if necessary, the names of the characters.

[0477] Step 2:

[0478] The user enters the required information into the input form mentioned above and submits the completed data. The input content is sent to the server in JSON format.

[0479] Step 3:

[0480] The server parses the received JSON data and generates an AI prompt based on the input. This prompt contains the necessary instructions for both the image generation AI and the text generation AI.

[0481] Step 4:

[0482] The server invokes an image generation AI and generates an image in the specified style according to the prompts. The generated image is temporarily stored in cloud storage.

[0483] Step 5:

[0484] The server invokes a text generation AI to generate sentences suitable for the story based on prompts. The generated text is checked to ensure it is easy to understand and matches the child's vocabulary level.

[0485] Step 6:

[0486] The server combines the generated images and text to create a digital picture book manuscript in HTML or PDF format. It then arranges the layout page by page to complete the final manuscript.

[0487] Step 7:

[0488] The server generates a download link for the completed digital picture book and provides it to the user's device. The user can then download the picture book from this link and view it on a tablet or other device.

[0489] Step 8:

[0490] If the user selects the bookbinding option, they will enter the additional information required for bookbinding and delivery on their terminal and send that information to the server. The server will then arrange the bookbinding process and delivery.

[0491] Step 9:

[0492] The system records children's thoughts and reactions to the generated picture books via devices or stuffed animals, and sends this feedback to a server. This information is then considered when creating future picture books, and used to provide content that is appropriate for the child's development.

[0493] (Example 1)

[0494] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0495] Traditional picture book creation systems had problems such as insufficient customization options and inadequate incorporation of feedback on the generated products. Furthermore, the options for providing the generated digital books in physical form were limited. As a result, it was difficult for users to create picture books that met their individual needs and to make subsequent improvements.

[0496] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0497] In this invention, the server includes means for receiving input from the user regarding the story content, image style, number of pages, and the message to be conveyed; means for converting the input information into a data format and analyzing it; and means for creating instructions for an image generation AI model and a text generation AI model based on the data analysis. This enables the creation of customizable picture books that meet the diverse needs of users and allows for continuous quality improvement through feedback on the generated content.

[0498] "User" refers to an individual or user who uses the system to input customized information for stories and images and receives the generated content.

[0499] A "server" is an information processing device that receives and analyzes input data, sends commands to the generating AI model, and stores and provides the generated content.

[0500] "Data format" refers to a format that structures information and converts it into a form that can be easily parsed, and generally includes formats such as JSON and XML.

[0501] "Data analysis" is the process of processing input information and extracting useful commands and patterns.

[0502] An "image generation AI model" refers to an algorithm that generates images in an artistic or requested style based on given instructions.

[0503] A "text generation AI model" refers to an algorithm that generates natural language stories and explanatory texts based on given instructions.

[0504] A "command" refers to a command or prompt that instructs an AI model to generate content in a specific form or with specific content.

[0505] A "digital book" refers to a book in electronic format that is accessible on a computer or electronic device, and is presented as multimedia content including text and images.

[0506] "Feedback" refers to evaluations and opinions provided by users, and is information used to improve the system and for future updates.

[0507] A "physical book" refers to a product in book form created by printing digitally generated content and using paper or other physical materials.

[0508] This invention is a system that automatically generates picture books customized by users via a web interface. The system mainly consists of a server, a terminal, and a generation AI model.

[0509] The device displays an input form to the user, which is necessary for customizing the story. This form is used to input the story's content, image style, number of pages, message to convey, character names, and other information. After the user enters this information, the device uses JavaScript and HTML to convert the data into JSON format and send it to the server.

[0510] The server analyzes the received data and creates instructions to pass to the image generation AI model and the text generation AI model. A Python script is used for this analysis, and the instructions are formed as specific prompts. For example, the image generation AI is given a prompt such as "Generate a watercolor-style adventure scene with a starry sky background," and the text generation AI is given a prompt such as "Create a story about children adventuring under a starry sky."

[0511] The server then uses this prompt to call an image generation AI model to generate an image in the specified style, and a text generation AI model to generate narrative text that takes vocabulary and expression into consideration. The specific processing utilizes APIs for the generation AI. For example, a Generative Adversarial Network (GAN) is used for image generation, and a natural language processing model such as GPT-3 is used for text generation.

[0512] The generated images and text are combined by the server in HTML or PDF format and saved as a digital book. The Python PIL library and ReportLab are used in this process.

[0513] Finally, the server provides the terminal with a download link for the finished product. Users can use the download link to view the picture book on their digital device and, if necessary, send feedback back to the server through a dedicated form. This feedback will be considered in the next generation process and used to improve the system.

[0514] This system allows users to easily customize and create picture books tailored to their individual needs, providing an experience that enhances learning effectiveness, especially for children.

[0515] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0516] Step 1:

[0517] The device displays a customization input form for the story to the user. Here, the user can enter details such as the story's content, image style, number of pages, message, and character names. This form is implemented using HTML and JavaScript. Upon completion, the device converts the input data into JSON format and sends it to the server.

[0518] Step 2:

[0519] The server parses the JSON data received from the terminal. This parsing process uses Python, and based on the provided data, it creates instructions for the image generation AI model and the text generation AI model. Specifically, the data processing involves parsing user input to generate prompt sentences, such as "Generate a watercolor-style adventure scene with a starry sky background."

[0520] Step 3:

[0521] The server calls an image generation AI model based on the generated prompt text and generates an image in the specified style. In this step, the prompt text is passed to the image generation AI via API, the AI ​​generates an appropriate image, and the image data is returned to the server. For example, a "scene of an adventurer with stars shining in the night sky" generated using GAN is one such example.

[0522] Step 4:

[0523] The server invokes a text generation AI model according to a prompt and generates narrative text using natural language processing. A prompt such as "Create a story about children adventuring under the starry sky" is used, and the AI ​​generates narration accordingly, outputting text such as "That night, under the shining starry sky, the adventure began..."

[0524] Step 5:

[0525] The server combines the generated images and text and records them as a digital book. This step uses Python's PIL library or ReportLab to arrange the images and text page by page and save them in HTML or PDF format. The output is an ebook file, ready for user viewing.

[0526] Step 6:

[0527] The server provides the terminal with a download link for the completed digital book. Users can download the picture book via this link and view it on their device. Furthermore, users can provide feedback to the server after use, which can be used in the next generation process. This feedback is accumulated to improve the system.

[0528] (Application Example 1)

[0529] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0530] In recent years, with the widespread adoption of digital devices, the demand for digital content that parents and children can enjoy together has increased. However, existing picture book-related content has difficulty reflecting the individual preferences of users, and there are limited means of providing interactive experiences such as creating stories together as a family. This invention aims to solve these problems by allowing users to customize the content of stories, supporting collaborative creation between parents and children, and making it instantly viewable on digital devices.

[0531] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0532] In this invention, the server includes means for receiving input from the user regarding the content of the document, the style of the images, the number of pages, and the message to be conveyed; means for making the generated electronic document viewable on a mobile information terminal; and means for presenting options for a digital version and a printed version of the generated document. This makes it possible for parents and children to collaboratively create original stories and instantly share and enjoy them on digital devices.

[0533] "Users" refers to individuals who use this system to customize documents and images and generate personalized digital content.

[0534] "Document content" refers to the text information of the story and explanations included in the generated digital picture book.

[0535] "Image style" refers to the artistic expression method of the generated visual content, and includes styles such as watercolor.

[0536] A "generative AI model" refers to an artificial intelligence system that automatically generates specific images and text based on user input.

[0537] A "prompt statement" refers to an input statement used to instruct a generative AI model on specific outputs.

[0538] "Personal information terminals" refer to computing devices that users can carry with them, such as smartphones and tablets.

[0539] The "feedback function" refers to a feature that collects user opinions and reactions to generated documents and uses them to improve the next generation process.

[0540] "Created jointly by parent and child" refers to the process where a parent and child work together to conceive the content and create a single digital piece of content.

[0541] This invention provides a specific embodiment of an interactive picture book generation system that allows parents and children to instantly enjoy stories they have collaboratively edited on a digital device.

[0542] The server displays a form on the terminal via a web interface to receive input information from the user, such as the content of the document, image format, number of pages, and the message to be conveyed. This information is sent to the server in JSON format. The server uses a generative AI model to dynamically generate images and text based on the received information. Specifically, prompt statements are used to instruct the AI ​​on what kind of images and text it should generate. For example, a prompt statement such as "Draw a watercolor-style picture of a beach and adventurers" might be used.

[0543] The terminal makes the generated electronic document viewable on a mobile device. This allows users to instantly enjoy the original story they created together as a family. Furthermore, the server presents users with the option of choosing between a digital version and a printed version of the generated document. Users can also obtain a physical printed picture book upon request.

[0544] The device also has a feedback function that sends comments and reactions to the generated document to the server. This feedback is reflected in the next story generation, allowing for content tailored to the child's interests and development. For example, when creating a picture book with a summer vacation adventure theme, the server generates images and text suitable for the document content, such as "adventurers visiting a seaside town."

[0545] In this way, the system of this invention makes it possible for parents and children to jointly generate customized digital content and provide enjoyable learning and experiences through it.

[0546] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0547] Step 1:

[0548] The terminal displays a web interface form for the user to input information such as story content, image style, number of pages, and the message to be conveyed. The entered information is organized as structured data and prepared for later transmission to the server. In this step, the input is information entered by the user, and the output is the organized data.

[0549] Step 2:

[0550] When a user presses the submit button on an input form, the terminal converts the prepared data into JSON format and sends it to the server. In this step, the input is prepared data, and the output is data encoded in JSON format.

[0551] Step 3:

[0552] The server receives JSON data sent from the terminal and analyzes it. Based on the received information, it creates the prompts necessary for the image generation AI and text generation AI. The input for this step is JSON data, and the output is prompts.

[0553] Step 4:

[0554] The server invokes a generative AI model, which generates an image in the specified style and constructs text according to the created prompt. The input for this step is the prompt, and the output is the generated image data and text data. Specifically, the image generation AI model receives instructions such as "Please draw a watercolor-style picture of a beach and adventurers" and generates an image.

[0555] Step 5:

[0556] The server combines the generated images and text to create the manuscript for the digital picture book. This manuscript is then formatted into a digital format such as HTML or PDF, making it viewable by users. The input for this step is image data and text data, and the output is the manuscript for the digital picture book.

[0557] Step 6:

[0558] The server generates and presents a link on the terminal that allows the user to download the completed digital picture book manuscript. The input for this step is the digital picture book manuscript, and the output is the download link.

[0559] Step 7:

[0560] The user uses a download link to save the e-book to their device and begins viewing it on their mobile device. In this step, the digital content, the e-book, is utilized at the user level.

[0561] Step 8:

[0562] After the user experiences the digital picture book, the device displays a feedback form, allowing them to enter their thoughts and reactions to the book. The entered feedback is sent to the server to be used in generating future stories. The input in this step is user feedback, and the output is feedback data.

[0563] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0564] This invention is a system that recognizes the user's emotions using an emotion engine and adjusts the content of the digital picture book based on that data to provide the most suitable picture book experience for each individual user. A specific embodiment of this system is described below.

[0565] First, the device provides the user with an input interface to customize the story content, image style, and number of pages. The user enters this information, and the data is sent to the server in JSON format.

[0566] The server analyzes the received data and generates AI prompts based on the input information. These generated prompts are passed to image and text generation AIs, which then generate materials based on the specified style and content. The generated illustrations and text are assembled into a story and saved in the form of a digital picture book.

[0567] Next, when a picture book is read aloud, the device activates an emotion engine. The emotion engine analyzes the user's, especially the child's, facial expressions and voice through the camera and microphone, and evaluates their current emotional state in real time. For example, if a child shows a surprised expression, the reaction can be visualized, and the next development of the story can be adjusted to increase engagement.

[0568] For example, if a child feels anxious during a tense scene in a story, the emotion engine can detect this and make adjustments such as inserting a softer, calmer scene on the next page.

[0569] Furthermore, the data accumulated by the emotion engine is stored on a server, and user preferences and reaction patterns are analyzed. This information is then fed back into future story generation, resulting in the creation of more personalized picture books.

[0570] Furthermore, user-customized digital picture books are provided via a downloadable link from the server, and if printed copies are desired as an option, production and delivery procedures are arranged.

[0571] This invention allows parents to read aloud in a way that resonates with their child's emotions, and allows children to enjoy stories tailored to their own needs.

[0572] The following describes the processing flow.

[0573] Step 1:

[0574] The device provides a web interface and displays a form where the user can input the story content, image style, number of pages, and the message they want to convey.

[0575] Step 2:

[0576] The user fills in the required information in the input form and completes the input. The data is then sent to the server in JSON format.

[0577] Step 3:

[0578] The server analyzes the received input data and creates prompts necessary for story generation. These prompts contain the information needed to generate both images and text.

[0579] Step 4:

[0580] The server invokes an image generation AI to generate an image in the specified style based on prompts. The generated image is temporarily stored in cloud storage.

[0581] Step 5:

[0582] The server invokes a text generation AI to generate sentences appropriate to the story from the prompts. This includes content tailored to a child's vocabulary level.

[0583] Step 6:

[0584] The server combines the generated images and text to create a digital picture book manuscript in HTML or PDF format. This manuscript includes the overall story structure and page layout.

[0585] Step 7:

[0586] The device activates an emotion engine and uses a camera and microphone to recognize the user's emotions in real time while reading picture books aloud.

[0587] Step 8:

[0588] The emotion engine detects emotional responses from the user's facial expressions and voice, and if, for example, a child is surprised, it receives data to adjust the content placed on the next page.

[0589] Step 9:

[0590] The server analyzes emotional data and adjusts the content of the picture book in real time based on the results. This provides a story experience tailored to the child's emotions.

[0591] Step 10:

[0592] The emotional data obtained by the emotion engine will be stored on a server and used to analyze the preferences of individual users in future picture book creation.

[0593] Step 11:

[0594] The server generates a link for the user to download the e-book and displays it on the device. If the user selects the printed book option, the server then handles information to arrange the shipping process.

[0595] (Example 2)

[0596] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0597] Existing digital picture books make it difficult to customize stories to suit the individual emotions and preferences of users. As a result, they cannot consistently provide users with an engaging and personalized storytelling experience, and there is a particular problem in that they cannot adequately deliver emotionally impactful read-aloud sessions, especially to children.

[0598] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0599] This invention includes a server that analyzes user input information to generate commands for a generation AI model and uses those commands to produce comprehensive story material; a server that analyzes the user's expressions and voice in real time and evaluates their emotions; and a server that dynamically adjusts the progression of the story according to the emotional evaluation to provide an adaptive electronic picture book experience. This makes it possible to dynamically adjust the story based on the user's individual emotions and preferences and provide an optimized picture book experience at all times.

[0600] "Users" refers to individual people who use the system to customize stories and receive an optimized picture book experience.

[0601] "The content of the story" refers to all elements related to the development of the story, such as the storyline, theme, and characters.

[0602] "Image style" refers to the artistic or design direction of the illustrations and visual materials used in a story.

[0603] "Page count" refers to the total number of pages that make up an electronic picture book, and is an element that indicates the length or volume of the story.

[0604] A "generative AI model" refers to artificial intelligence technology that automatically generates images and text based on input prompts.

[0605] A "prompt statement" is a text containing instructions given to a generative AI model, specifically a statement that directs the model to generate concrete content.

[0606] "Emotional assessment" refers to the act of analyzing and evaluating a user's current emotional state in real time based on their facial expressions and voice.

[0607] An "adaptive digital picture book experience" refers to a function that dynamically adjusts the story's progression and content according to the user's emotions and preferences, providing a personalized read-aloud experience.

[0608] This invention is a system that allows users to customize digital picture books based on their own emotions and preferences. First, the terminal provides the user with an intuitive interface where they can input the story content, image style, number of pages, etc. This allows the user to select specific keywords and create a customized story outline based on them.

[0609] The server receives the input data and analyzes it using programming languages ​​such as Python or JavaScript. From this analysis, it creates prompts to supply to the generating AI model. These prompts are provided in text format and might include something like, "Generate a space adventure story for children, focusing on friendship with colorful illustrations." Based on these prompts, AI technologies such as Stable Diffusion for image generation and the GPT series for text generation are used to specifically generate the story and illustrations.

[0610] The server integrates the generated illustrations and text to construct the manuscript for the digital picture book. This results in a digital work with narrative coherence and visual appeal.

[0611] Furthermore, during story time, the device uses its camera and microphone to transmit the user's facial expressions and voice to the emotion engine in real time. The emotion engine analyzes this data and evaluates the user's emotional state, dynamically adjusting the story's progression. For example, if a child shows a surprised expression, the story's development shifts towards a more reassuring direction to increase engagement.

[0612] The generated digital picture books are provided as downloadable links via a server, and printed and delivered versions are available upon request. In this way, users can enjoy a unique storytelling experience tailored to their individual emotions.

[0613] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0614] Step 1:

[0615] The terminal provides the user with an input interface. Here, the user can select or input information such as the story content, image style, and number of pages. This input information is collected as data in JSON format. Specifically, data is acquired using touchscreen or voice recognition technology.

[0616] Step 2:

[0617] The server receives JSON data sent from the terminal and performs analysis. A Python script is used for data analysis to extract necessary information and generate prompts required for the AI ​​model. At this stage, the data is structured, and specific keywords such as "space exploration" are extracted.

[0618] Step 3:

[0619] The server creates instructions for text and image generation AI based on the generated prompt text. These instructions are then passed to the generation AI model (e.g., Stable Diffusion or GPT series) to generate illustrations and narrative text based on the specified style and theme. Using the AI ​​prompt as input, the system obtains theme-appropriate images and text as output.

[0620] Step 4:

[0621] The server integrates the generated illustrations and text to create an electronic picture book. This creation utilizes programmatic digital layout, carefully considering the story's order and page count. As a result, a visually engaging and consistent work is produced.

[0622] Step 5:

[0623] The device uses a camera and microphone to capture the user's facial expressions and voice while a picture book is being read aloud, and transmits this data to the emotion engine. Real-time video and audio data are used as input for emotion analysis. Specifically, it performs facial recognition and voice analysis.

[0624] Step 6:

[0625] The server receives the results of emotion analysis and adjusts the story's progression accordingly. For example, if the user feels anxious, a calming scene is automatically inserted on the next page. This dynamic adjustment provides an optimal story experience that resonates with the user's emotions.

[0626] Step 7:

[0627] The server then provides the user with a customized digital picture book and generates a downloadable link. If the user wishes to have the book printed or delivered, the server handles those procedures. The final output is the individual picture book data, which is shared with the user.

[0628] (Application Example 2)

[0629] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0630] There is a challenge in providing individual users, especially children, with an appropriate and enjoyable story experience tailored to their emotions at any given time when enjoying ebooks. Furthermore, static content makes it difficult to fully personalize the user experience, resulting in a lack of increased user engagement.

[0631] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0632] In this invention, the server includes means for receiving input from the user regarding the story content, image style, number of pages, and message to be conveyed; means for creating an ebook manuscript by combining the generated images and text; and means for sensing the user's emotional state and dynamically adjusting the story content accordingly. This makes it possible to provide a personalized story experience that responds to the user's real-time emotions.

[0633] "User" refers to an individual or organization that uses the system to generate and customize stories.

[0634] A "story" refers to a series of texts and images combined to create an experience for the user, such as a narrative or episodes.

[0635] "Images" refer to the visual elements included in an e-book, specifically illustrations and diagrams generated to visually represent the story's content and style.

[0636] "Style" refers to the elements that characterize the appearance and design of a story or image, and it is the criterion that users use to select a particular theme or atmosphere.

[0637] "Page count" refers to the total number of pages in which the story unfolds within an ebook, and is an indicator used by users to determine the length of the story.

[0638] A "message" refers to the intention or meaning conveyed to the user through a story, and may include specific values ​​or lessons.

[0639] "Instructions" refer to data containing specific commands and settings necessary for generating a story, and are generated by the system.

[0640] "E-books" refer to a form of book that is stored in digital format and can be viewed by users through electronic devices.

[0641] "Emotional state" refers to data that indicates a user's psychological or emotional response, and includes information collected from the user's facial expressions and voice while they are browsing.

[0642] "Customization" refers to the process of adjusting the content and format of a story to suit the specific needs and preferences of a particular user.

[0643] This invention is an e-book system that dynamically adjusts the content of a story according to the user's emotional state. When a user inputs the story content, image style, number of pages, and message they want to convey into their device, this information is sent to a server. Based on this input information, the server creates the necessary instructions for story generation and generates prompt sentences using a generation AI model. These prompt sentences trigger image and text generation, creating the material for the story.

[0644] The device uses a camera and microphone to collect the user's facial expressions and voice in real time, and an emotion engine recognizes their emotional state. For example, OpenCV captures camera footage, and the microphone library acquires audio data. The server takes this emotional state into consideration and dynamically adjusts the content of the generated story. Depending on the emotional change, the next scene can be switched to content that provides a sense of security.

[0645] For example, if a user feels anxious during a tense scene in a story, a friendly character who alleviates the tension can be introduced in the next scene, creating a reassuring development. This allows users to receive a personalized experience tailored to their individual circumstances.

[0646] An example of a prompt to input into a generation AI model is, "The user is currently feeling anxious. Please take this emotion into consideration and generate a storyline in the next scene that will provide a sense of reassurance." By providing stories that resonate with the user in this way, the enjoyment of ebooks can be maximized.

[0647] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0648] Step 1:

[0649] The user inputs the story content, image style, number of pages, and message via their device. The entered data is temporarily stored on the device and converted to JSON format. This is the stage where the user's intentions and requests are concretized.

[0650] Step 2:

[0651] The terminal sends the converted JSON data to the server. The server parses the received data and creates the instructions necessary for story generation. Through the parsing process, the server extracts each element and reformats it into the appropriate format. The output is a prompt sentence to be input to the generating AI model.

[0652] Step 3:

[0653] The server inputs prompt text into the generative AI model, which then generates image and text materials. The generative AI model performs data calculations based on the instructions and constructs illustrations and text to match the settings. The generated materials are obtained as output.

[0654] Step 4:

[0655] The generated materials are combined on the server to create the manuscript for the ebook. The server integrates the materials to construct the overall story and compiles it into the final digital format. The output is the data of the completed ebook.

[0656] Step 5:

[0657] The device collects the user's facial expressions and voice using a camera and microphone, and analyzes them with an emotion engine. It captures video using OpenCV and records audio using the microphone library. The input consists of audio and video data, and the analysis yields the user's emotional state.

[0658] Step 6:

[0659] The server adjusts the content of the ebook based on the emotional state it receives. Using the results of the emotion engine, it replaces or rearranges text and images. As a result, the output is a story that corresponds to the user's emotions.

[0660] Step 7:

[0661] The adjusted ebook is returned to the user's device and becomes available for download. The final ebook data is sent to the device, and the user can use it for viewing or reading aloud. The output is a usable ebook.

[0662] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0663] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0664] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0665] [Fourth Embodiment]

[0666] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0667] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0668] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0669] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0670] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0671] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0672] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0673] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0674] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0675] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0676] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0677] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0678] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0679] The system of the present invention automatically generates picture books via a web interface. Specific embodiments are described below.

[0680] First, the device displays an input form for the user to customize the content of the picture book. This form can accept input such as the story content, image style, number of pages, message to be conveyed, and, if necessary, the names of the characters.

[0681] Users enter this information based on their preferences. Once the input is complete, this information is sent to the server in JSON format.

[0682] The server analyzes the received information and creates the necessary instructions for generating images and text. These instructions are then passed to the image generation AI and text generation AI, which carry out the specific generation processes.

[0683] The server invokes an image generation AI based on prompts and generates an image in the specified style. For example, an illustration of a starry sky in a watercolor style is generated. Additionally, a text generation AI is used to generate narration that takes vocabulary level into consideration. For example, a sentence like, "That night, under countless stars, the adventurers..." is generated.

[0684] The generated images and text are combined by the server to create a picture book manuscript in HTML or PDF format. This harmonizes the images and text, resulting in a digitally completed, original picture book based on the user's specified content.

[0685] Ultimately, the server makes this digital picture book available for download on devices, providing users with free access to it. Users can view this picture book on digital devices such as tablets and read it aloud to their children, and they can also optionally order a printed version.

[0686] Furthermore, feedback and reactions to the generated picture books are sent back to the server using a dedicated input device or stuffed animal. Based on this feedback, it becomes possible to provide content tailored to the child's interests and development when creating the next picture book.

[0687] This system allows parents to easily create customized picture books, providing their children with an enjoyable learning experience.

[0688] The following describes the processing flow.

[0689] Step 1:

[0690] The device provides the user with a web interface, displaying a form to enter the picture book's content, image style, number of pages, message to be conveyed, and, if necessary, the names of the characters.

[0691] Step 2:

[0692] The user enters the required information into the input form mentioned above and submits the completed data. The input content is sent to the server in JSON format.

[0693] Step 3:

[0694] The server parses the received JSON data and generates an AI prompt based on the input. This prompt contains the necessary instructions for both the image generation AI and the text generation AI.

[0695] Step 4:

[0696] The server invokes an image generation AI and generates an image in the specified style according to the prompts. The generated image is temporarily stored in cloud storage.

[0697] Step 5:

[0698] The server invokes a text generation AI to generate sentences suitable for the story based on prompts. The generated text is checked to ensure it is easy to understand and matches the child's vocabulary level.

[0699] Step 6:

[0700] The server combines the generated images and text to create a digital picture book manuscript in HTML or PDF format. It then arranges the layout page by page to complete the final manuscript.

[0701] Step 7:

[0702] The server generates a download link for the completed digital picture book and provides it to the user's device. The user can then download the picture book from this link and view it on a tablet or other device.

[0703] Step 8:

[0704] If the user selects the bookbinding option, they will enter the additional information required for bookbinding and delivery on their terminal and send that information to the server. The server will then arrange the bookbinding process and delivery.

[0705] Step 9:

[0706] The system records children's thoughts and reactions to the generated picture books via devices or stuffed animals, and sends this feedback to a server. This information is then considered when creating future picture books, and used to provide content that is appropriate for the child's development.

[0707] (Example 1)

[0708] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0709] Traditional picture book creation systems had problems such as insufficient customization options and inadequate incorporation of feedback on the generated products. Furthermore, the options for providing the generated digital books in physical form were limited. As a result, it was difficult for users to create picture books that met their individual needs and to make subsequent improvements.

[0710] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0711] In this invention, the server includes means for receiving input from the user regarding the story content, image style, number of pages, and the message to be conveyed; means for converting the input information into a data format and analyzing it; and means for creating instructions for an image generation AI model and a text generation AI model based on the data analysis. This enables the creation of customizable picture books that meet the diverse needs of users and allows for continuous quality improvement through feedback on the generated content.

[0712] "User" refers to an individual or user who uses the system to input customized information for stories and images and receives the generated content.

[0713] A "server" is an information processing device that receives and analyzes input data, sends commands to the generating AI model, and stores and provides the generated content.

[0714] "Data format" refers to a format that structures information and converts it into a form that can be easily parsed, and generally includes formats such as JSON and XML.

[0715] "Data analysis" is the process of processing input information and extracting useful commands and patterns.

[0716] An "image generation AI model" refers to an algorithm that generates images in an artistic or requested style based on given instructions.

[0717] A "text generation AI model" refers to an algorithm that generates natural language stories and explanatory texts based on given instructions.

[0718] A "command" refers to a command or prompt that instructs an AI model to generate content in a specific form or with specific content.

[0719] A "digital book" refers to a book in electronic format that is accessible on a computer or electronic device, and is presented as multimedia content including text and images.

[0720] "Feedback" refers to evaluations and opinions provided by users, and is information used to improve the system and for future updates.

[0721] A "physical book" refers to a product in book form created by printing digitally generated content and using paper or other physical materials.

[0722] This invention is a system that automatically generates picture books customized by users via a web interface. The system mainly consists of a server, a terminal, and a generation AI model.

[0723] The device displays an input form to the user, which is necessary for customizing the story. This form is used to input the story's content, image style, number of pages, message to convey, character names, and other information. After the user enters this information, the device uses JavaScript and HTML to convert the data into JSON format and send it to the server.

[0724] The server analyzes the received data and creates instructions to pass to the image generation AI model and the text generation AI model. A Python script is used for this analysis, and the instructions are formed as specific prompts. For example, the image generation AI is given a prompt such as "Generate a watercolor-style adventure scene with a starry sky background," and the text generation AI is given a prompt such as "Create a story about children adventuring under a starry sky."

[0725] The server then uses this prompt to call an image generation AI model to generate an image in the specified style, and a text generation AI model to generate narrative text that takes vocabulary and expression into consideration. The specific processing utilizes APIs for the generation AI. For example, a Generative Adversarial Network (GAN) is used for image generation, and a natural language processing model such as GPT-3 is used for text generation.

[0726] The generated images and text are combined by the server in HTML or PDF format and saved as a digital book. The Python PIL library and ReportLab are used in this process.

[0727] Finally, the server provides the terminal with a download link for the finished product. Users can use the download link to view the picture book on their digital device and, if necessary, send feedback back to the server through a dedicated form. This feedback will be considered in the next generation process and used to improve the system.

[0728] This system allows users to easily customize and create picture books tailored to their individual needs, providing an experience that enhances learning effectiveness, especially for children.

[0729] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0730] Step 1:

[0731] The device displays a customization input form for the story to the user. Here, the user can enter details such as the story's content, image style, number of pages, message, and character names. This form is implemented using HTML and JavaScript. Upon completion, the device converts the input data into JSON format and sends it to the server.

[0732] Step 2:

[0733] The server parses the JSON data received from the terminal. This parsing process uses Python, and based on the provided data, it creates instructions for the image generation AI model and the text generation AI model. Specifically, the data processing involves parsing user input to generate prompt sentences, such as "Generate a watercolor-style adventure scene with a starry sky background."

[0734] Step 3:

[0735] The server calls an image generation AI model based on the generated prompt text and generates an image in the specified style. In this step, the prompt text is passed to the image generation AI via API, the AI ​​generates an appropriate image, and the image data is returned to the server. For example, a "scene of an adventurer with stars shining in the night sky" generated using GAN is one such example.

[0736] Step 4:

[0737] The server invokes a text generation AI model according to a prompt and generates narrative text using natural language processing. A prompt such as "Create a story about children adventuring under the starry sky" is used, and the AI ​​generates narration accordingly, outputting text such as "That night, under the shining starry sky, the adventure began..."

[0738] Step 5:

[0739] The server combines the generated images and text and records them as a digital book. This step uses Python's PIL library or ReportLab to arrange the images and text page by page and save them in HTML or PDF format. The output is an ebook file, ready for user viewing.

[0740] Step 6:

[0741] The server provides the terminal with a download link for the completed digital book. Users can download the picture book via this link and view it on their device. Furthermore, users can provide feedback to the server after use, which can be used in the next generation process. This feedback is accumulated to improve the system.

[0742] (Application Example 1)

[0743] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0744] In recent years, with the widespread adoption of digital devices, the demand for digital content that parents and children can enjoy together has increased. However, existing picture book-related content has difficulty reflecting the individual preferences of users, and there are limited means of providing interactive experiences such as creating stories together as a family. This invention aims to solve these problems by allowing users to customize the content of stories, supporting collaborative creation between parents and children, and making it instantly viewable on digital devices.

[0745] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0746] In this invention, the server includes means for receiving input from the user regarding the content of the document, the style of the images, the number of pages, and the message to be conveyed; means for making the generated electronic document viewable on a mobile information terminal; and means for presenting options for a digital version and a printed version of the generated document. This makes it possible for parents and children to collaboratively create original stories and instantly share and enjoy them on digital devices.

[0747] "Users" refers to individuals who use this system to customize documents and images and generate personalized digital content.

[0748] "Document content" refers to the text information of the story and explanations included in the generated digital picture book.

[0749] "Image style" refers to the artistic expression method of the generated visual content, and includes styles such as watercolor.

[0750] A "generative AI model" refers to an artificial intelligence system that automatically generates specific images and text based on user input.

[0751] A "prompt statement" refers to an input statement used to instruct a generative AI model on specific outputs.

[0752] "Personal information terminals" refer to computing devices that users can carry with them, such as smartphones and tablets.

[0753] The "feedback function" refers to a feature that collects user opinions and reactions to generated documents and uses them to improve the next generation process.

[0754] "Created jointly by parent and child" refers to the process where a parent and child work together to conceive the content and create a single digital piece of content.

[0755] This invention provides a specific embodiment of an interactive picture book generation system that allows parents and children to instantly enjoy stories they have collaboratively edited on a digital device.

[0756] The server displays a form on the terminal via a web interface to receive input information from the user, such as the content of the document, image format, number of pages, and the message to be conveyed. This information is sent to the server in JSON format. The server uses a generative AI model to dynamically generate images and text based on the received information. Specifically, prompt statements are used to instruct the AI ​​on what kind of images and text it should generate. For example, a prompt statement such as "Draw a watercolor-style picture of a beach and adventurers" might be used.

[0757] The terminal makes the generated electronic document viewable on a mobile device. This allows users to instantly enjoy the original story they created together as a family. Furthermore, the server presents users with the option of choosing between a digital version and a printed version of the generated document. Users can also obtain a physical printed picture book upon request.

[0758] The device also has a feedback function that sends comments and reactions to the generated document to the server. This feedback is reflected in the next story generation, allowing for content tailored to the child's interests and development. For example, when creating a picture book with a summer vacation adventure theme, the server generates images and text suitable for the document content, such as "adventurers visiting a seaside town."

[0759] In this way, the system of this invention makes it possible for parents and children to jointly generate customized digital content and provide enjoyable learning and experiences through it.

[0760] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0761] Step 1:

[0762] The terminal displays a web interface form for the user to input information such as story content, image style, number of pages, and the message to be conveyed. The entered information is organized as structured data and prepared for later transmission to the server. In this step, the input is information entered by the user, and the output is the organized data.

[0763] Step 2:

[0764] When a user presses the submit button on an input form, the terminal converts the prepared data into JSON format and sends it to the server. In this step, the input is prepared data, and the output is data encoded in JSON format.

[0765] Step 3:

[0766] The server receives JSON data sent from the terminal and analyzes it. Based on the received information, it creates the prompts necessary for the image generation AI and text generation AI. The input for this step is JSON data, and the output is prompts.

[0767] Step 4:

[0768] The server invokes a generative AI model, which generates an image in the specified style and constructs text according to the created prompt. The input for this step is the prompt, and the output is the generated image data and text data. Specifically, the image generation AI model receives instructions such as "Please draw a watercolor-style picture of a beach and adventurers" and generates an image.

[0769] Step 5:

[0770] The server combines the generated images and text to create the manuscript for the digital picture book. This manuscript is then formatted into a digital format such as HTML or PDF, making it viewable by users. The input for this step is image data and text data, and the output is the manuscript for the digital picture book.

[0771] Step 6:

[0772] The server generates and presents a link on the terminal that allows the user to download the completed digital picture book manuscript. The input for this step is the digital picture book manuscript, and the output is the download link.

[0773] Step 7:

[0774] The user uses a download link to save the e-book to their device and begins viewing it on their mobile device. In this step, the digital content, the e-book, is utilized at the user level.

[0775] Step 8:

[0776] After the user experiences the digital picture book, the device displays a feedback form, allowing them to enter their thoughts and reactions to the book. The entered feedback is sent to the server to be used in generating future stories. The input in this step is user feedback, and the output is feedback data.

[0777] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0778] This invention is a system that recognizes the user's emotions using an emotion engine and adjusts the content of the digital picture book based on that data to provide the most suitable picture book experience for each individual user. A specific embodiment of this system is described below.

[0779] First, the device provides the user with an input interface to customize the story content, image style, and number of pages. The user enters this information, and the data is sent to the server in JSON format.

[0780] The server analyzes the received data and generates AI prompts based on the input information. These generated prompts are passed to image and text generation AIs, which then generate materials based on the specified style and content. The generated illustrations and text are assembled into a story and saved in the form of a digital picture book.

[0781] Next, when a picture book is read aloud, the device activates an emotion engine. The emotion engine analyzes the user's, especially the child's, facial expressions and voice through the camera and microphone, and evaluates their current emotional state in real time. For example, if a child shows a surprised expression, the reaction can be visualized, and the next development of the story can be adjusted to increase engagement.

[0782] For example, if a child feels anxious during a tense scene in a story, the emotion engine can detect this and make adjustments such as inserting a softer, calmer scene on the next page.

[0783] Furthermore, the data accumulated by the emotion engine is stored on a server, and user preferences and reaction patterns are analyzed. This information is then fed back into future story generation, resulting in the creation of more personalized picture books.

[0784] Furthermore, user-customized digital picture books are provided via a downloadable link from the server, and if printed copies are desired as an option, production and delivery procedures are arranged.

[0785] This invention allows parents to read aloud in a way that resonates with their child's emotions, and allows children to enjoy stories tailored to their own needs.

[0786] The following describes the processing flow.

[0787] Step 1:

[0788] The device provides a web interface and displays a form where the user can input the story content, image style, number of pages, and the message they want to convey.

[0789] Step 2:

[0790] The user fills in the required information in the input form and completes the input. The data is then sent to the server in JSON format.

[0791] Step 3:

[0792] The server analyzes the received input data and creates prompts necessary for story generation. These prompts contain the information needed to generate both images and text.

[0793] Step 4:

[0794] The server invokes an image generation AI to generate an image in the specified style based on prompts. The generated image is temporarily stored in cloud storage.

[0795] Step 5:

[0796] The server invokes a text generation AI to generate sentences appropriate to the story from the prompts. This includes content tailored to a child's vocabulary level.

[0797] Step 6:

[0798] The server combines the generated images and text to create a digital picture book manuscript in HTML or PDF format. This manuscript includes the overall story structure and page layout.

[0799] Step 7:

[0800] The device activates an emotion engine and uses a camera and microphone to recognize the user's emotions in real time while reading picture books aloud.

[0801] Step 8:

[0802] The emotion engine detects emotional responses from the user's facial expressions and voice, and if, for example, a child is surprised, it receives data to adjust the content placed on the next page.

[0803] Step 9:

[0804] The server analyzes emotional data and adjusts the content of the picture book in real time based on the results. This provides a story experience tailored to the child's emotions.

[0805] Step 10:

[0806] The emotional data obtained by the emotion engine will be stored on a server and used to analyze the preferences of individual users in future picture book creation.

[0807] Step 11:

[0808] The server generates a link for the user to download the e-book and displays it on the device. If the user selects the printed book option, the server then handles information to arrange the shipping process.

[0809] (Example 2)

[0810] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0811] Existing digital picture books make it difficult to customize stories to suit the individual emotions and preferences of users. As a result, they cannot consistently provide users with an engaging and personalized storytelling experience, and there is a particular problem in that they cannot adequately deliver emotionally impactful read-aloud sessions, especially to children.

[0812] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0813] This invention includes a server that analyzes user input information to generate commands for a generation AI model and uses those commands to produce comprehensive story material; a server that analyzes the user's expressions and voice in real time and evaluates their emotions; and a server that dynamically adjusts the progression of the story according to the emotional evaluation to provide an adaptive electronic picture book experience. This makes it possible to dynamically adjust the story based on the user's individual emotions and preferences and provide an optimized picture book experience at all times.

[0814] "Users" refers to individual people who use the system to customize stories and receive an optimized picture book experience.

[0815] "The content of the story" refers to all elements related to the development of the story, such as the storyline, theme, and characters.

[0816] "Image style" refers to the artistic or design direction of the illustrations and visual materials used in a story.

[0817] "Page count" refers to the total number of pages that make up an electronic picture book, and is an element that indicates the length or volume of the story.

[0818] A "generative AI model" refers to artificial intelligence technology that automatically generates images and text based on input prompts.

[0819] A "prompt statement" is a text containing instructions given to a generative AI model, specifically a statement that directs the model to generate concrete content.

[0820] "Emotional assessment" refers to the act of analyzing and evaluating a user's current emotional state in real time based on their facial expressions and voice.

[0821] An "adaptive digital picture book experience" refers to a function that dynamically adjusts the story's progression and content according to the user's emotions and preferences, providing a personalized read-aloud experience.

[0822] This invention is a system that allows users to customize digital picture books based on their own emotions and preferences. First, the terminal provides the user with an intuitive interface where they can input the story content, image style, number of pages, etc. This allows the user to select specific keywords and create a customized story outline based on them.

[0823] The server receives the input data and analyzes it using programming languages ​​such as Python or JavaScript. From this analysis, it creates prompts to supply to the generating AI model. These prompts are provided in text format and might include something like, "Generate a space adventure story for children, focusing on friendship with colorful illustrations." Based on these prompts, AI technologies such as Stable Diffusion for image generation and the GPT series for text generation are used to specifically generate the story and illustrations.

[0824] The server integrates the generated illustrations and text to construct the manuscript for the digital picture book. This results in a digital work with narrative coherence and visual appeal.

[0825] Furthermore, during story time, the device uses its camera and microphone to transmit the user's facial expressions and voice to the emotion engine in real time. The emotion engine analyzes this data and evaluates the user's emotional state, dynamically adjusting the story's progression. For example, if a child shows a surprised expression, the story's development shifts towards a more reassuring direction to increase engagement.

[0826] The generated digital picture books are provided as downloadable links via a server, and printed and delivered versions are available upon request. In this way, users can enjoy a unique storytelling experience tailored to their individual emotions.

[0827] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0828] Step 1:

[0829] The terminal provides the user with an input interface. Here, the user can select or input information such as the story content, image style, and number of pages. This input information is collected as data in JSON format. Specifically, data is acquired using touchscreen or voice recognition technology.

[0830] Step 2:

[0831] The server receives JSON data sent from the terminal and performs analysis. A Python script is used for data analysis to extract necessary information and generate prompts required for the AI ​​model. At this stage, the data is structured, and specific keywords such as "space exploration" are extracted.

[0832] Step 3:

[0833] The server creates instructions for text and image generation AI based on the generated prompt text. These instructions are then passed to the generation AI model (e.g., Stable Diffusion or GPT series) to generate illustrations and narrative text based on the specified style and theme. Using the AI ​​prompt as input, the system obtains theme-appropriate images and text as output.

[0834] Step 4:

[0835] The server integrates the generated illustrations and text to create an electronic picture book. This creation utilizes programmatic digital layout, carefully considering the story's order and page count. As a result, a visually engaging and consistent work is produced.

[0836] Step 5:

[0837] The device uses a camera and microphone to capture the user's facial expressions and voice while a picture book is being read aloud, and transmits this data to the emotion engine. Real-time video and audio data are used as input for emotion analysis. Specifically, it performs facial recognition and voice analysis.

[0838] Step 6:

[0839] The server receives the results of emotion analysis and adjusts the story's progression accordingly. For example, if the user feels anxious, a calming scene is automatically inserted on the next page. This dynamic adjustment provides an optimal story experience that resonates with the user's emotions.

[0840] Step 7:

[0841] The server then provides the user with a customized digital picture book and generates a downloadable link. If the user wishes to have the book printed or delivered, the server handles those procedures. The final output is the individual picture book data, which is shared with the user.

[0842] (Application Example 2)

[0843] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0844] There is a challenge in providing individual users, especially children, with an appropriate and enjoyable story experience tailored to their emotions at any given time when enjoying ebooks. Furthermore, static content makes it difficult to fully personalize the user experience, resulting in a lack of increased user engagement.

[0845] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0846] In this invention, the server includes means for receiving input from the user regarding the story content, image style, number of pages, and message to be conveyed; means for creating an ebook manuscript by combining the generated images and text; and means for sensing the user's emotional state and dynamically adjusting the story content accordingly. This makes it possible to provide a personalized story experience that responds to the user's real-time emotions.

[0847] "User" refers to an individual or organization that uses the system to generate and customize stories.

[0848] A "story" refers to a series of texts and images combined to create an experience for the user, such as a narrative or episodes.

[0849] "Images" refer to the visual elements included in an e-book, specifically illustrations and diagrams generated to visually represent the story's content and style.

[0850] "Style" refers to the elements that characterize the appearance and design of a story or image, and it is the criterion that users use to select a particular theme or atmosphere.

[0851] "Page count" refers to the total number of pages in which the story unfolds within an ebook, and is an indicator used by users to determine the length of the story.

[0852] A "message" refers to the intention or meaning conveyed to the user through a story, and may include specific values ​​or lessons.

[0853] "Instructions" refer to data containing specific commands and settings necessary for generating a story, and are generated by the system.

[0854] "E-books" refer to a form of book that is stored in digital format and can be viewed by users through electronic devices.

[0855] "Emotional state" refers to data that indicates a user's psychological or emotional response, and includes information collected from the user's facial expressions and voice while they are browsing.

[0856] "Customization" refers to the process of adjusting the content and format of a story to suit the specific needs and preferences of a particular user.

[0857] This invention is an e-book system that dynamically adjusts the content of a story according to the user's emotional state. When a user inputs the story content, image style, number of pages, and message they want to convey into their device, this information is sent to a server. Based on this input information, the server creates the necessary instructions for story generation and generates prompt sentences using a generation AI model. These prompt sentences trigger image and text generation, creating the material for the story.

[0858] The device uses a camera and microphone to collect the user's facial expressions and voice in real time, and an emotion engine recognizes their emotional state. For example, OpenCV captures camera footage, and the microphone library acquires audio data. The server takes this emotional state into consideration and dynamically adjusts the content of the generated story. Depending on the emotional change, the next scene can be switched to content that provides a sense of security.

[0859] For example, if a user feels anxious during a tense scene in a story, a friendly character who alleviates the tension can be introduced in the next scene, creating a reassuring development. This allows users to receive a personalized experience tailored to their individual circumstances.

[0860] An example of a prompt to input into a generation AI model is, "The user is currently feeling anxious. Please take this emotion into consideration and generate a storyline in the next scene that will provide a sense of reassurance." By providing stories that resonate with the user in this way, the enjoyment of ebooks can be maximized.

[0861] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0862] Step 1:

[0863] The user inputs the story content, image style, number of pages, and message via their device. The entered data is temporarily stored on the device and converted to JSON format. This is the stage where the user's intentions and requests are concretized.

[0864] Step 2:

[0865] The terminal sends the converted JSON data to the server. The server parses the received data and creates the instructions necessary for story generation. Through the parsing process, the server extracts each element and reformats it into the appropriate format. The output is a prompt sentence to be input to the generating AI model.

[0866] Step 3:

[0867] The server inputs prompt text into the generative AI model, which then generates image and text materials. The generative AI model performs data calculations based on the instructions and constructs illustrations and text to match the settings. The generated materials are obtained as output.

[0868] Step 4:

[0869] The generated materials are combined on the server to create the manuscript for the ebook. The server integrates the materials to construct the overall story and compiles it into the final digital format. The output is the data of the completed ebook.

[0870] Step 5:

[0871] The device collects the user's facial expressions and voice using a camera and microphone, and analyzes them with an emotion engine. It captures video using OpenCV and records audio using the microphone library. The input consists of audio and video data, and the analysis yields the user's emotional state.

[0872] Step 6:

[0873] The server adjusts the content of the ebook based on the emotional state it receives. Using the results of the emotion engine, it replaces or rearranges text and images. As a result, the output is a story that corresponds to the user's emotions.

[0874] Step 7:

[0875] The adjusted ebook is returned to the user's device and becomes available for download. The final ebook data is sent to the device, and the user can use it for viewing or reading aloud. The output is a usable ebook.

[0876] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0877] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0878] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0879] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0880] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0881] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0882] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0883] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0884] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0885] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0886] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0887] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0888] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0889] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0890] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0891] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0892] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0893] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0894] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0895] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0896] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0897] The following is further disclosed regarding the embodiments described above.

[0898] (Claim 1)

[0899] A means of receiving input from users regarding the story content, image style, number of pages, and the message they want to convey.

[0900] A means for creating instructions necessary for story generation based on the above input information,

[0901] A means for generating images and text based on the above instructions,

[0902] A means of creating a manuscript for an electronic picture book by combining the generated images and text,

[0903] A means to enable users to download the above-mentioned digital picture book,

[0904] A system that includes this.

[0905] (Claim 2)

[0906] The system according to claim 1, further comprising means for providing options for binding and delivering the above-mentioned electronic picture book.

[0907] (Claim 3)

[0908] The system according to claim 1, further comprising a feedback function for obtaining feedback on the generated digital picture book and reflecting it in the generation of the next story.

[0909] "Example 1"

[0910] (Claim 1)

[0911] A means of receiving input from users regarding the story content, image style, number of pages, and the message they want to convey.

[0912] A means for converting the above input information into a data format and analyzing it,

[0913] A means for creating instructions for an image generation AI model and a text generation AI model based on data analysis,

[0914] A means for generating images and text using the above instructions,

[0915] A means of integrating generated images and text to create a record in the form of a digital book,

[0916] A means to enable users to download the above digital books,

[0917] A means of receiving feedback through the user interface and reflecting it in the next data generation,

[0918] A system that includes this.

[0919] (Claim 2)

[0920] We offer further delivery options for providing the above digital books as physical books.

[0921] The system according to claim 1.

[0922] (Claim 3)

[0923] Includes functionality to obtain evaluations of generated digital books, store that evaluation information, and apply it to future data generation processes.

[0924] The system according to claim 1.

[0925] "Application Example 1"

[0926] (Claim 1)

[0927] A means of receiving input from users regarding the content of the document, the format of the images, the number of pages, and the message to be conveyed.

[0928] A means for creating instructions necessary for document generation based on the above input information,

[0929] A means to make the generated electronic document viewable on a mobile device,

[0930] A means of presenting options for digital and printed versions of the generated document,

[0931] A means to enable users to download the above electronic document,

[0932] A means for dynamically generating images and text in response to a prompt using a generative AI model,

[0933] A system that includes this.

[0934] (Claim 2)

[0935] The system according to claim 1, further comprising a feedback function for obtaining user evaluations of the above-mentioned electronic document and reflecting them in the next document generation.

[0936] (Claim 3)

[0937] The system according to claim 1, which includes a function to enable parents and children to jointly create documents and immediately view them on a mobile device.

[0938] "Example 2 of combining an emotion engine"

[0939] (Claim 1)

[0940] A means of receiving input from users regarding the story content, image style, number of pages, and the message they want to convey.

[0941] A method for analyzing the above input information to generate commands for the AI ​​model, and using those commands to produce comprehensive narrative material,

[0942] A means of analyzing users' expressions and voices in real time and evaluating their emotions,

[0943] A means to dynamically adjust the progression of the story in response to emotional evaluations and provide an adaptive digital picture book experience,

[0944] A means of creating a manuscript for an electronic picture book by combining the generated images and text,

[0945] A means to enable users to download the above-mentioned digital picture book,

[0946] A system that includes this.

[0947] (Claim 2)

[0948] The system according to claim 1, further comprising means for providing options for binding and delivering the above-mentioned electronic picture book.

[0949] (Claim 3)

[0950] The system according to claim 1, which includes a feedback function for accumulating acquired emotion analysis data and reflecting it in the next story generation.

[0951] "Application example 2 when combining with an emotional engine"

[0952] (Claim 1)

[0953] A means of receiving input from users regarding the story content, image style, number of pages, and the message they want to convey.

[0954] A means for creating instructions necessary for story generation based on the above input information,

[0955] A means for generating images and text based on the above instructions,

[0956] A means of creating an ebook manuscript by combining the generated images and text,

[0957] The means by which the above e-book can be obtained by users,

[0958] A means of sensing the user's emotional state and dynamically adjusting the story content accordingly,

[0959] A system that includes this.

[0960] (Claim 2)

[0961] The system according to claim 1, further comprising means for providing options for binding and transporting the above-mentioned e-books.

[0962] (Claim 3)

[0963] The system according to claim 1, comprising a feedback function for obtaining opinions on the generated e-book and reflecting them in the next story generation. [Explanation of Symbols]

[0964] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of receiving input from users regarding the story content, image style, number of pages, and the message they want to convey. A means for creating instructions necessary for story generation based on the above input information, A means for generating images and text based on the above instructions, A means of creating a manuscript for an electronic picture book by combining the generated images and text, A means to enable users to download the above-mentioned digital picture book, A system that includes this.

2. The system according to claim 1, further comprising means for providing options for binding and delivering the above-mentioned electronic picture book.

3. The system according to claim 1, further comprising a feedback function for obtaining feedback on the generated digital picture book and reflecting it in the generation of the next story.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A