System
The system addresses the lack of personalization in traditional picture books by generating customized stories and illustrations based on user prompts, providing an interactive and enjoyable reading experience.
Patent Information
- Application Number
- JP2024123990
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
Traditional picture books lack personalization and interactivity, making it difficult to match a child's preferences and moods, and existing systems struggle to quickly generate customized stories and illustrations.
A system that receives user prompts, analyzes them to extract key information, generates personalized stories, automatically creates illustrations, and combines them into a digital picture book, enabling quick and interactive content creation.
Enables the generation of personalized digital picture books that match a child's preferences and mood, enhancing the interactive and enjoyable reading experience.
Smart Images

Figure 2026022473000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] With picture books purchased at traditional bookstores, it was difficult to provide stories that matched a child's preferences and moods, and families lacked the means to turn storytime into a personalized, interactive experience. Furthermore, existing picture books have fixed content, making it difficult to change the story or illustrations to meet individual requests. This made it difficult for children to maintain their interest, and storytime became monotonous. [Means for solving the problem]
[0005] To solve the above problems, the present invention provides the following means: a system including means for receiving prompts from a user indicating the direction of a story, means for analyzing the prompts to extract key information, means for generating a story based on the key information, means for automatically generating illustrations corresponding to the story, means for combining the story and the illustrations to create a digital picture book, and means for providing the digital picture book to the user. This system makes it possible to quickly generate a personalized story and illustrations that match a child's preferences and mood, providing an interactive and enjoyable reading aloud experience.
[0006] A "user" is an individual or group that uses the system to direct the creation of stories and illustrations.
[0007] A "prompt" is text or audio data that a user inputs into the system with information such as the direction, theme, characters, and setting of the story.
[0008] "Key information" refers to important words and phrases extracted from the prompt, and serves as a guide for generating stories and illustrations.
[0009] A "narrative generation engine" is an algorithm or program for generating the components of a story (introduction, development, climax, and conclusion) based on key information.
[0010] An "illustration generation algorithm" is an AI algorithm or program that automatically draws illustrations to match a generated story.
[0011] A "digital picture book" is a picture book provided in electronic format that combines generated stories and illustrations.
[0012] "Personalization" refers to customizing content according to a user's individual preferences and needs.
[0013] "Natural language processing algorithm" is an AI technology that analyzes the text data of prompts and extracts important words and phrases.
[0014] "Providing means" refers to a method or program for transmitting the generated digital picture book to the user's terminal and displaying it in a viewable format. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0023] [First embodiment]
[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0036] This invention relates to a system that receives prompts from a user indicating the direction of a story, generates a story based on an analysis of the prompts, and automatically generates illustrations corresponding to the story. A specific embodiment of this system is described below.
[0037] 1. User prompt input
[0038] The terminal provides an interface that accepts prompt input from the user, for example, by displaying an input form on the terminal's display, allowing the user to enter text about themes, characters, settings, etc.
[0039] Example: User types "Pirate Adventures", "Captain Rose", and "Treasure Hunt".
[0040] 2. Sending a prompt
[0041] The terminal sends the prompt entered by the user to the server. Specifically, it sends the entered text data to the server as an HTTP request.
[0042] 3. Parsing the prompt
[0043] The server parses the prompts it receives, using natural language processing algorithms to extract key phrases from the input data (e.g., "pirate," "adventure," "treasure hunt," "Captain Rose," etc.).
[0044] 4. Story Generation
[0045] The server uses a story generation engine to build a story based on the analysis results. For example, it generates a story about Captain Rose and his companions' adventure in search of lost treasure. In this process, the components of the story (introduction, development, climax, and conclusion) are generated step by step.
[0046] 5. Illustration Generation
[0047] Based on the generated story, the server uses an AI image generation algorithm to create illustrations, such as scenes of Captain Rose and her ship, treasure maps, and seascapes.
[0048] 6. Picture book structure
[0049] The server combines the generated stories and illustrations to create a digital picture book, and determines the page layout by appropriately placing the story and illustrations on each page.
[0050] 7. Sending picture books
[0051] The server sends the completed digital picture book data to the device. Specifically, it compresses the generated picture book data and sends it to the device in an HTTP response.
[0052] 8. Provision to Users
[0053] The device then provides the user with the digital data of the picture book, which is then displayed in a format that the user can view through a dedicated viewer application. It also provides the option to print the book if necessary.
[0054] The above is a concrete example for carrying out the invention. This system can quickly generate and provide a personalized picture book that matches a child's preferences and mood, making story time a more interactive and enjoyable experience.
[0055] The processing flow will be explained below.
[0056] Step 1:
[0057] The user uses the terminal to input prompts such as the theme and characters of the story. Specifically, the user enters text such as "Pirate Adventure," "Captain Rose," or "Treasure Hunt" into an input form displayed on the terminal.
[0058] Step 2:
[0059] The terminal receives the prompt entered by the user and sends it to the server. Specifically, the terminal sends the prompt to the server as an HTTP request.
[0060] Step 3:
[0061] The server parses the received prompt and uses natural language processing algorithms to extract important key phrases from the prompt (e.g., pirate, adventure, treasure hunt, Captain Rose).
[0062] Step 4:
[0063] The server starts a story generation engine based on the extracted key phrases, which generates story components such as introduction, development, climax, and conclusion step by step.
[0064] Step 5:
[0065] Based on the generated story, the server runs an illustration generation algorithm, which creates scenes of Captain Rose, her ship, treasure maps, seascapes, and more.
[0066] Step 6:
[0067] The server combines the generated story and illustrations to create the pages of the digital picture book. It then arranges the text and illustrations on each page according to the story's development and determines the page layout.
[0068] Step 7:
[0069] The server compresses the completed digital picture book data and sends it to the terminal as an HTTP response. Specifically, the data for each page of the picture book is consolidated and provided in a compressed format.
[0070] Step 8:
[0071] The device unpacks the received data and displays it to the user through a dedicated viewer application, allowing the user to view the picture book. It also provides a printing option if desired.
[0072] The system allows users to enjoy personalized stories and illustrations in a short amount of time, making storytime a more interactive and enjoyable experience.
[0073] Example 1
[0074] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0075] Traditional methods for providing personalized stories and illustrations to today's children are time-consuming and labor-intensive, making it difficult to instantly generate content that matches a child's preferences. This is especially true when providing interactive and intuitive digital picture books, making it difficult for parents and educators to provide individually customized picture books for children.
[0076] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0077] In this invention, the server includes means for receiving prompts from a user indicating the direction of a story, means for analyzing the prompts and extracting key information, means for generating a story based on the key information, means for automatically generating illustrations corresponding to the story, means for combining the story and the illustrations to create a digital picture book, means for providing the digital picture book to the user, means for compressing and transmitting the digital picture book data to be provided to the user, and means for providing an interface to be displayed on the user's terminal. This enables a user to easily and instantly generate and provide a personalized digital picture book tailored to the preferences of their child.
[0078] A "prompt" refers to a narrative direction text input received from a user.
[0079] "Parsing" refers to the process of extracting important key phrases and information from the entered prompt.
[0080] "Key information" refers to important phrases and words necessary for story generation that are extracted through prompt analysis.
[0081] "Narrative generation" refers to the process of constructing a coherent story based on analyzed key information.
[0082] "Illustration generation" refers to the process of automatically creating visual images that correspond to a generated story.
[0083] A "digital picture book" refers to a picture book in electronic format that combines a story and illustrations.
[0084] "Compression" refers to data processing performed to reduce the volume of digital picture book data.
[0085] "Transmission" refers to the act of transferring compressed digital picture book data to the user's terminal.
[0086] "Interface" refers to tools such as screens and input forms that allow users to interact with a system.
[0087] "Terminal" refers to a device through which a user can enter prompts and view the generated digital picture book.
[0088] "Server" refers to the computer system that analyzes prompts, generates stories and illustrations, and assembles and transmits the digital picture book.
[0089] This invention relates to a system that receives prompts from a user indicating the direction of a story, generates a story based on an analysis of the prompts, and automatically generates illustrations corresponding to the story. A specific embodiment of this system will be described below.
[0090] First, the user uses a terminal to input a prompt. The terminal provides an interface for accepting prompt input from the user. The interface displays an input form on the display, allowing the user to input the theme, characters, setting, etc. in text. For example, the user might input prompts such as "Pirate Adventure," "Captain Rose," or "Treasure Hunt."
[0091] Next, the terminal transmits the input prompt to the server. Specifically, the terminal transmits the prompt as text data in the form of an HTTP request to the server.
[0092] The server analyzes the received prompts using a natural language processing algorithm (e.g., GPT-4), extracts important key phrases from the prompts, and stores them as data points needed to generate the story.
[0093] Next, the server uses a story generation engine to generate a story based on the analysis results. This engine assembles the story's components (introduction, development, climax, and conclusion) step by step. For example, it generates a story about Captain Rose and his companions' adventure in search of lost treasure.
[0094] The server then uses an AI image generation algorithm (e.g., DALL-E) to automatically generate illustrations based on the generated story. Visuals of each character and scene are automatically drawn, such as Captain Rose and her ship, treasure maps, and seascapes.
[0095] The server then combines the generated story and illustrations to create a digital picture book. Using a page layout tool (e.g., Adobe InDesign API), the server arranges the text and illustrations appropriately, creating each page of the digital picture book.
[0096] The completed digital picture book is compressed and sent to the device. The server compresses the generated picture book data and sends it to the device as an HTTP response.
[0097] Finally, the device provides the received digital picture book to the user through a dedicated viewer application, allowing the user to view the book and, if necessary, use the print function.
[0098] As described above, the present invention enables users to easily and instantly create and provide personalized digital picture books that match the preferences of their children.
[0099] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0100] Step 1:
[0101] The user inputs prompts using a terminal. The terminal displays an input form on the display, allowing the user to input themes, characters, settings, etc. as text. The prompts generated by the input are treated as text data.
[0102] Input: User types "Pirate Adventures", "Captain Rose", and "Treasure Hunt".
[0103] Output: Text data of the prompt entered on the terminal.
[0104] Step 2:
[0105] The terminal sends the entered prompt to the server. Specifically, the prompt is converted into JSON format as text data and sent to the server in the form of an HTTP POST request.
[0106] Input: The text data of the prompt entered on the terminal.
[0107] Output: Text data as an HTTP request sent to the server.
[0108] Step 3:
[0109] The server parses the received prompt and uses a natural language processing algorithm (e.g., GPT-4) to extract important key phrases from within the prompt.
[0110] Input: The prompt text data sent to the server as an HTTP request.
[0111] Output: Key phrases parsed from the prompt (e.g. "pirates", "adventure", "treasure hunt", "Captain Rose").
[0112] Step 4:
[0113] The server uses a story generation engine (e.g., OpenAI model) to construct a story based on the analysis results. The story generation engine gradually assembles the components of the story (introduction, development, climax, and conclusion).
[0114] Input: Key phrase (e.g. "pirate", "adventure", "treasure hunt", "Captain Rose").
[0115] Output: A completed story (e.g., Captain Rose's quest for lost treasure).
[0116] Step 5:
[0117] Based on the generated story, the server uses an AI image generation algorithm (e.g., DALL-E) to create illustrations, generating visuals that correspond to prompts and scenes in the story.
[0118] Input: Narrative text.
[0119] Output: Illustrations for each scene in the story (e.g. Captain Rose and her ship, a treasure map, a seascape).
[0120] Step 6:
[0121] The server combines the generated story and illustrations to create a digital picture book. It uses a page layout tool (e.g., Adobe InDesign API) to arrange the text and illustrations appropriately.
[0122] Input: narrative text and corresponding illustrations.
[0123] Output: Completed digital picture book data.
[0124] Step 7:
[0125] The server compresses the data of the completed digital picture book and sends it to the device. The data is compressed in ZIP format and sent to the device as an HTTP response.
[0126] Input: Completed digital picture book data.
[0127] Output: HTTP response containing the digital picture book data compressed in ZIP format.
[0128] Step 8:
[0129] The device then provides the received digital picture book to the user through a dedicated viewer application, allowing the user to view the book and, if necessary, use the print function.
[0130] Input: HTTP response containing the digital picture book data compressed in ZIP format.
[0131] Output: A digital picture book displayed in a dedicated viewer application, and a printed picture book if desired.
[0132] (Application example 1)
[0133] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0134] Conventional digital content generation systems have issues with being unable to adequately generate stories based on user intent or automatically generate corresponding visual illustrations. Furthermore, it is difficult to personalize the generated content, preventing an improved user experience. In particular, there has been a lack of systems that can be easily applied to a variety of devices, such as smartphones, smart glasses, head-mounted displays, and robots.
[0135] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0136] In this invention, the server includes means for receiving prompts from a user indicating the direction of a story, means for analyzing the prompts and extracting key information, means for generating a story based on the key information, means for automatically generating illustrations corresponding to the story, means for combining the story and the illustrations to form digital content, means for providing the digital content to a user, and means for providing the digital content as an application installed on a smartphone, smart glasses, a head-mounted display, or a robot. This enables the automatic generation of a story and visual illustrations based on the user's prompts, and further enables personalization to suit the preferences of each individual user.
[0137] A "prompt indicating the direction of the story" is text data entered by the user, and includes information such as the theme, characters, and setting of the story.
[0138] "Means for parsing prompts" refers to natural language processing algorithms that use generative AI models to extract key phrases and important information from user prompts.
[0139] The "means for generating a story" refers to a story generation engine that generates the elements of a story (introduction, development, climax, conclusion, etc.) step by step based on the extracted key information.
[0140] "Means for automatically generating illustrations" refers to a method of generating visual illustrations using an AI image generation algorithm based on the generated story.
[0141] "Means of constructing digital content" refers to the method of combining the generated story and illustrations to create a single unified piece of digital content (e.g., a digital picture book), including page layout and design elements.
[0142] The "means for providing digital content to a user" refers to a method for compressing the completed digital content and transferring it to the user's device over a network.
[0143] An "application installed on a smartphone, smart glasses, head-mounted display, or robot" is software installed on a specific device that allows a user to view and interact with generated digital content.
[0144] "Personalization" refers to individually adjusting the content and design of generated content based on the user's preferences, mood, past usage history, etc.
[0145] This invention relates to a system that receives prompts from a user indicating the direction of a story, generates a story based on an analysis of the prompts, and automatically generates illustrations corresponding to the story. This system includes the steps of inputting a prompt from a user, analyzing the prompts, generating a story, automatically generating illustrations, and composing and providing digital content.
[0146] User prompt input
[0147] The user inputs the story's theme, characters, setting, etc. as text through an application installed on a smartphone, smart glasses, head-mounted display, or robot. For example, the user might input "adventure," "hero," or "fighting a dragon." This input interface utilizes a user interface and viewer application.
[0148] Sending and parsing prompts
[0149] User input is sent from the device to the server as an HTTP request. The server receives the prompt using a communication library such as the "requests" library and parses it using a generative AI model and a natural language processing algorithm (e.g., GPT-4). Key phrases (e.g., "adventure," "hero," and "dragon") are extracted from the input data.
[0150] Story Generation
[0151] The server runs a story generation engine based on the analysis results to generate a story. The generated story is then completed in stages, with components such as an introduction, development, climax, and conclusion.
[0152] Illustration generation
[0153] Based on the story, the server automatically generates illustrations using an AI image generation algorithm (e.g., Stable Diffusion). For example, it depicts a scene of a hero fighting a dragon or an adventure scene. These generated illustrations are placed on each page of the story.
[0154] Digital content configuration
[0155] The generated stories and illustrations are combined to form digital storybooks and other digital content, which are organized into a unified format, including page layout and design elements.
[0156] Providing digital content
[0157] The server then sends the completed digital content to the device, where it is compressed and returned as an HTTP response. The device then receives this data and displays it in a dedicated viewer application, allowing the user to browse the content. Additionally, the device offers the option to personalize the content based on the user's preferences and mood.
[0158] Prompt Sentence Examples
[0159] "Adventure Hero Fights Dragon"
[0160] This allows the system to enable users to easily create personalized stories and corresponding illustrations that can be enjoyed on a variety of devices.
[0161] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0162] Step 1:
[0163] The user inputs story prompts using an application installed on a smartphone, smart glasses, a head-mounted display, or a robot. The user inputs themes, characters, and settings in text format, such as "adventure," "hero," or "fighting a dragon." The input data is then saved on the device via the user interface.
[0164] Input: Text data such as story theme, characters, and setting
[0165] Output: Input data saved on the device
[0166] Step 2:
[0167] The terminal sends the entered prompt to the server, sends the prompt text data as an HTTP request to the server, and adds appropriate header information, and the sent data is temporarily stored on the server.
[0168] Input: Input data stored on the device
[0169] Output: Data sent to the server
[0170] Step 3:
[0171] The server analyzes the received prompt and uses a natural language processing algorithm (e.g., GPT-4) to extract keywords (e.g., "adventure," "hero," "dragon") from the prompt. The analysis results are used as key information for generating the story.
[0172] Input: Data sent to the server
[0173] Output: Parsed key information
[0174] Step 4:
[0175] The server then uses the analysis results to run a story generation engine to generate a story. Using a generative AI model, it creates a complete story, including components such as an introduction, development, climax, and conclusion. During this process, keywords entered by the user are reflected in each part of the story.
[0176] Input: Parsed key information
[0177] Output: Generated narrative text
[0178] Step 5:
[0179] The server automatically generates illustrations based on the generated story using an AI image generation algorithm (e.g., Stable Diffusion). An image corresponding to each story scene is generated. The generated illustrations are stored on the server.
[0180] Input: Generated narrative text
[0181] Output: Generated illustration image
[0182] Step 6:
[0183] The server then combines the generated story text and illustrations to create digital content. It then creates a digital picture book by designing a page layout and appropriately placing illustrations corresponding to the story on each page.
[0184] Input: Generated story text and illustration images
[0185] Output: Digital picture book data
[0186] Step 7:
[0187] The server sends the completed digital picture book data to the device. The data is compressed and returned as an HTTP response. The device saves the received digital picture book data.
[0188] Input: Digital picture book data
[0189] Output: Digital picture book data saved on the device
[0190] Step 8:
[0191] The device provides the received digital picture book data to the user. Using a dedicated viewer application, the device displays the digital picture book so that the user can view it. If necessary, the device can also personalize the book to suit the user's preferences and mood.
[0192] Input: Digital picture book data stored on the device
[0193] Output: A digital picture book that can be viewed by the user
[0194] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0195] This invention relates to a system that receives prompts from a user indicating the direction of a story, generates a story based on an analysis of the prompts, and automatically generates illustrations corresponding to the story. Furthermore, it is equipped with a function that combines an emotion engine that recognizes the user's emotions and adjusts the content of the generated story and illustrations according to the user's emotions. A specific embodiment of this system is described below.
[0196] 1. User Emotion Recognition
[0197] When the user inputs a prompt, the device uses an emotion engine to analyze the user's emotions. Specifically, it uses facial recognition and voice analysis technologies to capture the user's emotional state (happiness, sadness, excitement, etc.).
[0198] 2. User prompt input
[0199] The terminal provides an interface that accepts prompt input from the user, a process in which the user inputs themes, characters, settings, etc.
[0200] Example: When a user types "pirate adventure," "Captain Rose," or "treasure hunt," the emotion engine analyzes the user's emotional state.
[0201] 3. Sending prompts and emotion data
[0202] The terminal transmits the prompt input by the user and the analyzed emotion data to the server. Specifically, the terminal transmits an HTTP request including the text data and the emotion data to the server.
[0203] 4. Prompt and Emotion Data Analysis
[0204] The server analyzes the received prompt, using natural language processing algorithms to extract key phrases from the prompt and analyze them based on sentiment data.
[0205] 5. Story Generation
[0206] The server then launches a story generation engine based on the analysis results and emotional data. It generates the story's components (introduction, development, climax, and conclusion) step by step, adjusting the content and tone of the story according to the emotional data.
[0207] Example: If the user is in an excited state, a story with an emphasis on action and adventure may be generated.
[0208] 6. Illustration Generation
[0209] The server runs an AI image generation algorithm based on the generated story, adjusting the color and style of the resulting illustration depending on the emotional data.
[0210] Example: If the user is happy, the illustration will be bright and colorful.
[0211] 7. Picture Book Structure
[0212] The server combines the generated stories and illustrations to create pages for the digital picture book, and determines the page layout by appropriately arranging the stories and illustrations on each page.
[0213] 8. Sending picture books
[0214] The server compresses the data of the completed digital picture book and sends it to the terminal as an HTTP response, providing the generated picture book data in a compressed format.
[0215] 9. Provision to Users
[0216] The device unpacks the received data and displays it to the user through a dedicated viewer application, allowing the user to view the picture book. It also provides a printing option if desired.
[0217] This system allows users to enjoy personalized stories and illustrations based on their emotional state, making storytime a more interactive and enjoyable experience.
[0218] The processing flow will be explained below.
[0219] Step 1:
[0220] The user uses the terminal to input prompts such as the theme and characters of the story. Specifically, the user enters text such as "Pirate Adventure," "Captain Rose," and "Treasure Hunt" into an input form displayed on the terminal.
[0221] Step 2:
[0222] The device activates an emotion engine to analyze the user's emotions while the user is entering prompts. Specifically, the device's camera and microphone are used to capture the user's emotional state (e.g., joy, sadness, excitement) in real time using facial recognition or voice analysis technology.
[0223] Step 3:
[0224] The device sends the prompt entered by the user and the analyzed emotion data to the server. Specifically, it sends an HTTP request including the text data and emotion data to the server.
[0225] Step 4:
[0226] The server analyzes the received prompt and uses natural language processing algorithms to extract key phrases from the prompt, while also analyzing emotional data to recognize the user's emotional state.
[0227] Step 5:
[0228] The server launches a story generation engine based on the extracted key phrases and emotional data. It generates the elements of the story (introduction, development, climax, and conclusion) and adjusts the content and tone of the story according to the user's emotional data. Specifically, if the user is excited, a story with an emphasis on action and adventure will be generated.
[0229] Step 6:
[0230] The server runs an AI image generation algorithm based on the generated story. The AI adjusts the color and style of illustrations related to the scenario (e.g., Captain Rose, her ship, treasure map, seascape) according to the user's emotional data. Specifically, if the user is happy, illustrations with bright colors and tones are generated.
[0231] Step 7:
[0232] The server combines the generated story and illustrations to create pages of a digital picture book. It determines the page layout by appropriately arranging the story text and illustrations on each page.
[0233] Step 8:
[0234] The server compresses the completed digital picture book data and sends it to the terminal as an HTTP response. Specifically, the data for each page of the picture book is consolidated and provided in a compressed format.
[0235] Step 9:
[0236] The device unpacks the received data and displays it to the user through a dedicated viewer application, allowing the user to view the picture book. It also provides a printing option if desired.
[0237] This system allows users to enjoy personalized stories and illustrations based on their emotional state, making storytime a more interactive and enjoyable experience.
[0238] Example 2
[0239] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0240] Conventional story generation systems lack personalization based on user emotions and preferences, resulting in uniform generated content and low user satisfaction. Furthermore, the story and illustration generation processes are independent, resulting in a lack of consistency between the two. Furthermore, the lack of a way to reflect user emotions in real time compromises the interactive experience.
[0241] The specification process by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving a prompt from the user indicating the direction of the story, means for using a natural language processing algorithm to analyze the prompt and extract key information, emotion recognition means for analyzing the user's emotions, means for generating a story based on the key information and the emotions, means for automatically generating illustrations corresponding to the story in a color tone and style according to the user's emotions, means for combining the story and the illustrations to create a digital picture book, and means for providing the digital picture book to the user. This makes it possible to consistently provide stories and illustrations personalized according to the user's emotions and preferences.
[0242] A "user" is someone who utilizes the system to provide prompts that direct the story.
[0243] A "prompt" is information that indicates the direction of the story's theme, characters, setting, etc., entered by the user.
[0244] A "natural language processing algorithm" is a technology that analyzes text data and extracts meaning and key information from it.
[0245] "Emotion recognition means" is a technology that analyzes a user's facial expressions and voice and classifies the user's emotional state.
[0246] A "story generator" is an algorithm that generates story components based on prompt and emotion data.
[0247] The "illustration generation means" is a technology that automatically generates illustrations in a color tone and style that corresponds to the user's emotions based on the generated story.
[0248] A "digital picture book" is a digital picture book that combines generated stories and illustrations.
[0249] The "server" is an information processing device that acts as the backend of this system, analyzing prompts, recognizing emotions, generating stories and illustrations, and providing digital picture books.
[0250] This invention relates to a system that receives prompts from a user indicating the direction of a story, generates a story based on an analysis of the prompts, and automatically generates illustrations corresponding to the story. Furthermore, this system is equipped with an emotion engine that recognizes the user's emotions, and has the function of adjusting the content of the generated story and illustrations according to the user's emotions.
[0251] Hardware and software used
[0252] Device: A device on which a user can input prompts and view the generated digital picture book. Examples include a PC, tablet, or smartphone.
[0253] Server: A central information processing device that analyzes data, generates stories and illustrations, and creates and provides digital picture books. Specifically, it includes cloud servers and local servers.
[0254] Emotion recognition technology: Facial recognition technology (e.g., OpenCV) and voice analysis technology (e.g., Google Speech-to-Text API) are used.
[0255] Natural language processing algorithms: Algorithms such as spaCy and BERT are used to analyze text data and extract key information.
[0256] Story generation engine: A generative AI model such as GPT-3 is used to generate a story based on prompts and sentiment data.
[0257] Illustration generation algorithm: AI image generation algorithms (e.g., DALL-E, MidJourney) are used.
[0258] Examples of data processing and data calculation
[0259] 1. User Emotion Recognition
[0260] When the user inputs a prompt, the device activates the emotion engine, collects the user's facial expressions and voice through the camera and microphone, and uses facial recognition and voice analysis technology to analyze the user's emotions (happiness, sadness, excitement, etc.) in real time.
[0261] 2. User prompt input
[0262] The device provides an interface for the user to enter prompts that provide direction to the story, such as text boxes and selection menus that allow the user to enter theme, characters, and setting.
[0263] Specific examples
[0264] When the user inputs prompts such as "Pirate Adventure," "Captain Rose," or "Treasure Hunt," the device uses its emotion engine to analyze the user's emotional state and recognize, for example, that the user is excited.
[0265] 3. Sending prompts and emotion data
[0266] The device sends the input prompt and the analyzed emotion data to the server as an HTTP POST request, which includes both text data and emotion data.
[0267] 4. Prompt and Emotion Data Analysis
[0268] The server uses natural language processing algorithms to analyze the received prompts and extract important key information, while also expanding and adjusting the meaning of the prompts based on emotional data.
[0269] 5. Story Generation
[0270] The server then activates a story generation engine based on the analyzed prompt data and emotion data to generate a story, the content and tone of which are adjusted according to the user's emotional state.
[0271] Specific examples
[0272] If the user is in an excited state, the generated story will emphasize elements of action and adventure.
[0273] 6. Illustration Generation
[0274] The server then runs an AI image generation algorithm based on the generated story, generating illustrations with adjusted color tone and style depending on the emotional data.
[0275] Specific examples
[0276] If the user is happy, the illustrations generated will be colorful and bright in tone.
[0277] 7. Digital Picture Book Structure
[0278] The server combines the generated story and illustrations to create pages for a digital picture book. The page layout is adjusted to make it easy to read by adjusting font size and line spacing.
[0279] 8. Sending digital books
[0280] The server sends the completed digital picture book to the terminal in a compressed format, such as ZIP.
[0281] 9. Provision to Users
[0282] The device decompresses the received data and displays it to the user through a dedicated viewer, where the user can view the picture book and print it if necessary.
[0283] Through the above process, users can enjoy a digital picture book with stories and illustrations personalized according to their emotions and preferences.
[0284] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0285] Step 1: Recognizing user emotions
[0286] Input: User's facial expression data, voice data
[0287] Specific operation: The device collects the user's facial expressions and voice through the camera and microphone. Using facial recognition technology (e.g., OpenCV) and voice analysis technology (e.g., Google Speech-to-Text API), it analyzes the user's emotional state (e.g., joy, sadness, excitement) in real time. The analysis results are recorded as emotion data.
[0288] Output: User's emotional state (e.g., excitement level)
[0289] Step 2: User prompt input
[0290] Input: User-provided narrative prompts (theme, characters, setting)
[0291] Specific behavior: The device displays an interface (text boxes and / or selection menus) for the user to enter a prompt. The user enters a prompt that indicates the direction of the story they want to take.
[0292] Output: Prompt data (e.g. "Pirate Adventure", "Captain Rose", "Treasure Hunt")
[0293] Step 3: Sending prompts and emotion data
[0294] Input: prompt data, emotion data
[0295] How it works: The device sends the entered prompt and analyzed emotion data to the server as an HTTP POST request. This communication is performed using the secure HTTPS protocol.
[0296] Output: HTTP POST request (including prompt and emotion data)
[0297] Step 4: Analyze prompts and sentiment data
[0298] Input: HTTP POST request (prompt data, emotion data)
[0299] What happens: The server parses the incoming request, uses natural language processing algorithms (e.g., spaCy, BERT) to extract key phrases from the prompt, and optimizes the prompt content based on sentiment data.
[0300] Output: Parsed key information, prompt data with sentiment information
[0301] Step 5: Story Generation
[0302] Input: Parsed key information, prompt data with emotional information
[0303] Specific operation: The server launches a story generation engine (e.g., GPT-3) to generate a story based on the prompt and emotion data. It generates the story components (introduction, development, climax, and conclusion) step by step and adjusts them according to the user's emotions.
[0304] Output: Generated narrative data
[0305] Step 6: Illustration generation
[0306] Input: Generated story data, emotion information
[0307] How it works: The server runs an AI image generation algorithm (e.g., DALL-E, MidJourney) to generate illustrations based on the story content and the user's emotions. Color tone and style are adjusted depending on the emotion.
[0308] Output: Generated illustration data
[0309] Step 7: Composing your digital storybook
[0310] Input: Generated story data, generated illustration data
[0311] Specific operation: The server combines the story and illustrations to create the pages of the digital picture book. On each page, the server appropriately arranges the story text and corresponding illustrations and determines the page layout.
[0312] Output: Completed digital picture book data
[0313] Step 8: Submit your digital book
[0314] Input: Completed digital picture book data
[0315] Specific operation: The server compresses the completed digital picture book data and sends it to the terminal as an HTTP response. This data is often provided in ZIP format.
[0316] Output: HTTP response (compressed digital picture book data)
[0317] Step 9: Provide to users
[0318] Input: HTTP response (compressed digital picture book data)
[0319] Specific operation: The device decompresses the received data and displays it to the user through a dedicated viewer application. The user can view the picture book with this viewer and print it if necessary.
[0320] Output: A digital storybook that users can view
[0321] (Application example 2)
[0322] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0323] Conventional digital picture book generation systems could generate stories and illustrations based on user input, but they were unable to personalize this content based on the user's emotional state. This made it difficult to provide the optimal reading experience for users. New methods were needed to increase user satisfaction, especially for subjects such as children, whose emotions are strongly influenced by the content.
[0324] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0325] In this invention, the server includes means for receiving prompts from a user indicating the direction of the story, means for analyzing the prompts and extracting key information, means for generating a story based on the key information, means for automatically generating illustrations corresponding to the story, means for combining the story and the illustrations to create a digital picture book, means for providing the digital picture book to the user, means for analyzing the user's emotions, and means for adjusting the content of the story and illustrations based on the analyzed emotional data. This allows the story and illustrations to be personalized according to the user's emotional state, enabling a more satisfying interactive story-telling experience.
[0326] A "user" is someone using a terminal to enter story prompts.
[0327] A "prompt" is a string or sentence that the user enters to indicate the direction or theme of the story.
[0328] "Key information" is important information necessary for generating a story that is extracted by analyzing a prompt.
[0329] "Narrative generation" is the process of constructing the content of a story based on key information.
[0330] "Automatic illustration generation" is the act of automatically generating visual content corresponding to a generated story using technologies such as AI.
[0331] A "digital picture book" is an electronic picture book that combines generated stories and illustrations.
[0332] "Emotion analysis" is the process of analyzing a user's emotional state using facial recognition and voice analysis techniques.
[0333] "Emotion data" is data that represents the emotional state of a user obtained by emotion analysis.
[0334] "Personalization" means tailoring content to a user based on their preferences and emotional state.
[0335] A "natural language processing algorithm" is a computer algorithm that analyzes text data and understands its meaning and context.
[0336] This invention is a system that generates a story based on prompts from a user and generates illustrations corresponding to that story. It also includes a function to recognize the user's emotions and adjust the content of the story and illustrations accordingly. A specific embodiment of this system is described below.
[0337] 1. User prompt input
[0338] The user uses a terminal to input prompts that indicate the direction of the story. The input prompts are text data that includes the theme, characters, setting, etc. of the story. Examples of prompts include "Pirate Adventure," "Captain Rose," and "Treasure Hunt."
[0339] 2. Emotion analysis
[0340] The server uses an emotion engine to analyze the user's emotions when the user enters a prompt. The analysis uses facial recognition and voice analysis technologies to capture the user's emotional state (happiness, sadness, excitement, etc.). The emotion engine can use a facial recognition library (OpenCV) or a voice analysis library (Google Cloud Speech-to-Text API).
[0341] 3. Sending prompts and emotion data
[0342] The terminal transmits the prompt input by the user and the analyzed emotion data to the server, which then transmits the data as an HTTP request including text data and emotion data.
[0343] 4. Prompt and Emotion Data Analysis
[0344] The server analyzes the received prompts. It uses natural language processing algorithms to extract key phrases from the prompts and analyze them based on sentiment data. For natural language processing, open source NLP tools (e.g., spaCy or NLTK) can be used.
[0345] 5. Story Generation
[0346] The server launches a story generation engine based on the analysis results and emotional data. It generates the story's components (introduction, development, climax, and conclusion) step by step, adjusting the content and tone of the story according to the emotional data. A generative AI model (such as GPT-3) is used to generate the story.
[0347] 6. Illustration Generation
[0348] The server runs an AI image generation algorithm based on the generated story. The color and style of the generated illustration are adjusted according to the emotional data. AI image generation can use technologies such as DeepArt and DALL-E.
[0349] 7. Digital Picture Book Structure
[0350] The server combines the generated story and illustrations to create the pages of the digital picture book. It then appropriately arranges the story and illustrations on each page and determines the page layout. Software used includes PIL (Python Imaging Library).
[0351] 8. Sending digital books
[0352] The server compresses the completed digital picture book data and sends it to the terminal as an HTTP response. The generated picture book data is provided in a compressed format (e.g., ZIP format).
[0353] 9. Provision to Users
[0354] The device unpacks the received data and displays it to the user through a dedicated viewer application, allowing the user to view the digital picture book. It also provides a printing option if desired.
[0355] This system allows users to enjoy personalized stories and illustrations based on their emotional state, making story time a more interactive and enjoyable experience.
[0356] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0357] Step 1:
[0358] The user uses a terminal to input prompts that indicate the direction of the story. The input prompts are text data that includes the theme, characters, and setting of the story. Examples of inputs include "Pirate Adventure," "Captain Rose," and "Treasure Hunt." This input data is used for subsequent analysis.
[0359] Step 2:
[0360] The device uses an emotion engine to analyze the user's emotions as they input prompts. It uses facial recognition technology (e.g., OpenCV) and voice analysis technology (e.g., Google Cloud Speech-to-Text API) to capture the user's emotional state (e.g., joy, sadness, excitement, etc.). This emotional data influences the generation of the story and illustrations.
[0361] Step 3:
[0362] The device sends the prompt entered by the user and the analyzed emotion data to the server. This data is sent to the server as an HTTP request containing text data and emotion data. In order to send the input data (prompt and emotion data), the data is encoded and sent in this step.
[0363] Step 4:
[0364] The server parses the received prompt. It uses natural language processing algorithms (e.g., spaCy or NLTK) to extract important key phrases from the prompt. The extracted key phrases are used as the basis for narrative generation. This step involves text analysis and key phrase extraction.
[0365] Step 5:
[0366] The server launches a story generation engine based on the analysis results and emotional data. It uses a generative AI model (such as GPT-3) to generate story components (introduction, development, climax, and conclusion) step by step, and adjusts the content and tone of the story according to the emotional data. For example, if the user is excited about the prompt "Pirate adventure," a story with an emphasis on action and adventure will be generated.
[0367] Step 6:
[0368] The server runs an AI image generation algorithm (such as DeepArt or DALL-E) based on the generated story. The color and style of the generated illustration are adjusted according to the emotional data. For example, if the user is happy, the illustration will be colorful and bright. In this step, the process of generating visual content based on the story text is carried out.
[0369] Step 7:
[0370] The server combines the generated story and illustrations to create the pages of the digital picture book. Using software such as PIL (Python Imaging Library), the story and illustrations are laid out and beautiful page designs are created. In this step, the page layout is determined and content is integrated.
[0371] Step 8:
[0372] The server compresses the data of the completed digital picture book and sends it to the terminal as an HTTP response. The generated picture book data is provided in a compressed format (e.g., ZIP format) so that the user can download it. In this step, data compression and communication take place.
[0373] Step 9:
[0374] The device decompresses the received data and displays it to the user through a dedicated viewer application. The user can view the digital picture book using this viewer. It also provides a printing option if necessary. In this step, the data is decompressed and displayed.
[0375] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0376] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0377] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0378] [Second embodiment]
[0379] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0380] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0381] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0382] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0383] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0384] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0385] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0386] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0387] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0388] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0389] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0390] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0391] This invention relates to a system that receives prompts from a user indicating the direction of a story, generates a story based on an analysis of the prompts, and automatically generates illustrations corresponding to the story. A specific embodiment of this system is described below.
[0392] 1. User prompt input
[0393] The terminal provides an interface that accepts prompt input from the user, for example, by displaying an input form on the terminal's display, allowing the user to enter text about themes, characters, settings, etc.
[0394] Example: User types "Pirate Adventures", "Captain Rose", and "Treasure Hunt".
[0395] 2. Sending a prompt
[0396] The terminal sends the prompt entered by the user to the server. Specifically, it sends the entered text data to the server as an HTTP request.
[0397] 3. Parsing the prompt
[0398] The server parses the prompts it receives, using natural language processing algorithms to extract key phrases from the input data (e.g., "pirate," "adventure," "treasure hunt," "Captain Rose," etc.).
[0399] 4. Story Generation
[0400] The server uses a story generation engine to build a story based on the analysis results. For example, it generates a story about Captain Rose and his companions' adventure in search of lost treasure. In this process, the components of the story (introduction, development, climax, and conclusion) are generated step by step.
[0401] 5. Illustration Generation
[0402] Based on the generated story, the server uses an AI image generation algorithm to create illustrations, such as scenes of Captain Rose and her ship, treasure maps, and seascapes.
[0403] 6. Picture book structure
[0404] The server combines the generated stories and illustrations to create a digital picture book, and determines the page layout by appropriately placing the story and illustrations on each page.
[0405] 7. Sending picture books
[0406] The server sends the completed digital picture book data to the device. Specifically, it compresses the generated picture book data and sends it to the device in an HTTP response.
[0407] 8. Provision to Users
[0408] The device then provides the user with the digital data of the picture book, which is then displayed in a format that the user can view through a dedicated viewer application. It also provides the option to print the book if necessary.
[0409] The above is a concrete example for carrying out the invention. This system can quickly generate and provide a personalized picture book that matches a child's preferences and mood, making story time a more interactive and enjoyable experience.
[0410] The processing flow will be explained below.
[0411] Step 1:
[0412] The user uses the terminal to input prompts such as the theme and characters of the story. Specifically, the user enters text such as "Pirate Adventure," "Captain Rose," or "Treasure Hunt" into an input form displayed on the terminal.
[0413] Step 2:
[0414] The terminal receives the prompt entered by the user and sends it to the server. Specifically, the terminal sends the prompt to the server as an HTTP request.
[0415] Step 3:
[0416] The server parses the received prompt and uses natural language processing algorithms to extract important key phrases from the prompt (e.g., pirate, adventure, treasure hunt, Captain Rose).
[0417] Step 4:
[0418] The server starts a story generation engine based on the extracted key phrases, which generates story components such as introduction, development, climax, and conclusion step by step.
[0419] Step 5:
[0420] Based on the generated story, the server runs an illustration generation algorithm, which creates scenes of Captain Rose, her ship, treasure maps, seascapes, and more.
[0421] Step 6:
[0422] The server combines the generated story and illustrations to create the pages of the digital picture book. It then arranges the text and illustrations on each page according to the story's development and determines the page layout.
[0423] Step 7:
[0424] The server compresses the completed digital picture book data and sends it to the terminal as an HTTP response. Specifically, the data for each page of the picture book is consolidated and provided in a compressed format.
[0425] Step 8:
[0426] The device unpacks the received data and displays it to the user through a dedicated viewer application, allowing the user to view the picture book. It also provides a printing option if desired.
[0427] The system allows users to enjoy personalized stories and illustrations in a short amount of time, making storytime a more interactive and enjoyable experience.
[0428] Example 1
[0429] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0430] Traditional methods for providing personalized stories and illustrations to today's children are time-consuming and labor-intensive, making it difficult to instantly generate content that matches a child's preferences. This is especially true when providing interactive and intuitive digital picture books, making it difficult for parents and educators to provide individually customized picture books for children.
[0431] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0432] In this invention, the server includes means for receiving prompts from a user indicating the direction of a story, means for analyzing the prompts and extracting key information, means for generating a story based on the key information, means for automatically generating illustrations corresponding to the story, means for combining the story and the illustrations to create a digital picture book, means for providing the digital picture book to the user, means for compressing and transmitting the digital picture book data to be provided to the user, and means for providing an interface to be displayed on the user's terminal. This enables a user to easily and instantly generate and provide a personalized digital picture book tailored to the preferences of their child.
[0433] A "prompt" refers to a narrative direction text input received from a user.
[0434] "Parsing" refers to the process of extracting important key phrases and information from the entered prompt.
[0435] "Key information" refers to important phrases and words necessary for story generation that are extracted through prompt analysis.
[0436] "Narrative generation" refers to the process of constructing a coherent story based on analyzed key information.
[0437] "Illustration generation" refers to the process of automatically creating visual images that correspond to a generated story.
[0438] A "digital picture book" refers to a picture book in electronic format that combines a story and illustrations.
[0439] "Compression" refers to data processing performed to reduce the volume of digital picture book data.
[0440] "Transmission" refers to the act of transferring compressed digital picture book data to the user's terminal.
[0441] "Interface" refers to tools such as screens and input forms that allow users to interact with a system.
[0442] "Terminal" refers to a device through which a user can enter prompts and view the generated digital picture book.
[0443] "Server" refers to the computer system that analyzes prompts, generates stories and illustrations, and assembles and transmits the digital picture book.
[0444] This invention relates to a system that receives prompts from a user indicating the direction of a story, generates a story based on an analysis of the prompts, and automatically generates illustrations corresponding to the story. A specific embodiment of this system will be described below.
[0445] First, the user uses a terminal to input a prompt. The terminal provides an interface for accepting prompt input from the user. The interface displays an input form on the display, allowing the user to input the theme, characters, setting, etc. in text. For example, the user might input prompts such as "Pirate Adventure," "Captain Rose," or "Treasure Hunt."
[0446] Next, the terminal transmits the input prompt to the server. Specifically, the terminal transmits the prompt as text data in the form of an HTTP request to the server.
[0447] The server analyzes the received prompts using a natural language processing algorithm (e.g., GPT-4), extracts important key phrases from the prompts, and stores them as data points needed to generate the story.
[0448] Next, the server uses a story generation engine to generate a story based on the analysis results. This engine assembles the story's components (introduction, development, climax, and conclusion) step by step. For example, it generates a story about Captain Rose and his companions' adventure in search of lost treasure.
[0449] The server then uses an AI image generation algorithm (e.g., DALL-E) to automatically generate illustrations based on the generated story. Visuals of each character and scene are automatically drawn, such as Captain Rose and her ship, treasure maps, and seascapes.
[0450] The server then combines the generated story and illustrations to create a digital picture book. Using a page layout tool (e.g., Adobe InDesign API), the server arranges the text and illustrations appropriately, creating each page of the digital picture book.
[0451] The completed digital picture book is compressed and sent to the device. The server compresses the generated picture book data and sends it to the device as an HTTP response.
[0452] Finally, the device provides the received digital picture book to the user through a dedicated viewer application, allowing the user to view the book and, if necessary, use the print function.
[0453] As described above, the present invention enables users to easily and instantly create and provide personalized digital picture books that match the preferences of their children.
[0454] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0455] Step 1:
[0456] The user inputs prompts using a terminal. The terminal displays an input form on the display, allowing the user to input themes, characters, settings, etc. as text. The prompts generated by the input are treated as text data.
[0457] Input: User types "Pirate Adventures", "Captain Rose", and "Treasure Hunt".
[0458] Output: Text data of the prompt entered on the terminal.
[0459] Step 2:
[0460] The terminal sends the entered prompt to the server. Specifically, the prompt is converted into JSON format as text data and sent to the server in the form of an HTTP POST request.
[0461] Input: The text data of the prompt entered on the terminal.
[0462] Output: Text data as an HTTP request sent to the server.
[0463] Step 3:
[0464] The server parses the received prompt and uses a natural language processing algorithm (e.g., GPT-4) to extract important key phrases from within the prompt.
[0465] Input: The prompt text data sent to the server as an HTTP request.
[0466] Output: Key phrases parsed from the prompt (e.g. "pirates", "adventure", "treasure hunt", "Captain Rose").
[0467] Step 4:
[0468] The server uses a story generation engine (e.g., OpenAI model) to construct a story based on the analysis results. The story generation engine gradually assembles the components of the story (introduction, development, climax, and conclusion).
[0469] Input: Key phrase (e.g. "pirate", "adventure", "treasure hunt", "Captain Rose").
[0470] Output: A completed story (e.g., Captain Rose's quest for lost treasure).
[0471] Step 5:
[0472] Based on the generated story, the server uses an AI image generation algorithm (e.g., DALL-E) to create illustrations, generating visuals that correspond to prompts and scenes in the story.
[0473] Input: Narrative text.
[0474] Output: Illustrations for each scene in the story (e.g. Captain Rose and her ship, a treasure map, a seascape).
[0475] Step 6:
[0476] The server combines the generated story and illustrations to create a digital picture book. It uses a page layout tool (e.g., Adobe InDesign API) to arrange the text and illustrations appropriately.
[0477] Input: narrative text and corresponding illustrations.
[0478] Output: Completed digital picture book data.
[0479] Step 7:
[0480] The server compresses the data of the completed digital picture book and sends it to the device. The data is compressed in ZIP format and sent to the device as an HTTP response.
[0481] Input: Completed digital picture book data.
[0482] Output: HTTP response containing the digital picture book data compressed in ZIP format.
[0483] Step 8:
[0484] The device then provides the received digital picture book to the user through a dedicated viewer application, allowing the user to view the book and, if necessary, use the print function.
[0485] Input: HTTP response containing the digital picture book data compressed in ZIP format.
[0486] Output: A digital picture book displayed in a dedicated viewer application, and a printed picture book if desired.
[0487] (Application example 1)
[0488] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0489] Conventional digital content generation systems have issues with being unable to adequately generate stories based on user intent or automatically generate corresponding visual illustrations. Furthermore, it is difficult to personalize the generated content, preventing an improved user experience. In particular, there has been a lack of systems that can be easily applied to a variety of devices, such as smartphones, smart glasses, head-mounted displays, and robots.
[0490] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0491] In this invention, the server includes means for receiving prompts from a user indicating the direction of a story, means for analyzing the prompts and extracting key information, means for generating a story based on the key information, means for automatically generating illustrations corresponding to the story, means for combining the story and the illustrations to form digital content, means for providing the digital content to a user, and means for providing the digital content as an application installed on a smartphone, smart glasses, a head-mounted display, or a robot. This enables the automatic generation of a story and visual illustrations based on the user's prompts, and further enables personalization to suit the preferences of each individual user.
[0492] A "prompt indicating the direction of the story" is text data entered by the user, and includes information such as the theme, characters, and setting of the story.
[0493] "Means for parsing prompts" refers to natural language processing algorithms that use generative AI models to extract key phrases and important information from user prompts.
[0494] The "means for generating a story" refers to a story generation engine that generates the elements of a story (introduction, development, climax, conclusion, etc.) step by step based on the extracted key information.
[0495] "Means for automatically generating illustrations" refers to a method of generating visual illustrations using an AI image generation algorithm based on the generated story.
[0496] "Means of constructing digital content" refers to the method of combining the generated story and illustrations to create a single unified piece of digital content (e.g., a digital picture book), including page layout and design elements.
[0497] The "means for providing digital content to a user" refers to a method for compressing the completed digital content and transferring it to the user's device over a network.
[0498] An "application installed on a smartphone, smart glasses, head-mounted display, or robot" is software installed on a specific device that allows a user to view and interact with generated digital content.
[0499] "Personalization" refers to individually adjusting the content and design of generated content based on the user's preferences, mood, past usage history, etc.
[0500] This invention relates to a system that receives prompts from a user indicating the direction of a story, generates a story based on an analysis of the prompts, and automatically generates illustrations corresponding to the story. This system includes the steps of inputting a prompt from a user, analyzing the prompts, generating a story, automatically generating illustrations, and composing and providing digital content.
[0501] User prompt input
[0502] The user inputs the story's theme, characters, setting, etc. as text through an application installed on a smartphone, smart glasses, head-mounted display, or robot. For example, the user might input "adventure," "hero," or "fighting a dragon." This input interface utilizes a user interface and viewer application.
[0503] Sending and parsing prompts
[0504] User input is sent from the device to the server as an HTTP request. The server receives the prompt using a communication library such as the "requests" library and parses it using a generative AI model and a natural language processing algorithm (e.g., GPT-4). Key phrases (e.g., "adventure," "hero," and "dragon") are extracted from the input data.
[0505] Story Generation
[0506] The server runs a story generation engine based on the analysis results to generate a story. The generated story is then completed in stages, with components such as an introduction, development, climax, and conclusion.
[0507] Illustration generation
[0508] Based on the story, the server automatically generates illustrations using an AI image generation algorithm (e.g., Stable Diffusion). For example, it depicts a scene of a hero fighting a dragon or an adventure scene. These generated illustrations are placed on each page of the story.
[0509] Digital content configuration
[0510] The generated stories and illustrations are combined to form digital storybooks and other digital content, which are organized into a unified format, including page layout and design elements.
[0511] Providing digital content
[0512] The server then sends the completed digital content to the device, where it is compressed and returned as an HTTP response. The device then receives this data and displays it in a dedicated viewer application, allowing the user to browse the content. Additionally, the device offers the option to personalize the content based on the user's preferences and mood.
[0513] Prompt Sentence Examples
[0514] "Adventure Hero Fights Dragon"
[0515] This allows the system to enable users to easily create personalized stories and corresponding illustrations that can be enjoyed on a variety of devices.
[0516] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0517] Step 1:
[0518] The user inputs story prompts using an application installed on a smartphone, smart glasses, a head-mounted display, or a robot. The user inputs themes, characters, and settings in text format, such as "adventure," "hero," or "fighting a dragon." The input data is then saved on the device via the user interface.
[0519] Input: Text data such as story theme, characters, and setting
[0520] Output: Input data saved on the device
[0521] Step 2:
[0522] The terminal sends the entered prompt to the server, sends the prompt text data as an HTTP request to the server, and adds appropriate header information, and the sent data is temporarily stored on the server.
[0523] Input: Input data stored on the device
[0524] Output: Data sent to the server
[0525] Step 3:
[0526] The server analyzes the received prompt and uses a natural language processing algorithm (e.g., GPT-4) to extract keywords (e.g., "adventure," "hero," "dragon") from the prompt. The analysis results are used as key information for generating the story.
[0527] Input: Data sent to the server
[0528] Output: Parsed key information
[0529] Step 4:
[0530] The server then uses the analysis results to run a story generation engine to generate a story. Using a generative AI model, it creates a complete story, including components such as an introduction, development, climax, and conclusion. During this process, keywords entered by the user are reflected in each part of the story.
[0531] Input: Parsed key information
[0532] Output: Generated narrative text
[0533] Step 5:
[0534] The server automatically generates illustrations based on the generated story using an AI image generation algorithm (e.g., Stable Diffusion). An image corresponding to each story scene is generated. The generated illustrations are stored on the server.
[0535] Input: Generated narrative text
[0536] Output: Generated illustration image
[0537] Step 6:
[0538] The server then combines the generated story text and illustrations to create digital content. It then creates a digital picture book by designing a page layout and appropriately placing illustrations corresponding to the story on each page.
[0539] Input: Generated story text and illustration images
[0540] Output: Digital picture book data
[0541] Step 7:
[0542] The server sends the completed digital picture book data to the device. The data is compressed and returned as an HTTP response. The device saves the received digital picture book data.
[0543] Input: Digital picture book data
[0544] Output: Digital picture book data saved on the device
[0545] Step 8:
[0546] The device provides the received digital picture book data to the user. Using a dedicated viewer application, the device displays the digital picture book so that the user can view it. If necessary, the device can also personalize the book to suit the user's preferences and mood.
[0547] Input: Digital picture book data stored on the device
[0548] Output: A digital picture book that can be viewed by the user
[0549] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0550] This invention relates to a system that receives prompts from a user indicating the direction of a story, generates a story based on an analysis of the prompts, and automatically generates illustrations corresponding to the story. Furthermore, it is equipped with a function that combines an emotion engine that recognizes the user's emotions and adjusts the content of the generated story and illustrations according to the user's emotions. A specific embodiment of this system is described below.
[0551] 1. User Emotion Recognition
[0552] When the user inputs a prompt, the device uses an emotion engine to analyze the user's emotions. Specifically, it uses facial recognition and voice analysis technologies to capture the user's emotional state (happiness, sadness, excitement, etc.).
[0553] 2. User prompt input
[0554] The terminal provides an interface that accepts prompt input from the user, a process in which the user inputs themes, characters, settings, etc.
[0555] Example: When a user types "pirate adventure," "Captain Rose," or "treasure hunt," the emotion engine analyzes the user's emotional state.
[0556] 3. Sending prompts and emotion data
[0557] The terminal transmits the prompt input by the user and the analyzed emotion data to the server. Specifically, the terminal transmits an HTTP request including the text data and the emotion data to the server.
[0558] 4. Prompt and Emotion Data Analysis
[0559] The server analyzes the received prompt, using natural language processing algorithms to extract key phrases from the prompt and analyze them based on sentiment data.
[0560] 5. Story Generation
[0561] The server then launches a story generation engine based on the analysis results and emotional data. It generates the story's components (introduction, development, climax, and conclusion) step by step, adjusting the content and tone of the story according to the emotional data.
[0562] Example: If the user is in an excited state, a story with an emphasis on action and adventure may be generated.
[0563] 6. Illustration Generation
[0564] The server runs an AI image generation algorithm based on the generated story, adjusting the color and style of the resulting illustration depending on the emotional data.
[0565] Example: If the user is happy, the illustration will be bright and colorful.
[0566] 7. Picture Book Structure
[0567] The server combines the generated stories and illustrations to create pages for the digital picture book, and determines the page layout by appropriately arranging the stories and illustrations on each page.
[0568] 8. Sending picture books
[0569] The server compresses the data of the completed digital picture book and sends it to the terminal as an HTTP response, providing the generated picture book data in a compressed format.
[0570] 9. Provision to Users
[0571] The device unpacks the received data and displays it to the user through a dedicated viewer application, allowing the user to view the picture book. It also provides a printing option if desired.
[0572] This system allows users to enjoy personalized stories and illustrations based on their emotional state, making storytime a more interactive and enjoyable experience.
[0573] The processing flow will be explained below.
[0574] Step 1:
[0575] The user uses the terminal to input prompts such as the theme and characters of the story. Specifically, the user enters text such as "Pirate Adventure," "Captain Rose," and "Treasure Hunt" into an input form displayed on the terminal.
[0576] Step 2:
[0577] The device activates an emotion engine to analyze the user's emotions while the user is entering prompts. Specifically, the device's camera and microphone are used to capture the user's emotional state (e.g., joy, sadness, excitement) in real time using facial recognition or voice analysis technology.
[0578] Step 3:
[0579] The device sends the prompt entered by the user and the analyzed emotion data to the server. Specifically, it sends an HTTP request including the text data and emotion data to the server.
[0580] Step 4:
[0581] The server analyzes the received prompt and uses natural language processing algorithms to extract key phrases from the prompt, while also analyzing emotional data to recognize the user's emotional state.
[0582] Step 5:
[0583] The server launches a story generation engine based on the extracted key phrases and emotional data. It generates the elements of the story (introduction, development, climax, and conclusion) and adjusts the content and tone of the story according to the user's emotional data. Specifically, if the user is excited, a story with an emphasis on action and adventure will be generated.
[0584] Step 6:
[0585] The server runs an AI image generation algorithm based on the generated story. The AI adjusts the color and style of illustrations related to the scenario (e.g., Captain Rose, her ship, treasure map, seascape) according to the user's emotional data. Specifically, if the user is happy, illustrations with bright colors and tones are generated.
[0586] Step 7:
[0587] The server combines the generated story and illustrations to create pages of a digital picture book. It determines the page layout by appropriately arranging the story text and illustrations on each page.
[0588] Step 8:
[0589] The server compresses the completed digital picture book data and sends it to the terminal as an HTTP response. Specifically, the data for each page of the picture book is consolidated and provided in a compressed format.
[0590] Step 9:
[0591] The device unpacks the received data and displays it to the user through a dedicated viewer application, allowing the user to view the picture book. It also provides a printing option if desired.
[0592] This system allows users to enjoy personalized stories and illustrations based on their emotional state, making storytime a more interactive and enjoyable experience.
[0593] Example 2
[0594] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0595] Conventional story generation systems lack personalization based on user emotions and preferences, resulting in uniform generated content and low user satisfaction. Furthermore, the story and illustration generation processes are independent, resulting in a lack of consistency between the two. Furthermore, the lack of a way to reflect user emotions in real time compromises the interactive experience.
[0596] The specification process by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving a prompt from the user indicating the direction of the story, means for using a natural language processing algorithm to analyze the prompt and extract key information, emotion recognition means for analyzing the user's emotions, means for generating a story based on the key information and the emotions, means for automatically generating illustrations corresponding to the story in a color tone and style according to the user's emotions, means for combining the story and the illustrations to create a digital picture book, and means for providing the digital picture book to the user. This makes it possible to consistently provide stories and illustrations personalized according to the user's emotions and preferences.
[0597] A "user" is someone who utilizes the system to provide prompts that direct the story.
[0598] A "prompt" is information that indicates the direction of the story's theme, characters, setting, etc., entered by the user.
[0599] A "natural language processing algorithm" is a technology that analyzes text data and extracts meaning and key information from it.
[0600] "Emotion recognition means" is a technology that analyzes a user's facial expressions and voice and classifies the user's emotional state.
[0601] A "story generator" is an algorithm that generates story components based on prompt and emotion data.
[0602] The "illustration generation means" is a technology that automatically generates illustrations in a color tone and style that corresponds to the user's emotions based on the generated story.
[0603] A "digital picture book" is a digital picture book that combines generated stories and illustrations.
[0604] The "server" is an information processing device that acts as the backend of this system, analyzing prompts, recognizing emotions, generating stories and illustrations, and providing digital picture books.
[0605] This invention relates to a system that receives prompts from a user indicating the direction of a story, generates a story based on an analysis of the prompts, and automatically generates illustrations corresponding to the story. Furthermore, this system is equipped with an emotion engine that recognizes the user's emotions, and has the function of adjusting the content of the generated story and illustrations according to the user's emotions.
[0606] Hardware and software used
[0607] Device: A device on which a user can input prompts and view the generated digital picture book. Examples include a PC, tablet, or smartphone.
[0608] Server: A central information processing device that analyzes data, generates stories and illustrations, and creates and provides digital picture books. Specifically, it includes cloud servers and local servers.
[0609] Emotion recognition technology: Facial recognition technology (e.g., OpenCV) and voice analysis technology (e.g., Google Speech-to-Text API) are used.
[0610] Natural language processing algorithms: Algorithms such as spaCy and BERT are used to analyze text data and extract key information.
[0611] Story generation engine: A generative AI model such as GPT-3 is used to generate a story based on prompts and sentiment data.
[0612] Illustration generation algorithm: AI image generation algorithms (e.g., DALL-E, MidJourney) are used.
[0613] Examples of data processing and data calculation
[0614] 1. User Emotion Recognition
[0615] When the user inputs a prompt, the device activates the emotion engine, collects the user's facial expressions and voice through the camera and microphone, and uses facial recognition and voice analysis technology to analyze the user's emotions (happiness, sadness, excitement, etc.) in real time.
[0616] 2. User prompt input
[0617] The device provides an interface for the user to enter prompts that provide direction to the story, such as text boxes and selection menus that allow the user to enter theme, characters, and setting.
[0618] Specific examples
[0619] When the user inputs prompts such as "Pirate Adventure," "Captain Rose," or "Treasure Hunt," the device uses its emotion engine to analyze the user's emotional state and recognize, for example, that the user is excited.
[0620] 3. Sending prompts and emotion data
[0621] The device sends the input prompt and the analyzed emotion data to the server as an HTTP POST request, which includes both text data and emotion data.
[0622] 4. Prompt and Emotion Data Analysis
[0623] The server uses natural language processing algorithms to analyze the received prompts and extract important key information, while also expanding and adjusting the meaning of the prompts based on emotional data.
[0624] 5. Story Generation
[0625] The server then activates a story generation engine based on the analyzed prompt data and emotion data to generate a story, the content and tone of which are adjusted according to the user's emotional state.
[0626] Specific examples
[0627] If the user is in an excited state, the generated story will emphasize elements of action and adventure.
[0628] 6. Illustration Generation
[0629] The server then runs an AI image generation algorithm based on the generated story, generating illustrations with adjusted color tone and style depending on the emotional data.
[0630] Specific examples
[0631] If the user is happy, the illustrations generated will be colorful and bright in tone.
[0632] 7. Digital Picture Book Structure
[0633] The server combines the generated story and illustrations to create pages for a digital picture book. The page layout is adjusted to make it easy to read by adjusting font size and line spacing.
[0634] 8. Sending digital books
[0635] The server sends the completed digital picture book to the terminal in a compressed format, such as ZIP.
[0636] 9. Provision to Users
[0637] The device decompresses the received data and displays it to the user through a dedicated viewer, where the user can view the picture book and print it if necessary.
[0638] Through the above process, users can enjoy a digital picture book with stories and illustrations personalized according to their emotions and preferences.
[0639] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0640] Step 1: Recognizing user emotions
[0641] Input: User's facial expression data, voice data
[0642] Specific operation: The device collects the user's facial expressions and voice through the camera and microphone. Using facial recognition technology (e.g., OpenCV) and voice analysis technology (e.g., Google Speech-to-Text API), it analyzes the user's emotional state (e.g., joy, sadness, excitement) in real time. The analysis results are recorded as emotion data.
[0643] Output: User's emotional state (e.g., excitement level)
[0644] Step 2: User prompt input
[0645] Input: User-provided narrative prompts (theme, characters, setting)
[0646] Specific behavior: The device displays an interface (text boxes and / or selection menus) for the user to enter a prompt. The user enters a prompt that indicates the direction of the story they want to take.
[0647] Output: Prompt data (e.g. "Pirate Adventure", "Captain Rose", "Treasure Hunt")
[0648] Step 3: Sending prompts and emotion data
[0649] Input: prompt data, emotion data
[0650] How it works: The device sends the entered prompt and analyzed emotion data to the server as an HTTP POST request. This communication is performed using the secure HTTPS protocol.
[0651] Output: HTTP POST request (including prompt and emotion data)
[0652] Step 4: Analyze prompts and sentiment data
[0653] Input: HTTP POST request (prompt data, emotion data)
[0654] What happens: The server parses the incoming request, uses natural language processing algorithms (e.g., spaCy, BERT) to extract key phrases from the prompt, and optimizes the prompt content based on sentiment data.
[0655] Output: Parsed key information, prompt data with sentiment information
[0656] Step 5: Story Generation
[0657] Input: Parsed key information, prompt data with emotional information
[0658] Specific operation: The server launches a story generation engine (e.g., GPT-3) to generate a story based on the prompt and emotion data. It generates the story components (introduction, development, climax, and conclusion) step by step and adjusts them according to the user's emotions.
[0659] Output: Generated narrative data
[0660] Step 6: Illustration generation
[0661] Input: Generated story data, emotion information
[0662] How it works: The server runs an AI image generation algorithm (e.g., DALL-E, MidJourney) to generate illustrations based on the story content and the user's emotions. Color tone and style are adjusted depending on the emotion.
[0663] Output: Generated illustration data
[0664] Step 7: Composing your digital storybook
[0665] Input: Generated story data, generated illustration data
[0666] Specific operation: The server combines the story and illustrations to create the pages of the digital picture book. On each page, the server appropriately arranges the story text and corresponding illustrations and determines the page layout.
[0667] Output: Completed digital picture book data
[0668] Step 8: Submit your digital book
[0669] Input: Completed digital picture book data
[0670] Specific operation: The server compresses the completed digital picture book data and sends it to the terminal as an HTTP response. This data is often provided in ZIP format.
[0671] Output: HTTP response (compressed digital picture book data)
[0672] Step 9: Provide to users
[0673] Input: HTTP response (compressed digital picture book data)
[0674] Specific operation: The device decompresses the received data and displays it to the user through a dedicated viewer application. The user can view the picture book with this viewer and print it if necessary.
[0675] Output: A digital storybook that users can view
[0676] (Application example 2)
[0677] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0678] Conventional digital picture book generation systems could generate stories and illustrations based on user input, but they were unable to personalize this content based on the user's emotional state. This made it difficult to provide the optimal reading experience for users. New methods were needed to increase user satisfaction, especially for subjects such as children, whose emotions are strongly influenced by the content.
[0679] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0680] In this invention, the server includes means for receiving prompts from a user indicating the direction of the story, means for analyzing the prompts and extracting key information, means for generating a story based on the key information, means for automatically generating illustrations corresponding to the story, means for combining the story and the illustrations to create a digital picture book, means for providing the digital picture book to the user, means for analyzing the user's emotions, and means for adjusting the content of the story and illustrations based on the analyzed emotional data. This allows the story and illustrations to be personalized according to the user's emotional state, enabling a more satisfying interactive story-telling experience.
[0681] A "user" is someone using a terminal to enter story prompts.
[0682] A "prompt" is a string or sentence that the user enters to indicate the direction or theme of the story.
[0683] "Key information" is important information necessary for generating a story that is extracted by analyzing a prompt.
[0684] "Narrative generation" is the process of constructing the content of a story based on key information.
[0685] "Automatic illustration generation" is the act of automatically generating visual content corresponding to a generated story using technologies such as AI.
[0686] A "digital picture book" is an electronic picture book that combines generated stories and illustrations.
[0687] "Emotion analysis" is the process of analyzing a user's emotional state using facial recognition and voice analysis techniques.
[0688] "Emotion data" is data that represents the emotional state of a user obtained by emotion analysis.
[0689] "Personalization" means tailoring content to a user based on their preferences and emotional state.
[0690] A "natural language processing algorithm" is a computer algorithm that analyzes text data and understands its meaning and context.
[0691] This invention is a system that generates a story based on prompts from a user and generates illustrations corresponding to that story. It also includes a function to recognize the user's emotions and adjust the content of the story and illustrations accordingly. A specific embodiment of this system is described below.
[0692] 1. User prompt input
[0693] The user uses a terminal to input prompts that indicate the direction of the story. The input prompts are text data that includes the theme, characters, setting, etc. of the story. Examples of prompts include "Pirate Adventure," "Captain Rose," and "Treasure Hunt."
[0694] 2. Emotion analysis
[0695] The server uses an emotion engine to analyze the user's emotions when the user enters a prompt. The analysis uses facial recognition and voice analysis technologies to capture the user's emotional state (happiness, sadness, excitement, etc.). The emotion engine can use a facial recognition library (OpenCV) or a voice analysis library (Google Cloud Speech-to-Text API).
[0696] 3. Sending prompts and emotion data
[0697] The terminal transmits the prompt input by the user and the analyzed emotion data to the server, which then transmits the data as an HTTP request including text data and emotion data.
[0698] 4. Prompt and Emotion Data Analysis
[0699] The server analyzes the received prompts. It uses natural language processing algorithms to extract key phrases from the prompts and analyze them based on sentiment data. For natural language processing, open source NLP tools (e.g., spaCy or NLTK) can be used.
[0700] 5. Story Generation
[0701] The server launches a story generation engine based on the analysis results and emotional data. It generates the story's components (introduction, development, climax, and conclusion) step by step, adjusting the content and tone of the story according to the emotional data. A generative AI model (such as GPT-3) is used to generate the story.
[0702] 6. Illustration Generation
[0703] The server runs an AI image generation algorithm based on the generated story. The color and style of the generated illustration are adjusted according to the emotional data. AI image generation can use technologies such as DeepArt and DALL-E.
[0704] 7. Digital Picture Book Structure
[0705] The server combines the generated story and illustrations to create the pages of the digital picture book. It then appropriately arranges the story and illustrations on each page and determines the page layout. Software used includes PIL (Python Imaging Library).
[0706] 8. Sending digital books
[0707] The server compresses the completed digital picture book data and sends it to the terminal as an HTTP response. The generated picture book data is provided in a compressed format (e.g., ZIP format).
[0708] 9. Provision to Users
[0709] The device unpacks the received data and displays it to the user through a dedicated viewer application, allowing the user to view the digital picture book. It also provides a printing option if desired.
[0710] This system allows users to enjoy personalized stories and illustrations based on their emotional state, making story time a more interactive and enjoyable experience.
[0711] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0712] Step 1:
[0713] The user uses a terminal to input prompts that indicate the direction of the story. The input prompts are text data that includes the theme, characters, and setting of the story. Examples of inputs include "Pirate Adventure," "Captain Rose," and "Treasure Hunt." This input data is used for subsequent analysis.
[0714] Step 2:
[0715] The device uses an emotion engine to analyze the user's emotions as they input prompts. It uses facial recognition technology (e.g., OpenCV) and voice analysis technology (e.g., Google Cloud Speech-to-Text API) to capture the user's emotional state (e.g., joy, sadness, excitement, etc.). This emotional data influences the generation of the story and illustrations.
[0716] Step 3:
[0717] The device sends the prompt entered by the user and the analyzed emotion data to the server. This data is sent to the server as an HTTP request containing text data and emotion data. In order to send the input data (prompt and emotion data), the data is encoded and sent in this step.
[0718] Step 4:
[0719] The server parses the received prompt. It uses natural language processing algorithms (e.g., spaCy or NLTK) to extract important key phrases from the prompt. The extracted key phrases are used as the basis for narrative generation. This step involves text analysis and key phrase extraction.
[0720] Step 5:
[0721] The server launches a story generation engine based on the analysis results and emotional data. It uses a generative AI model (such as GPT-3) to generate story components (introduction, development, climax, and conclusion) step by step, and adjusts the content and tone of the story according to the emotional data. For example, if the user is excited about the prompt "Pirate adventure," a story with an emphasis on action and adventure will be generated.
[0722] Step 6:
[0723] The server runs an AI image generation algorithm (such as DeepArt or DALL-E) based on the generated story. The color and style of the generated illustration are adjusted according to the emotional data. For example, if the user is happy, the illustration will be colorful and bright. In this step, the process of generating visual content based on the story text is carried out.
[0724] Step 7:
[0725] The server combines the generated story and illustrations to create the pages of the digital picture book. Using software such as PIL (Python Imaging Library), the story and illustrations are laid out and beautiful page designs are created. In this step, the page layout is determined and content is integrated.
[0726] Step 8:
[0727] The server compresses the data of the completed digital picture book and sends it to the terminal as an HTTP response. The generated picture book data is provided in a compressed format (e.g., ZIP format) so that the user can download it. In this step, data compression and communication take place.
[0728] Step 9:
[0729] The device decompresses the received data and displays it to the user through a dedicated viewer application. The user can view the digital picture book using this viewer. It also provides a printing option if necessary. In this step, the data is decompressed and displayed.
[0730] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0731] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0732] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0733] [Third embodiment]
[0734] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0735] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0736] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0737] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0738] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0739] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0740] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0741] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0742] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0743] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0744] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0745] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0746] This invention relates to a system that receives prompts from a user indicating the direction of a story, generates a story based on an analysis of the prompts, and automatically generates illustrations corresponding to the story. A specific embodiment of this system is described below.
[0747] 1. User prompt input
[0748] The terminal provides an interface that accepts prompt input from the user, for example, by displaying an input form on the terminal's display, allowing the user to enter text about themes, characters, settings, etc.
[0749] Example: User types "Pirate Adventures", "Captain Rose", and "Treasure Hunt".
[0750] 2. Sending a prompt
[0751] The terminal sends the prompt entered by the user to the server. Specifically, it sends the entered text data to the server as an HTTP request.
[0752] 3. Parsing the prompt
[0753] The server parses the prompts it receives, using natural language processing algorithms to extract key phrases from the input data (e.g., "pirate," "adventure," "treasure hunt," "Captain Rose," etc.).
[0754] 4. Story Generation
[0755] The server uses a story generation engine to build a story based on the analysis results. For example, it generates a story about Captain Rose and his companions' adventure in search of lost treasure. In this process, the components of the story (introduction, development, climax, and conclusion) are generated step by step.
[0756] 5. Illustration Generation
[0757] Based on the generated story, the server uses an AI image generation algorithm to create illustrations, such as scenes of Captain Rose and her ship, treasure maps, and seascapes.
[0758] 6. Picture book structure
[0759] The server combines the generated stories and illustrations to create a digital picture book, and determines the page layout by appropriately placing the story and illustrations on each page.
[0760] 7. Sending picture books
[0761] The server sends the completed digital picture book data to the device. Specifically, it compresses the generated picture book data and sends it to the device in an HTTP response.
[0762] 8. Provision to Users
[0763] The device then provides the user with the digital data of the picture book, which is then displayed in a format that the user can view through a dedicated viewer application. It also provides the option to print the book if necessary.
[0764] The above is a concrete example for carrying out the invention. This system can quickly generate and provide a personalized picture book that matches a child's preferences and mood, making story time a more interactive and enjoyable experience.
[0765] The processing flow will be explained below.
[0766] Step 1:
[0767] The user uses the terminal to input prompts such as the theme and characters of the story. Specifically, the user enters text such as "Pirate Adventure," "Captain Rose," or "Treasure Hunt" into an input form displayed on the terminal.
[0768] Step 2:
[0769] The terminal receives the prompt entered by the user and sends it to the server. Specifically, the terminal sends the prompt to the server as an HTTP request.
[0770] Step 3:
[0771] The server parses the received prompt and uses natural language processing algorithms to extract important key phrases from the prompt (e.g., pirate, adventure, treasure hunt, Captain Rose).
[0772] Step 4:
[0773] The server starts a story generation engine based on the extracted key phrases, which generates story components such as introduction, development, climax, and conclusion step by step.
[0774] Step 5:
[0775] Based on the generated story, the server runs an illustration generation algorithm, which creates scenes of Captain Rose, her ship, treasure maps, seascapes, and more.
[0776] Step 6:
[0777] The server combines the generated story and illustrations to create the pages of the digital picture book. It then arranges the text and illustrations on each page according to the story's development and determines the page layout.
[0778] Step 7:
[0779] The server compresses the completed digital picture book data and sends it to the terminal as an HTTP response. Specifically, the data for each page of the picture book is consolidated and provided in a compressed format.
[0780] Step 8:
[0781] The device unpacks the received data and displays it to the user through a dedicated viewer application, allowing the user to view the picture book. It also provides a printing option if desired.
[0782] The system allows users to enjoy personalized stories and illustrations in a short amount of time, making storytime a more interactive and enjoyable experience.
[0783] Example 1
[0784] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0785] Traditional methods for providing personalized stories and illustrations to today's children are time-consuming and labor-intensive, making it difficult to instantly generate content that matches a child's preferences. This is especially true when providing interactive and intuitive digital picture books, making it difficult for parents and educators to provide individually customized picture books for children.
[0786] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0787] In this invention, the server includes means for receiving prompts from a user indicating the direction of a story, means for analyzing the prompts and extracting key information, means for generating a story based on the key information, means for automatically generating illustrations corresponding to the story, means for combining the story and the illustrations to create a digital picture book, means for providing the digital picture book to the user, means for compressing and transmitting the digital picture book data to be provided to the user, and means for providing an interface to be displayed on the user's terminal. This enables a user to easily and instantly generate and provide a personalized digital picture book tailored to the preferences of their child.
[0788] A "prompt" refers to a narrative direction text input received from a user.
[0789] "Parsing" refers to the process of extracting important key phrases and information from the entered prompt.
[0790] "Key information" refers to important phrases and words necessary for story generation that are extracted through prompt analysis.
[0791] "Narrative generation" refers to the process of constructing a coherent story based on analyzed key information.
[0792] "Illustration generation" refers to the process of automatically creating visual images that correspond to a generated story.
[0793] A "digital picture book" refers to a picture book in electronic format that combines a story and illustrations.
[0794] "Compression" refers to data processing performed to reduce the volume of digital picture book data.
[0795] "Transmission" refers to the act of transferring compressed digital picture book data to the user's terminal.
[0796] "Interface" refers to tools such as screens and input forms that allow users to interact with a system.
[0797] "Terminal" refers to a device through which a user can enter prompts and view the generated digital picture book.
[0798] "Server" refers to the computer system that analyzes prompts, generates stories and illustrations, and assembles and transmits the digital picture book.
[0799] This invention relates to a system that receives prompts from a user indicating the direction of a story, generates a story based on an analysis of the prompts, and automatically generates illustrations corresponding to the story. A specific embodiment of this system will be described below.
[0800] First, the user uses a terminal to input a prompt. The terminal provides an interface for accepting prompt input from the user. The interface displays an input form on the display, allowing the user to input the theme, characters, setting, etc. in text. For example, the user might input prompts such as "Pirate Adventure," "Captain Rose," or "Treasure Hunt."
[0801] Next, the terminal transmits the input prompt to the server. Specifically, the terminal transmits the prompt as text data in the form of an HTTP request to the server.
[0802] The server analyzes the received prompts using a natural language processing algorithm (e.g., GPT-4), extracts important key phrases from the prompts, and stores them as data points needed to generate the story.
[0803] Next, the server uses a story generation engine to generate a story based on the analysis results. This engine assembles the story's components (introduction, development, climax, and conclusion) step by step. For example, it generates a story about Captain Rose and his companions' adventure in search of lost treasure.
[0804] The server then uses an AI image generation algorithm (e.g., DALL-E) to automatically generate illustrations based on the generated story. Visuals of each character and scene are automatically drawn, such as Captain Rose and her ship, treasure maps, and seascapes.
[0805] The server then combines the generated story and illustrations to create a digital picture book. Using a page layout tool (e.g., Adobe InDesign API), the server arranges the text and illustrations appropriately, creating each page of the digital picture book.
[0806] The completed digital picture book is compressed and sent to the device. The server compresses the generated picture book data and sends it to the device as an HTTP response.
[0807] Finally, the device provides the received digital picture book to the user through a dedicated viewer application, allowing the user to view the book and, if necessary, use the print function.
[0808] As described above, the present invention enables users to easily and instantly create and provide personalized digital picture books that match the preferences of their children.
[0809] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0810] Step 1:
[0811] The user inputs prompts using a terminal. The terminal displays an input form on the display, allowing the user to input themes, characters, settings, etc. as text. The prompts generated by the input are treated as text data.
[0812] Input: User types "Pirate Adventures", "Captain Rose", and "Treasure Hunt".
[0813] Output: Text data of the prompt entered on the terminal.
[0814] Step 2:
[0815] The terminal sends the entered prompt to the server. Specifically, the prompt is converted into JSON format as text data and sent to the server in the form of an HTTP POST request.
[0816] Input: The text data of the prompt entered on the terminal.
[0817] Output: Text data as an HTTP request sent to the server.
[0818] Step 3:
[0819] The server parses the received prompt and uses a natural language processing algorithm (e.g., GPT-4) to extract important key phrases from within the prompt.
[0820] Input: The prompt text data sent to the server as an HTTP request.
[0821] Output: Key phrases parsed from the prompt (e.g. "pirates", "adventure", "treasure hunt", "Captain Rose").
[0822] Step 4:
[0823] The server uses a story generation engine (e.g., OpenAI model) to construct a story based on the analysis results. The story generation engine gradually assembles the components of the story (introduction, development, climax, and conclusion).
[0824] Input: Key phrase (e.g. "pirate", "adventure", "treasure hunt", "Captain Rose").
[0825] Output: A completed story (e.g., Captain Rose's quest for lost treasure).
[0826] Step 5:
[0827] Based on the generated story, the server uses an AI image generation algorithm (e.g., DALL-E) to create illustrations, generating visuals that correspond to prompts and scenes in the story.
[0828] Input: Narrative text.
[0829] Output: Illustrations for each scene in the story (e.g. Captain Rose and her ship, a treasure map, a seascape).
[0830] Step 6:
[0831] The server combines the generated story and illustrations to create a digital picture book. It uses a page layout tool (e.g., Adobe InDesign API) to arrange the text and illustrations appropriately.
[0832] Input: narrative text and corresponding illustrations.
[0833] Output: Completed digital picture book data.
[0834] Step 7:
[0835] The server compresses the data of the completed digital picture book and sends it to the device. The data is compressed in ZIP format and sent to the device as an HTTP response.
[0836] Input: Completed digital picture book data.
[0837] Output: HTTP response containing the digital picture book data compressed in ZIP format.
[0838] Step 8:
[0839] The device then provides the received digital picture book to the user through a dedicated viewer application, allowing the user to view the book and, if necessary, use the print function.
[0840] Input: HTTP response containing the digital picture book data compressed in ZIP format.
[0841] Output: A digital picture book displayed in a dedicated viewer application, and a printed picture book if desired.
[0842] (Application example 1)
[0843] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0844] Conventional digital content generation systems have issues with being unable to adequately generate stories based on user intent or automatically generate corresponding visual illustrations. Furthermore, it is difficult to personalize the generated content, preventing an improved user experience. In particular, there has been a lack of systems that can be easily applied to a variety of devices, such as smartphones, smart glasses, head-mounted displays, and robots.
[0845] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0846] In this invention, the server includes means for receiving prompts from a user indicating the direction of a story, means for analyzing the prompts and extracting key information, means for generating a story based on the key information, means for automatically generating illustrations corresponding to the story, means for combining the story and the illustrations to form digital content, means for providing the digital content to a user, and means for providing the digital content as an application installed on a smartphone, smart glasses, a head-mounted display, or a robot. This enables the automatic generation of a story and visual illustrations based on the user's prompts, and further enables personalization to suit the preferences of each individual user.
[0847] A "prompt indicating the direction of the story" is text data entered by the user, and includes information such as the theme, characters, and setting of the story.
[0848] "Means for parsing prompts" refers to natural language processing algorithms that use generative AI models to extract key phrases and important information from user prompts.
[0849] The "means for generating a story" refers to a story generation engine that generates the elements of a story (introduction, development, climax, conclusion, etc.) step by step based on the extracted key information.
[0850] "Means for automatically generating illustrations" refers to a method of generating visual illustrations using an AI image generation algorithm based on the generated story.
[0851] "Means of constructing digital content" refers to the method of combining the generated story and illustrations to create a single unified piece of digital content (e.g., a digital picture book), including page layout and design elements.
[0852] The "means for providing digital content to a user" refers to a method for compressing the completed digital content and transferring it to the user's device over a network.
[0853] An "application installed on a smartphone, smart glasses, head-mounted display, or robot" is software installed on a specific device that allows a user to view and interact with generated digital content.
[0854] "Personalization" refers to individually adjusting the content and design of generated content based on the user's preferences, mood, past usage history, etc.
[0855] This invention relates to a system that receives prompts from a user indicating the direction of a story, generates a story based on an analysis of the prompts, and automatically generates illustrations corresponding to the story. This system includes the steps of inputting a prompt from a user, analyzing the prompts, generating a story, automatically generating illustrations, and composing and providing digital content.
[0856] User prompt input
[0857] The user inputs the story's theme, characters, setting, etc. as text through an application installed on a smartphone, smart glasses, head-mounted display, or robot. For example, the user might input "adventure," "hero," or "fighting a dragon." This input interface utilizes a user interface and viewer application.
[0858] Sending and parsing prompts
[0859] User input is sent from the device to the server as an HTTP request. The server receives the prompt using a communication library such as the "requests" library and parses it using a generative AI model and a natural language processing algorithm (e.g., GPT-4). Key phrases (e.g., "adventure," "hero," and "dragon") are extracted from the input data.
[0860] Story Generation
[0861] The server runs a story generation engine based on the analysis results to generate a story. The generated story is then completed in stages, with components such as an introduction, development, climax, and conclusion.
[0862] Illustration generation
[0863] Based on the story, the server automatically generates illustrations using an AI image generation algorithm (e.g., Stable Diffusion). For example, it depicts a scene of a hero fighting a dragon or an adventure scene. These generated illustrations are placed on each page of the story.
[0864] Digital content configuration
[0865] The generated stories and illustrations are combined to form digital storybooks and other digital content, which are organized into a unified format, including page layout and design elements.
[0866] Providing digital content
[0867] The server then sends the completed digital content to the device, where it is compressed and returned as an HTTP response. The device then receives this data and displays it in a dedicated viewer application, allowing the user to browse the content. Additionally, the device offers the option to personalize the content based on the user's preferences and mood.
[0868] Prompt Sentence Examples
[0869] "Adventure Hero Fights Dragon"
[0870] This allows the system to enable users to easily create personalized stories and corresponding illustrations that can be enjoyed on a variety of devices.
[0871] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0872] Step 1:
[0873] The user inputs story prompts using an application installed on a smartphone, smart glasses, a head-mounted display, or a robot. The user inputs themes, characters, and settings in text format, such as "adventure," "hero," or "fighting a dragon." The input data is then saved on the device via the user interface.
[0874] Input: Text data such as story theme, characters, and setting
[0875] Output: Input data saved on the device
[0876] Step 2:
[0877] The terminal sends the entered prompt to the server, sends the prompt text data as an HTTP request to the server, and adds appropriate header information, and the sent data is temporarily stored on the server.
[0878] Input: Input data stored on the device
[0879] Output: Data sent to the server
[0880] Step 3:
[0881] The server analyzes the received prompt and uses a natural language processing algorithm (e.g., GPT-4) to extract keywords (e.g., "adventure," "hero," "dragon") from the prompt. The analysis results are used as key information for generating the story.
[0882] Input: Data sent to the server
[0883] Output: Parsed key information
[0884] Step 4:
[0885] The server then uses the analysis results to run a story generation engine to generate a story. Using a generative AI model, it creates a complete story, including components such as an introduction, development, climax, and conclusion. During this process, keywords entered by the user are reflected in each part of the story.
[0886] Input: Parsed key information
[0887] Output: Generated narrative text
[0888] Step 5:
[0889] The server automatically generates illustrations based on the generated story using an AI image generation algorithm (e.g., Stable Diffusion). An image corresponding to each story scene is generated. The generated illustrations are stored on the server.
[0890] Input: Generated narrative text
[0891] Output: Generated illustration image
[0892] Step 6:
[0893] The server then combines the generated story text and illustrations to create digital content. It then creates a digital picture book by designing a page layout and appropriately placing illustrations corresponding to the story on each page.
[0894] Input: Generated story text and illustration images
[0895] Output: Digital picture book data
[0896] Step 7:
[0897] The server sends the completed digital picture book data to the device. The data is compressed and returned as an HTTP response. The device saves the received digital picture book data.
[0898] Input: Digital picture book data
[0899] Output: Digital picture book data saved on the device
[0900] Step 8:
[0901] The device provides the received digital picture book data to the user. Using a dedicated viewer application, the device displays the digital picture book so that the user can view it. If necessary, the device can also personalize the book to suit the user's preferences and mood.
[0902] Input: Digital picture book data stored on the device
[0903] Output: A digital picture book that can be viewed by the user
[0904] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0905] This invention relates to a system that receives prompts from a user indicating the direction of a story, generates a story based on an analysis of the prompts, and automatically generates illustrations corresponding to the story. Furthermore, it is equipped with a function that combines an emotion engine that recognizes the user's emotions and adjusts the content of the generated story and illustrations according to the user's emotions. A specific embodiment of this system is described below.
[0906] 1. User Emotion Recognition
[0907] When the user inputs a prompt, the device uses an emotion engine to analyze the user's emotions. Specifically, it uses facial recognition and voice analysis technologies to capture the user's emotional state (happiness, sadness, excitement, etc.).
[0908] 2. User prompt input
[0909] The terminal provides an interface that accepts prompt input from the user, a process in which the user inputs themes, characters, settings, etc.
[0910] Example: When a user types "pirate adventure," "Captain Rose," or "treasure hunt," the emotion engine analyzes the user's emotional state.
[0911] 3. Sending prompts and emotion data
[0912] The terminal transmits the prompt input by the user and the analyzed emotion data to the server. Specifically, the terminal transmits an HTTP request including the text data and the emotion data to the server.
[0913] 4. Prompt and Emotion Data Analysis
[0914] The server analyzes the received prompt, using natural language processing algorithms to extract key phrases from the prompt and analyze them based on sentiment data.
[0915] 5. Story Generation
[0916] The server then launches a story generation engine based on the analysis results and emotional data. It generates the story's components (introduction, development, climax, and conclusion) step by step, adjusting the content and tone of the story according to the emotional data.
[0917] Example: If the user is in an excited state, a story with an emphasis on action and adventure may be generated.
[0918] 6. Illustration Generation
[0919] The server runs an AI image generation algorithm based on the generated story, adjusting the color and style of the resulting illustration depending on the emotional data.
[0920] Example: If the user is happy, the illustration will be bright and colorful.
[0921] 7. Picture Book Structure
[0922] The server combines the generated stories and illustrations to create pages for the digital picture book, and determines the page layout by appropriately arranging the stories and illustrations on each page.
[0923] 8. Sending picture books
[0924] The server compresses the data of the completed digital picture book and sends it to the terminal as an HTTP response, providing the generated picture book data in a compressed format.
[0925] 9. Provision to Users
[0926] The device unpacks the received data and displays it to the user through a dedicated viewer application, allowing the user to view the picture book. It also provides a printing option if desired.
[0927] This system allows users to enjoy personalized stories and illustrations based on their emotional state, making storytime a more interactive and enjoyable experience.
[0928] The processing flow will be explained below.
[0929] Step 1:
[0930] The user uses the terminal to input prompts such as the theme and characters of the story. Specifically, the user enters text such as "Pirate Adventure," "Captain Rose," and "Treasure Hunt" into an input form displayed on the terminal.
[0931] Step 2:
[0932] The device activates an emotion engine to analyze the user's emotions while the user is entering prompts. Specifically, the device's camera and microphone are used to capture the user's emotional state (e.g., joy, sadness, excitement) in real time using facial recognition or voice analysis technology.
[0933] Step 3:
[0934] The device sends the prompt entered by the user and the analyzed emotion data to the server. Specifically, it sends an HTTP request including the text data and emotion data to the server.
[0935] Step 4:
[0936] The server analyzes the received prompt and uses natural language processing algorithms to extract key phrases from the prompt, while also analyzing emotional data to recognize the user's emotional state.
[0937] Step 5:
[0938] The server launches a story generation engine based on the extracted key phrases and emotional data. It generates the elements of the story (introduction, development, climax, and conclusion) and adjusts the content and tone of the story according to the user's emotional data. Specifically, if the user is excited, a story with an emphasis on action and adventure will be generated.
[0939] Step 6:
[0940] The server runs an AI image generation algorithm based on the generated story. The AI adjusts the color and style of illustrations related to the scenario (e.g., Captain Rose, her ship, treasure map, seascape) according to the user's emotional data. Specifically, if the user is happy, illustrations with bright colors and tones are generated.
[0941] Step 7:
[0942] The server combines the generated story and illustrations to create pages of a digital picture book. It determines the page layout by appropriately arranging the story text and illustrations on each page.
[0943] Step 8:
[0944] The server compresses the completed digital picture book data and sends it to the terminal as an HTTP response. Specifically, the data for each page of the picture book is consolidated and provided in a compressed format.
[0945] Step 9:
[0946] The device unpacks the received data and displays it to the user through a dedicated viewer application, allowing the user to view the picture book. It also provides a printing option if desired.
[0947] This system allows users to enjoy personalized stories and illustrations based on their emotional state, making storytime a more interactive and enjoyable experience.
[0948] Example 2
[0949] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0950] Conventional story generation systems lack personalization based on user emotions and preferences, resulting in uniform generated content and low user satisfaction. Furthermore, the story and illustration generation processes are independent, resulting in a lack of consistency between the two. Furthermore, the lack of a way to reflect user emotions in real time compromises the interactive experience.
[0951] The specification process by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving a prompt from the user indicating the direction of the story, means for using a natural language processing algorithm to analyze the prompt and extract key information, emotion recognition means for analyzing the user's emotions, means for generating a story based on the key information and the emotions, means for automatically generating illustrations corresponding to the story in a color tone and style according to the user's emotions, means for combining the story and the illustrations to create a digital picture book, and means for providing the digital picture book to the user. This makes it possible to consistently provide stories and illustrations personalized according to the user's emotions and preferences.
[0952] A "user" is someone who utilizes the system to provide prompts that direct the story.
[0953] A "prompt" is information that indicates the direction of the story's theme, characters, setting, etc., entered by the user.
[0954] A "natural language processing algorithm" is a technology that analyzes text data and extracts meaning and key information from it.
[0955] "Emotion recognition means" is a technology that analyzes a user's facial expressions and voice and classifies the user's emotional state.
[0956] A "story generator" is an algorithm that generates story components based on prompt and emotion data.
[0957] The "illustration generation means" is a technology that automatically generates illustrations in a color tone and style that corresponds to the user's emotions based on the generated story.
[0958] A "digital picture book" is a digital picture book that combines generated stories and illustrations.
[0959] The "server" is an information processing device that acts as the backend of this system, analyzing prompts, recognizing emotions, generating stories and illustrations, and providing digital picture books.
[0960] This invention relates to a system that receives prompts from a user indicating the direction of a story, generates a story based on an analysis of the prompts, and automatically generates illustrations corresponding to the story. Furthermore, this system is equipped with an emotion engine that recognizes the user's emotions, and has the function of adjusting the content of the generated story and illustrations according to the user's emotions.
[0961] Hardware and software used
[0962] Device: A device on which a user can input prompts and view the generated digital picture book. Examples include a PC, tablet, or smartphone.
[0963] Server: A central information processing device that analyzes data, generates stories and illustrations, and creates and provides digital picture books. Specifically, it includes cloud servers and local servers.
[0964] Emotion recognition technology: Facial recognition technology (e.g., OpenCV) and voice analysis technology (e.g., Google Speech-to-Text API) are used.
[0965] Natural language processing algorithms: Algorithms such as spaCy and BERT are used to analyze text data and extract key information.
[0966] Story generation engine: A generative AI model such as GPT-3 is used to generate a story based on prompts and sentiment data.
[0967] Illustration generation algorithm: AI image generation algorithms (e.g., DALL-E, MidJourney) are used.
[0968] Examples of data processing and data calculation
[0969] 1. User Emotion Recognition
[0970] When the user inputs a prompt, the device activates the emotion engine, collects the user's facial expressions and voice through the camera and microphone, and uses facial recognition and voice analysis technology to analyze the user's emotions (happiness, sadness, excitement, etc.) in real time.
[0971] 2. User prompt input
[0972] The device provides an interface for the user to enter prompts that provide direction to the story, such as text boxes and selection menus that allow the user to enter theme, characters, and setting.
[0973] Specific examples
[0974] When the user inputs prompts such as "Pirate Adventure," "Captain Rose," or "Treasure Hunt," the device uses its emotion engine to analyze the user's emotional state and recognize, for example, that the user is excited.
[0975] 3. Sending prompts and emotion data
[0976] The device sends the input prompt and the analyzed emotion data to the server as an HTTP POST request, which includes both text data and emotion data.
[0977] 4. Prompt and Emotion Data Analysis
[0978] The server uses natural language processing algorithms to analyze the received prompts and extract important key information, while also expanding and adjusting the meaning of the prompts based on emotional data.
[0979] 5. Story Generation
[0980] The server then activates a story generation engine based on the analyzed prompt data and emotion data to generate a story, the content and tone of which are adjusted according to the user's emotional state.
[0981] Specific examples
[0982] If the user is in an excited state, the generated story will emphasize elements of action and adventure.
[0983] 6. Illustration Generation
[0984] The server then runs an AI image generation algorithm based on the generated story, generating illustrations with adjusted color tone and style depending on the emotional data.
[0985] Specific examples
[0986] If the user is happy, the illustrations generated will be colorful and bright in tone.
[0987] 7. Digital Picture Book Structure
[0988] The server combines the generated story and illustrations to create pages for a digital picture book. The page layout is adjusted to make it easy to read by adjusting font size and line spacing.
[0989] 8. Sending digital books
[0990] The server sends the completed digital picture book to the terminal in a compressed format, such as ZIP.
[0991] 9. Provision to Users
[0992] The device decompresses the received data and displays it to the user through a dedicated viewer, where the user can view the picture book and print it if necessary.
[0993] Through the above process, users can enjoy a digital picture book with stories and illustrations personalized according to their emotions and preferences.
[0994] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0995] Step 1: Recognizing user emotions
[0996] Input: User's facial expression data, voice data
[0997] Specific operation: The device collects the user's facial expressions and voice through the camera and microphone. Using facial recognition technology (e.g., OpenCV) and voice analysis technology (e.g., Google Speech-to-Text API), it analyzes the user's emotional state (e.g., joy, sadness, excitement) in real time. The analysis results are recorded as emotion data.
[0998] Output: User's emotional state (e.g., excitement level)
[0999] Step 2: User prompt input
[1000] Input: User-provided narrative prompts (theme, characters, setting)
[1001] Specific behavior: The device displays an interface (text boxes and / or selection menus) for the user to enter a prompt. The user enters a prompt that indicates the direction of the story they want to take.
[1002] Output: Prompt data (e.g. "Pirate Adventure", "Captain Rose", "Treasure Hunt")
[1003] Step 3: Sending prompts and emotion data
[1004] Input: prompt data, emotion data
[1005] How it works: The device sends the entered prompt and analyzed emotion data to the server as an HTTP POST request. This communication is performed using the secure HTTPS protocol.
[1006] Output: HTTP POST request (including prompt and emotion data)
[1007] Step 4: Analyze prompts and sentiment data
[1008] Input: HTTP POST request (prompt data, emotion data)
[1009] What happens: The server parses the incoming request, uses natural language processing algorithms (e.g., spaCy, BERT) to extract key phrases from the prompt, and optimizes the prompt content based on sentiment data.
[1010] Output: Parsed key information, prompt data with sentiment information
[1011] Step 5: Story Generation
[1012] Input: Parsed key information, prompt data with emotional information
[1013] Specific operation: The server launches a story generation engine (e.g., GPT-3) to generate a story based on the prompt and emotion data. It generates the story components (introduction, development, climax, and conclusion) step by step and adjusts them according to the user's emotions.
[1014] Output: Generated narrative data
[1015] Step 6: Illustration generation
[1016] Input: Generated story data, emotion information
[1017] How it works: The server runs an AI image generation algorithm (e.g., DALL-E, MidJourney) to generate illustrations based on the story content and the user's emotions. Color tone and style are adjusted depending on the emotion.
[1018] Output: Generated illustration data
[1019] Step 7: Composing your digital storybook
[1020] Input: Generated story data, generated illustration data
[1021] Specific operation: The server combines the story and illustrations to create the pages of the digital picture book. On each page, the server appropriately arranges the story text and corresponding illustrations and determines the page layout.
[1022] Output: Completed digital picture book data
[1023] Step 8: Submit your digital book
[1024] Input: Completed digital picture book data
[1025] Specific operation: The server compresses the completed digital picture book data and sends it to the terminal as an HTTP response. This data is often provided in ZIP format.
[1026] Output: HTTP response (compressed digital picture book data)
[1027] Step 9: Provide to users
[1028] Input: HTTP response (compressed digital picture book data)
[1029] Specific operation: The device decompresses the received data and displays it to the user through a dedicated viewer application. The user can view the picture book with this viewer and print it if necessary.
[1030] Output: A digital storybook that users can view
[1031] (Application example 2)
[1032] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1033] Conventional digital picture book generation systems could generate stories and illustrations based on user input, but they were unable to personalize this content based on the user's emotional state. This made it difficult to provide the optimal reading experience for users. New methods were needed to increase user satisfaction, especially for subjects such as children, whose emotions are strongly influenced by the content.
[1034] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1035] In this invention, the server includes means for receiving prompts from a user indicating the direction of the story, means for analyzing the prompts and extracting key information, means for generating a story based on the key information, means for automatically generating illustrations corresponding to the story, means for combining the story and the illustrations to create a digital picture book, means for providing the digital picture book to the user, means for analyzing the user's emotions, and means for adjusting the content of the story and illustrations based on the analyzed emotional data. This allows the story and illustrations to be personalized according to the user's emotional state, enabling a more satisfying interactive story-telling experience.
[1036] A "user" is someone using a terminal to enter story prompts.
[1037] A "prompt" is a string or sentence that the user enters to indicate the direction or theme of the story.
[1038] "Key information" is important information necessary for generating a story that is extracted by analyzing a prompt.
[1039] "Narrative generation" is the process of constructing the content of a story based on key information.
[1040] "Automatic illustration generation" is the act of automatically generating visual content corresponding to a generated story using technologies such as AI.
[1041] A "digital picture book" is an electronic picture book that combines generated stories and illustrations.
[1042] "Emotion analysis" is the process of analyzing a user's emotional state using facial recognition and voice analysis techniques.
[1043] "Emotion data" is data that represents the emotional state of a user obtained by emotion analysis.
[1044] "Personalization" means tailoring content to a user based on their preferences and emotional state.
[1045] A "natural language processing algorithm" is a computer algorithm that analyzes text data and understands its meaning and context.
[1046] This invention is a system that generates a story based on prompts from a user and generates illustrations corresponding to that story. It also includes a function to recognize the user's emotions and adjust the content of the story and illustrations accordingly. A specific embodiment of this system is described below.
[1047] 1. User prompt input
[1048] The user uses a terminal to input prompts that indicate the direction of the story. The input prompts are text data that includes the theme, characters, setting, etc. of the story. Examples of prompts include "Pirate Adventure," "Captain Rose," and "Treasure Hunt."
[1049] 2. Emotion analysis
[1050] The server uses an emotion engine to analyze the user's emotions when the user enters a prompt. The analysis uses facial recognition and voice analysis technologies to capture the user's emotional state (happiness, sadness, excitement, etc.). The emotion engine can use a facial recognition library (OpenCV) or a voice analysis library (Google Cloud Speech-to-Text API).
[1051] 3. Sending prompts and emotion data
[1052] The terminal transmits the prompt input by the user and the analyzed emotion data to the server, which then transmits the data as an HTTP request including text data and emotion data.
[1053] 4. Prompt and Emotion Data Analysis
[1054] The server analyzes the received prompts. It uses natural language processing algorithms to extract key phrases from the prompts and analyze them based on sentiment data. For natural language processing, open source NLP tools (e.g., spaCy or NLTK) can be used.
[1055] 5. Story Generation
[1056] The server launches a story generation engine based on the analysis results and emotional data. It generates the story's components (introduction, development, climax, and conclusion) step by step, adjusting the content and tone of the story according to the emotional data. A generative AI model (such as GPT-3) is used to generate the story.
[1057] 6. Illustration Generation
[1058] The server runs an AI image generation algorithm based on the generated story. The color and style of the generated illustration are adjusted according to the emotional data. AI image generation can use technologies such as DeepArt and DALL-E.
[1059] 7. Digital Picture Book Structure
[1060] The server combines the generated story and illustrations to create the pages of the digital picture book. It then appropriately arranges the story and illustrations on each page and determines the page layout. Software used includes PIL (Python Imaging Library).
[1061] 8. Sending digital books
[1062] The server compresses the completed digital picture book data and sends it to the terminal as an HTTP response. The generated picture book data is provided in a compressed format (e.g., ZIP format).
[1063] 9. Provision to Users
[1064] The device unpacks the received data and displays it to the user through a dedicated viewer application, allowing the user to view the digital picture book. It also provides a printing option if desired.
[1065] This system allows users to enjoy personalized stories and illustrations based on their emotional state, making story time a more interactive and enjoyable experience.
[1066] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1067] Step 1:
[1068] The user uses a terminal to input prompts that indicate the direction of the story. The input prompts are text data that includes the theme, characters, and setting of the story. Examples of inputs include "Pirate Adventure," "Captain Rose," and "Treasure Hunt." This input data is used for subsequent analysis.
[1069] Step 2:
[1070] The device uses an emotion engine to analyze the user's emotions as they input prompts. It uses facial recognition technology (e.g., OpenCV) and voice analysis technology (e.g., Google Cloud Speech-to-Text API) to capture the user's emotional state (e.g., joy, sadness, excitement, etc.). This emotional data influences the generation of the story and illustrations.
[1071] Step 3:
[1072] The device sends the prompt entered by the user and the analyzed emotion data to the server. This data is sent to the server as an HTTP request containing text data and emotion data. In order to send the input data (prompt and emotion data), the data is encoded and sent in this step.
[1073] Step 4:
[1074] The server parses the received prompt. It uses natural language processing algorithms (e.g., spaCy or NLTK) to extract important key phrases from the prompt. The extracted key phrases are used as the basis for narrative generation. This step involves text analysis and key phrase extraction.
[1075] Step 5:
[1076] The server launches a story generation engine based on the analysis results and emotional data. It uses a generative AI model (such as GPT-3) to generate story components (introduction, development, climax, and conclusion) step by step, and adjusts the content and tone of the story according to the emotional data. For example, if the user is excited about the prompt "Pirate adventure," a story with an emphasis on action and adventure will be generated.
[1077] Step 6:
[1078] The server runs an AI image generation algorithm (such as DeepArt or DALL-E) based on the generated story. The color and style of the generated illustration are adjusted according to the emotional data. For example, if the user is happy, the illustration will be colorful and bright. In this step, the process of generating visual content based on the story text is carried out.
[1079] Step 7:
[1080] The server combines the generated story and illustrations to create the pages of the digital picture book. Using software such as PIL (Python Imaging Library), the story and illustrations are laid out and beautiful page designs are created. In this step, the page layout is determined and content is integrated.
[1081] Step 8:
[1082] The server compresses the data of the completed digital picture book and sends it to the terminal as an HTTP response. The generated picture book data is provided in a compressed format (e.g., ZIP format) so that the user can download it. In this step, data compression and communication take place.
[1083] Step 9:
[1084] The device decompresses the received data and displays it to the user through a dedicated viewer application. The user can view the digital picture book using this viewer. It also provides a printing option if necessary. In this step, the data is decompressed and displayed.
[1085] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1086] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1087] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1088] [Fourth embodiment]
[1089] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1090] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1091] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1092] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1093] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1094] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1095] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1096] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1097] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1098] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1099] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1100] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1101] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1102] This invention relates to a system that receives prompts from a user indicating the direction of a story, generates a story based on an analysis of the prompts, and automatically generates illustrations corresponding to the story. A specific embodiment of this system is described below.
[1103] 1. User prompt input
[1104] The terminal provides an interface that accepts prompt input from the user, for example, by displaying an input form on the terminal's display, allowing the user to enter text about themes, characters, settings, etc.
[1105] Example: User types "Pirate Adventures", "Captain Rose", and "Treasure Hunt".
[1106] 2. Sending a prompt
[1107] The terminal sends the prompt entered by the user to the server. Specifically, it sends the entered text data to the server as an HTTP request.
[1108] 3. Parsing the prompt
[1109] The server parses the prompts it receives, using natural language processing algorithms to extract key phrases from the input data (e.g., "pirate," "adventure," "treasure hunt," "Captain Rose," etc.).
[1110] 4. Story Generation
[1111] The server uses a story generation engine to build a story based on the analysis results. For example, it generates a story about Captain Rose and his companions' adventure in search of lost treasure. In this process, the components of the story (introduction, development, climax, and conclusion) are generated step by step.
[1112] 5. Illustration Generation
[1113] Based on the generated story, the server uses an AI image generation algorithm to create illustrations, such as scenes of Captain Rose and her ship, treasure maps, and seascapes.
[1114] 6. Picture book structure
[1115] The server combines the generated stories and illustrations to create a digital picture book, and determines the page layout by appropriately placing the story and illustrations on each page.
[1116] 7. Sending picture books
[1117] The server sends the completed digital picture book data to the device. Specifically, it compresses the generated picture book data and sends it to the device in an HTTP response.
[1118] 8. Provision to Users
[1119] The device then provides the user with the digital data of the picture book, which is then displayed in a format that the user can view through a dedicated viewer application. It also provides the option to print the book if necessary.
[1120] The above is a concrete example for carrying out the invention. This system can quickly generate and provide a personalized picture book that matches a child's preferences and mood, making story time a more interactive and enjoyable experience.
[1121] The processing flow will be explained below.
[1122] Step 1:
[1123] The user uses the terminal to input prompts such as the theme and characters of the story. Specifically, the user enters text such as "Pirate Adventure," "Captain Rose," or "Treasure Hunt" into an input form displayed on the terminal.
[1124] Step 2:
[1125] The terminal receives the prompt entered by the user and sends it to the server. Specifically, the terminal sends the prompt to the server as an HTTP request.
[1126] Step 3:
[1127] The server parses the received prompt and uses natural language processing algorithms to extract important key phrases from the prompt (e.g., pirate, adventure, treasure hunt, Captain Rose).
[1128] Step 4:
[1129] The server starts a story generation engine based on the extracted key phrases, which generates story components such as introduction, development, climax, and conclusion step by step.
[1130] Step 5:
[1131] Based on the generated story, the server runs an illustration generation algorithm, which creates scenes of Captain Rose, her ship, treasure maps, seascapes, and more.
[1132] Step 6:
[1133] The server combines the generated story and illustrations to create the pages of the digital picture book. It then arranges the text and illustrations on each page according to the story's development and determines the page layout.
[1134] Step 7:
[1135] The server compresses the completed digital picture book data and sends it to the terminal as an HTTP response. Specifically, the data for each page of the picture book is consolidated and provided in a compressed format.
[1136] Step 8:
[1137] The device unpacks the received data and displays it to the user through a dedicated viewer application, allowing the user to view the picture book. It also provides a printing option if desired.
[1138] The system allows users to enjoy personalized stories and illustrations in a short amount of time, making storytime a more interactive and enjoyable experience.
[1139] Example 1
[1140] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1141] Traditional methods for providing personalized stories and illustrations to today's children are time-consuming and labor-intensive, making it difficult to instantly generate content that matches a child's preferences. This is especially true when providing interactive and intuitive digital picture books, making it difficult for parents and educators to provide individually customized picture books for children.
[1142] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1143] In this invention, the server includes means for receiving prompts from a user indicating the direction of a story, means for analyzing the prompts and extracting key information, means for generating a story based on the key information, means for automatically generating illustrations corresponding to the story, means for combining the story and the illustrations to create a digital picture book, means for providing the digital picture book to the user, means for compressing and transmitting the digital picture book data to be provided to the user, and means for providing an interface to be displayed on the user's terminal. This enables a user to easily and instantly generate and provide a personalized digital picture book tailored to the preferences of their child.
[1144] A "prompt" refers to a narrative direction text input received from a user.
[1145] "Parsing" refers to the process of extracting important key phrases and information from the entered prompt.
[1146] "Key information" refers to important phrases and words necessary for story generation that are extracted through prompt analysis.
[1147] "Narrative generation" refers to the process of constructing a coherent story based on analyzed key information.
[1148] "Illustration generation" refers to the process of automatically creating visual images that correspond to a generated story.
[1149] A "digital picture book" refers to a picture book in electronic format that combines a story and illustrations.
[1150] "Compression" refers to data processing performed to reduce the volume of digital picture book data.
[1151] "Transmission" refers to the act of transferring compressed digital picture book data to the user's terminal.
[1152] "Interface" refers to tools such as screens and input forms that allow users to interact with a system.
[1153] "Terminal" refers to a device through which a user can enter prompts and view the generated digital picture book.
[1154] "Server" refers to the computer system that analyzes prompts, generates stories and illustrations, and assembles and transmits the digital picture book.
[1155] This invention relates to a system that receives prompts from a user indicating the direction of a story, generates a story based on an analysis of the prompts, and automatically generates illustrations corresponding to the story. A specific embodiment of this system will be described below.
[1156] First, the user uses a terminal to input a prompt. The terminal provides an interface for accepting prompt input from the user. The interface displays an input form on the display, allowing the user to input the theme, characters, setting, etc. in text. For example, the user might input prompts such as "Pirate Adventure," "Captain Rose," or "Treasure Hunt."
[1157] Next, the terminal transmits the input prompt to the server. Specifically, the terminal transmits the prompt as text data in the form of an HTTP request to the server.
[1158] The server analyzes the received prompts using a natural language processing algorithm (e.g., GPT-4), extracts important key phrases from the prompts, and stores them as data points needed to generate the story.
[1159] Next, the server uses a story generation engine to generate a story based on the analysis results. This engine assembles the story's components (introduction, development, climax, and conclusion) step by step. For example, it generates a story about Captain Rose and his companions' adventure in search of lost treasure.
[1160] The server then uses an AI image generation algorithm (e.g., DALL-E) to automatically generate illustrations based on the generated story. Visuals of each character and scene are automatically drawn, such as Captain Rose and her ship, treasure maps, and seascapes.
[1161] The server then combines the generated story and illustrations to create a digital picture book. Using a page layout tool (e.g., Adobe InDesign API), the server arranges the text and illustrations appropriately, creating each page of the digital picture book.
[1162] The completed digital picture book is compressed and sent to the device. The server compresses the generated picture book data and sends it to the device as an HTTP response.
[1163] Finally, the device provides the received digital picture book to the user through a dedicated viewer application, allowing the user to view the book and, if necessary, use the print function.
[1164] As described above, the present invention enables users to easily and instantly create and provide personalized digital picture books that match the preferences of their children.
[1165] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1166] Step 1:
[1167] The user inputs prompts using a terminal. The terminal displays an input form on the display, allowing the user to input themes, characters, settings, etc. as text. The prompts generated by the input are treated as text data.
[1168] Input: User types "Pirate Adventures", "Captain Rose", and "Treasure Hunt".
[1169] Output: Text data of the prompt entered on the terminal.
[1170] Step 2:
[1171] The terminal sends the entered prompt to the server. Specifically, the prompt is converted into JSON format as text data and sent to the server in the form of an HTTP POST request.
[1172] Input: The text data of the prompt entered on the terminal.
[1173] Output: Text data as an HTTP request sent to the server.
[1174] Step 3:
[1175] The server parses the received prompt and uses a natural language processing algorithm (e.g., GPT-4) to extract important key phrases from within the prompt.
[1176] Input: The prompt text data sent to the server as an HTTP request.
[1177] Output: Key phrases parsed from the prompt (e.g. "pirates", "adventure", "treasure hunt", "Captain Rose").
[1178] Step 4:
[1179] The server uses a story generation engine (e.g., OpenAI model) to construct a story based on the analysis results. The story generation engine gradually assembles the components of the story (introduction, development, climax, and conclusion).
[1180] Input: Key phrase (e.g. "pirate", "adventure", "treasure hunt", "Captain Rose").
[1181] Output: A completed story (e.g., Captain Rose's quest for lost treasure).
[1182] Step 5:
[1183] Based on the generated story, the server uses an AI image generation algorithm (e.g., DALL-E) to create illustrations, generating visuals that correspond to prompts and scenes in the story.
[1184] Input: Narrative text.
[1185] Output: Illustrations for each scene in the story (e.g. Captain Rose and her ship, a treasure map, a seascape).
[1186] Step 6:
[1187] The server combines the generated story and illustrations to create a digital picture book. It uses a page layout tool (e.g., Adobe InDesign API) to arrange the text and illustrations appropriately.
[1188] Input: narrative text and corresponding illustrations.
[1189] Output: Completed digital picture book data.
[1190] Step 7:
[1191] The server compresses the data of the completed digital picture book and sends it to the device. The data is compressed in ZIP format and sent to the device as an HTTP response.
[1192] Input: Completed digital picture book data.
[1193] Output: HTTP response containing the digital picture book data compressed in ZIP format.
[1194] Step 8:
[1195] The device then provides the received digital picture book to the user through a dedicated viewer application, allowing the user to view the book and, if necessary, use the print function.
[1196] Input: HTTP response containing the digital picture book data compressed in ZIP format.
[1197] Output: A digital picture book displayed in a dedicated viewer application, and a printed picture book if desired.
[1198] (Application example 1)
[1199] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1200] Conventional digital content generation systems have issues with being unable to adequately generate stories based on user intent or automatically generate corresponding visual illustrations. Furthermore, it is difficult to personalize the generated content, preventing an improved user experience. In particular, there has been a lack of systems that can be easily applied to a variety of devices, such as smartphones, smart glasses, head-mounted displays, and robots.
[1201] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1202] In this invention, the server includes means for receiving prompts from a user indicating the direction of a story, means for analyzing the prompts and extracting key information, means for generating a story based on the key information, means for automatically generating illustrations corresponding to the story, means for combining the story and the illustrations to form digital content, means for providing the digital content to a user, and means for providing the digital content as an application installed on a smartphone, smart glasses, a head-mounted display, or a robot. This enables the automatic generation of a story and visual illustrations based on the user's prompts, and further enables personalization to suit the preferences of each individual user.
[1203] A "prompt indicating the direction of the story" is text data entered by the user, and includes information such as the theme, characters, and setting of the story.
[1204] "Means for parsing prompts" refers to natural language processing algorithms that use generative AI models to extract key phrases and important information from user prompts.
[1205] The "means for generating a story" refers to a story generation engine that generates the elements of a story (introduction, development, climax, conclusion, etc.) step by step based on the extracted key information.
[1206] "Means for automatically generating illustrations" refers to a method of generating visual illustrations using an AI image generation algorithm based on the generated story.
[1207] "Means of constructing digital content" refers to the method of combining the generated story and illustrations to create a single unified piece of digital content (e.g., a digital picture book), including page layout and design elements.
[1208] The "means for providing digital content to a user" refers to a method for compressing the completed digital content and transferring it to the user's device over a network.
[1209] An "application installed on a smartphone, smart glasses, head-mounted display, or robot" is software installed on a specific device that allows a user to view and interact with generated digital content.
[1210] "Personalization" refers to individually adjusting the content and design of generated content based on the user's preferences, mood, past usage history, etc.
[1211] This invention relates to a system that receives prompts from a user indicating the direction of a story, generates a story based on an analysis of the prompts, and automatically generates illustrations corresponding to the story. This system includes the steps of inputting a prompt from a user, analyzing the prompts, generating a story, automatically generating illustrations, and composing and providing digital content.
[1212] User prompt input
[1213] The user inputs the story's theme, characters, setting, etc. as text through an application installed on a smartphone, smart glasses, head-mounted display, or robot. For example, the user might input "adventure," "hero," or "fighting a dragon." This input interface utilizes a user interface and viewer application.
[1214] Sending and parsing prompts
[1215] User input is sent from the device to the server as an HTTP request. The server receives the prompt using a communication library such as the "requests" library and parses it using a generative AI model and a natural language processing algorithm (e.g., GPT-4). Key phrases (e.g., "adventure," "hero," and "dragon") are extracted from the input data.
[1216] Story Generation
[1217] The server runs a story generation engine based on the analysis results to generate a story. The generated story is then completed in stages, with components such as an introduction, development, climax, and conclusion.
[1218] Illustration generation
[1219] Based on the story, the server automatically generates illustrations using an AI image generation algorithm (e.g., Stable Diffusion). For example, it depicts a scene of a hero fighting a dragon or an adventure scene. These generated illustrations are placed on each page of the story.
[1220] Digital content configuration
[1221] The generated stories and illustrations are combined to form digital storybooks and other digital content, which are organized into a unified format, including page layout and design elements.
[1222] Providing digital content
[1223] The server then sends the completed digital content to the device, where it is compressed and returned as an HTTP response. The device then receives this data and displays it in a dedicated viewer application, allowing the user to browse the content. Additionally, the device offers the option to personalize the content based on the user's preferences and mood.
[1224] Prompt Sentence Examples
[1225] "Adventure Hero Fights Dragon"
[1226] This allows the system to enable users to easily create personalized stories and corresponding illustrations that can be enjoyed on a variety of devices.
[1227] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1228] Step 1:
[1229] The user inputs story prompts using an application installed on a smartphone, smart glasses, a head-mounted display, or a robot. The user inputs themes, characters, and settings in text format, such as "adventure," "hero," or "fighting a dragon." The input data is then saved on the device via the user interface.
[1230] Input: Text data such as story theme, characters, and setting
[1231] Output: Input data saved on the device
[1232] Step 2:
[1233] The terminal sends the entered prompt to the server, sends the prompt text data as an HTTP request to the server, and adds appropriate header information, and the sent data is temporarily stored on the server.
[1234] Input: Input data stored on the device
[1235] Output: Data sent to the server
[1236] Step 3:
[1237] The server analyzes the received prompt and uses a natural language processing algorithm (e.g., GPT-4) to extract keywords (e.g., "adventure," "hero," "dragon") from the prompt. The analysis results are used as key information for generating the story.
[1238] Input: Data sent to the server
[1239] Output: Parsed key information
[1240] Step 4:
[1241] The server then uses the analysis results to run a story generation engine to generate a story. Using a generative AI model, it creates a complete story, including components such as an introduction, development, climax, and conclusion. During this process, keywords entered by the user are reflected in each part of the story.
[1242] Input: Parsed key information
[1243] Output: Generated narrative text
[1244] Step 5:
[1245] The server automatically generates illustrations based on the generated story using an AI image generation algorithm (e.g., Stable Diffusion). An image corresponding to each story scene is generated. The generated illustrations are stored on the server.
[1246] Input: Generated narrative text
[1247] Output: Generated illustration image
[1248] Step 6:
[1249] The server then combines the generated story text and illustrations to create digital content. It then creates a digital picture book by designing a page layout and appropriately placing illustrations corresponding to the story on each page.
[1250] Input: Generated story text and illustration images
[1251] Output: Digital picture book data
[1252] Step 7:
[1253] The server sends the completed digital picture book data to the device. The data is compressed and returned as an HTTP response. The device saves the received digital picture book data.
[1254] Input: Digital picture book data
[1255] Output: Digital picture book data saved on the device
[1256] Step 8:
[1257] The device provides the received digital picture book data to the user. Using a dedicated viewer application, the device displays the digital picture book so that the user can view it. If necessary, the device can also personalize the book to suit the user's preferences and mood.
[1258] Input: Digital picture book data stored on the device
[1259] Output: A digital picture book that can be viewed by the user
[1260] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1261] This invention relates to a system that receives prompts from a user indicating the direction of a story, generates a story based on an analysis of the prompts, and automatically generates illustrations corresponding to the story. Furthermore, it is equipped with a function that combines an emotion engine that recognizes the user's emotions and adjusts the content of the generated story and illustrations according to the user's emotions. A specific embodiment of this system is described below.
[1262] 1. User Emotion Recognition
[1263] When the user inputs a prompt, the device uses an emotion engine to analyze the user's emotions. Specifically, it uses facial recognition and voice analysis technologies to capture the user's emotional state (happiness, sadness, excitement, etc.).
[1264] 2. User prompt input
[1265] The terminal provides an interface that accepts prompt input from the user, a process in which the user inputs themes, characters, settings, etc.
[1266] Example: When a user types "pirate adventure," "Captain Rose," or "treasure hunt," the emotion engine analyzes the user's emotional state.
[1267] 3. Sending prompts and emotion data
[1268] The terminal transmits the prompt input by the user and the analyzed emotion data to the server. Specifically, the terminal transmits an HTTP request including the text data and the emotion data to the server.
[1269] 4. Prompt and Emotion Data Analysis
[1270] The server analyzes the received prompt, using natural language processing algorithms to extract key phrases from the prompt and analyze them based on sentiment data.
[1271] 5. Story Generation
[1272] The server then launches a story generation engine based on the analysis results and emotional data. It generates the story's components (introduction, development, climax, and conclusion) step by step, adjusting the content and tone of the story according to the emotional data.
[1273] Example: If the user is in an excited state, a story with an emphasis on action and adventure may be generated.
[1274] 6. Illustration Generation
[1275] The server runs an AI image generation algorithm based on the generated story, adjusting the color and style of the resulting illustration depending on the emotional data.
[1276] Example: If the user is happy, the illustration will be bright and colorful.
[1277] 7. Picture Book Structure
[1278] The server combines the generated stories and illustrations to create pages for the digital picture book, and determines the page layout by appropriately arranging the stories and illustrations on each page.
[1279] 8. Sending picture books
[1280] The server compresses the data of the completed digital picture book and sends it to the terminal as an HTTP response, providing the generated picture book data in a compressed format.
[1281] 9. Provision to Users
[1282] The device unpacks the received data and displays it to the user through a dedicated viewer application, allowing the user to view the picture book. It also provides a printing option if desired.
[1283] This system allows users to enjoy personalized stories and illustrations based on their emotional state, making storytime a more interactive and enjoyable experience.
[1284] The processing flow will be explained below.
[1285] Step 1:
[1286] The user uses the terminal to input prompts such as the theme and characters of the story. Specifically, the user enters text such as "Pirate Adventure," "Captain Rose," and "Treasure Hunt" into an input form displayed on the terminal.
[1287] Step 2:
[1288] The device activates an emotion engine to analyze the user's emotions while the user is entering prompts. Specifically, the device's camera and microphone are used to capture the user's emotional state (e.g., joy, sadness, excitement) in real time using facial recognition or voice analysis technology.
[1289] Step 3:
[1290] The device sends the prompt entered by the user and the analyzed emotion data to the server. Specifically, it sends an HTTP request including the text data and emotion data to the server.
[1291] Step 4:
[1292] The server analyzes the received prompt and uses natural language processing algorithms to extract key phrases from the prompt, while also analyzing emotional data to recognize the user's emotional state.
[1293] Step 5:
[1294] The server launches a story generation engine based on the extracted key phrases and emotional data. It generates the elements of the story (introduction, development, climax, and conclusion) and adjusts the content and tone of the story according to the user's emotional data. Specifically, if the user is excited, a story with an emphasis on action and adventure will be generated.
[1295] Step 6:
[1296] The server runs an AI image generation algorithm based on the generated story. The AI adjusts the color and style of illustrations related to the scenario (e.g., Captain Rose, her ship, treasure map, seascape) according to the user's emotional data. Specifically, if the user is happy, illustrations with bright colors and tones are generated.
[1297] Step 7:
[1298] The server combines the generated story and illustrations to create pages of a digital picture book. It determines the page layout by appropriately arranging the story text and illustrations on each page.
[1299] Step 8:
[1300] The server compresses the completed digital picture book data and sends it to the terminal as an HTTP response. Specifically, the data for each page of the picture book is consolidated and provided in a compressed format.
[1301] Step 9:
[1302] The device unpacks the received data and displays it to the user through a dedicated viewer application, allowing the user to view the picture book. It also provides a printing option if desired.
[1303] This system allows users to enjoy personalized stories and illustrations based on their emotional state, making storytime a more interactive and enjoyable experience.
[1304] Example 2
[1305] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1306] Conventional story generation systems lack personalization based on user emotions and preferences, resulting in uniform generated content and low user satisfaction. Furthermore, the story and illustration generation processes are independent, resulting in a lack of consistency between the two. Furthermore, the lack of a way to reflect user emotions in real time compromises the interactive experience.
[1307] The specification process by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving a prompt from the user indicating the direction of the story, means for using a natural language processing algorithm to analyze the prompt and extract key information, emotion recognition means for analyzing the user's emotions, means for generating a story based on the key information and the emotions, means for automatically generating illustrations corresponding to the story in a color tone and style according to the user's emotions, means for combining the story and the illustrations to create a digital picture book, and means for providing the digital picture book to the user. This makes it possible to consistently provide stories and illustrations personalized according to the user's emotions and preferences.
[1308] A "user" is someone who utilizes the system to provide prompts that direct the story.
[1309] A "prompt" is information that indicates the direction of the story's theme, characters, setting, etc., entered by the user.
[1310] A "natural language processing algorithm" is a technology that analyzes text data and extracts meaning and key information from it.
[1311] "Emotion recognition means" is a technology that analyzes a user's facial expressions and voice and classifies the user's emotional state.
[1312] A "story generator" is an algorithm that generates story components based on prompt and emotion data.
[1313] The "illustration generation means" is a technology that automatically generates illustrations in a color tone and style that corresponds to the user's emotions based on the generated story.
[1314] A "digital picture book" is a digital picture book that combines generated stories and illustrations.
[1315] The "server" is an information processing device that acts as the backend of this system, analyzing prompts, recognizing emotions, generating stories and illustrations, and providing digital picture books.
[1316] This invention relates to a system that receives prompts from a user indicating the direction of a story, generates a story based on an analysis of the prompts, and automatically generates illustrations corresponding to the story. Furthermore, this system is equipped with an emotion engine that recognizes the user's emotions, and has the function of adjusting the content of the generated story and illustrations according to the user's emotions.
[1317] Hardware and software used
[1318] Device: A device on which a user can input prompts and view the generated digital picture book. Examples include a PC, tablet, or smartphone.
[1319] Server: A central information processing device that analyzes data, generates stories and illustrations, and creates and provides digital picture books. Specifically, it includes cloud servers and local servers.
[1320] Emotion recognition technology: Facial recognition technology (e.g., OpenCV) and voice analysis technology (e.g., Google Speech-to-Text API) are used.
[1321] Natural language processing algorithms: Algorithms such as spaCy and BERT are used to analyze text data and extract key information.
[1322] Story generation engine: A generative AI model such as GPT-3 is used to generate a story based on prompts and sentiment data.
[1323] Illustration generation algorithm: AI image generation algorithms (e.g., DALL-E, MidJourney) are used.
[1324] Examples of data processing and data calculation
[1325] 1. User Emotion Recognition
[1326] When the user inputs a prompt, the device activates the emotion engine, collects the user's facial expressions and voice through the camera and microphone, and uses facial recognition and voice analysis technology to analyze the user's emotions (happiness, sadness, excitement, etc.) in real time.
[1327] 2. User prompt input
[1328] The device provides an interface for the user to enter prompts that provide direction to the story, such as text boxes and selection menus that allow the user to enter theme, characters, and setting.
[1329] Specific examples
[1330] When the user inputs prompts such as "Pirate Adventure," "Captain Rose," or "Treasure Hunt," the device uses its emotion engine to analyze the user's emotional state and recognize, for example, that the user is excited.
[1331] 3. Sending prompts and emotion data
[1332] The device sends the input prompt and the analyzed emotion data to the server as an HTTP POST request, which includes both text data and emotion data.
[1333] 4. Prompt and Emotion Data Analysis
[1334] The server uses natural language processing algorithms to analyze the received prompts and extract important key information, while also expanding and adjusting the meaning of the prompts based on emotional data.
[1335] 5. Story Generation
[1336] The server then activates a story generation engine based on the analyzed prompt data and emotion data to generate a story, the content and tone of which are adjusted according to the user's emotional state.
[1337] Specific examples
[1338] If the user is in an excited state, the generated story will emphasize elements of action and adventure.
[1339] 6. Illustration Generation
[1340] The server then runs an AI image generation algorithm based on the generated story, generating illustrations with adjusted color tone and style depending on the emotional data.
[1341] Specific examples
[1342] If the user is happy, the illustrations generated will be colorful and bright in tone.
[1343] 7. Digital Picture Book Structure
[1344] The server combines the generated story and illustrations to create pages for a digital picture book. The page layout is adjusted to make it easy to read by adjusting font size and line spacing.
[1345] 8. Sending digital books
[1346] The server sends the completed digital picture book to the terminal in a compressed format, such as ZIP.
[1347] 9. Provision to Users
[1348] The device decompresses the received data and displays it to the user through a dedicated viewer, where the user can view the picture book and print it if necessary.
[1349] Through the above process, users can enjoy a digital picture book with stories and illustrations personalized according to their emotions and preferences.
[1350] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1351] Step 1: Recognizing user emotions
[1352] Input: User's facial expression data, voice data
[1353] Specific operation: The device collects the user's facial expressions and voice through the camera and microphone. Using facial recognition technology (e.g., OpenCV) and voice analysis technology (e.g., Google Speech-to-Text API), it analyzes the user's emotional state (e.g., joy, sadness, excitement) in real time. The analysis results are recorded as emotion data.
[1354] Output: User's emotional state (e.g., excitement level)
[1355] Step 2: User prompt input
[1356] Input: User-provided narrative prompts (theme, characters, setting)
[1357] Specific behavior: The device displays an interface (text boxes and / or selection menus) for the user to enter a prompt. The user enters a prompt that indicates the direction of the story they want to take.
[1358] Output: Prompt data (e.g. "Pirate Adventure", "Captain Rose", "Treasure Hunt")
[1359] Step 3: Sending prompts and emotion data
[1360] Input: prompt data, emotion data
[1361] How it works: The device sends the entered prompt and analyzed emotion data to the server as an HTTP POST request. This communication is performed using the secure HTTPS protocol.
[1362] Output: HTTP POST request (including prompt and emotion data)
[1363] Step 4: Analyze prompts and sentiment data
[1364] Input: HTTP POST request (prompt data, emotion data)
[1365] What happens: The server parses the incoming request, uses natural language processing algorithms (e.g., spaCy, BERT) to extract key phrases from the prompt, and optimizes the prompt content based on sentiment data.
[1366] Output: Parsed key information, prompt data with sentiment information
[1367] Step 5: Story Generation
[1368] Input: Parsed key information, prompt data with emotional information
[1369] Specific operation: The server launches a story generation engine (e.g., GPT-3) to generate a story based on the prompt and emotion data. It generates the story components (introduction, development, climax, and conclusion) step by step and adjusts them according to the user's emotions.
[1370] Output: Generated narrative data
[1371] Step 6: Illustration generation
[1372] Input: Generated story data, emotion information
[1373] How it works: The server runs an AI image generation algorithm (e.g., DALL-E, MidJourney) to generate illustrations based on the story content and the user's emotions. Color tone and style are adjusted depending on the emotion.
[1374] Output: Generated illustration data
[1375] Step 7: Composing your digital storybook
[1376] Input: Generated story data, generated illustration data
[1377] Specific operation: The server combines the story and illustrations to create the pages of the digital picture book. On each page, the server appropriately arranges the story text and corresponding illustrations and determines the page layout.
[1378] Output: Completed digital picture book data
[1379] Step 8: Submit your digital book
[1380] Input: Completed digital picture book data
[1381] Specific operation: The server compresses the completed digital picture book data and sends it to the terminal as an HTTP response. This data is often provided in ZIP format.
[1382] Output: HTTP response (compressed digital picture book data)
[1383] Step 9: Provide to users
[1384] Input: HTTP response (compressed digital picture book data)
[1385] Specific operation: The device decompresses the received data and displays it to the user through a dedicated viewer application. The user can view the picture book with this viewer and print it if necessary.
[1386] Output: A digital storybook that users can view
[1387] (Application example 2)
[1388] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1389] Conventional digital picture book generation systems could generate stories and illustrations based on user input, but they were unable to personalize this content based on the user's emotional state. This made it difficult to provide the optimal reading experience for users. New methods were needed to increase user satisfaction, especially for subjects such as children, whose emotions are strongly influenced by the content.
[1390] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1391] In this invention, the server includes means for receiving prompts from a user indicating the direction of the story, means for analyzing the prompts and extracting key information, means for generating a story based on the key information, means for automatically generating illustrations corresponding to the story, means for combining the story and the illustrations to create a digital picture book, means for providing the digital picture book to the user, means for analyzing the user's emotions, and means for adjusting the content of the story and illustrations based on the analyzed emotional data. This allows the story and illustrations to be personalized according to the user's emotional state, enabling a more satisfying interactive story-telling experience.
[1392] A "user" is someone using a terminal to enter story prompts.
[1393] A "prompt" is a string or sentence that the user enters to indicate the direction or theme of the story.
[1394] "Key information" is important information necessary for generating a story that is extracted by analyzing a prompt.
[1395] "Narrative generation" is the process of constructing the content of a story based on key information.
[1396] "Automatic illustration generation" is the act of automatically generating visual content corresponding to a generated story using technologies such as AI.
[1397] A "digital picture book" is an electronic picture book that combines generated stories and illustrations.
[1398] "Emotion analysis" is the process of analyzing a user's emotional state using facial recognition and voice analysis techniques.
[1399] "Emotion data" is data that represents the emotional state of a user obtained by emotion analysis.
[1400] "Personalization" means tailoring content to a user based on their preferences and emotional state.
[1401] A "natural language processing algorithm" is a computer algorithm that analyzes text data and understands its meaning and context.
[1402] This invention is a system that generates a story based on prompts from a user and generates illustrations corresponding to that story. It also includes a function to recognize the user's emotions and adjust the content of the story and illustrations accordingly. A specific embodiment of this system is described below.
[1403] 1. User prompt input
[1404] The user uses a terminal to input prompts that indicate the direction of the story. The input prompts are text data that includes the theme, characters, setting, etc. of the story. Examples of prompts include "Pirate Adventure," "Captain Rose," and "Treasure Hunt."
[1405] 2. Emotion analysis
[1406] The server uses an emotion engine to analyze the user's emotions when the user enters a prompt. The analysis uses facial recognition and voice analysis technologies to capture the user's emotional state (happiness, sadness, excitement, etc.). The emotion engine can use a facial recognition library (OpenCV) or a voice analysis library (Google Cloud Speech-to-Text API).
[1407] 3. Sending prompts and emotion data
[1408] The terminal transmits the prompt input by the user and the analyzed emotion data to the server, which then transmits the data as an HTTP request including text data and emotion data.
[1409] 4. Prompt and Emotion Data Analysis
[1410] The server analyzes the received prompts. It uses natural language processing algorithms to extract key phrases from the prompts and analyze them based on sentiment data. For natural language processing, open source NLP tools (e.g., spaCy or NLTK) can be used.
[1411] 5. Story Generation
[1412] The server launches a story generation engine based on the analysis results and emotional data. It generates the story's components (introduction, development, climax, and conclusion) step by step, adjusting the content and tone of the story according to the emotional data. A generative AI model (such as GPT-3) is used to generate the story.
[1413] 6. Illustration Generation
[1414] The server runs an AI image generation algorithm based on the generated story. The color and style of the generated illustration are adjusted according to the emotional data. AI image generation can use technologies such as DeepArt and DALL-E.
[1415] 7. Digital Picture Book Structure
[1416] The server combines the generated story and illustrations to create the pages of the digital picture book. It then appropriately arranges the story and illustrations on each page and determines the page layout. Software used includes PIL (Python Imaging Library).
[1417] 8. Sending digital books
[1418] The server compresses the completed digital picture book data and sends it to the terminal as an HTTP response. The generated picture book data is provided in a compressed format (e.g., ZIP format).
[1419] 9. Provision to Users
[1420] The device unpacks the received data and displays it to the user through a dedicated viewer application, allowing the user to view the digital picture book. It also provides a printing option if desired.
[1421] This system allows users to enjoy personalized stories and illustrations based on their emotional state, making story time a more interactive and enjoyable experience.
[1422] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1423] Step 1:
[1424] The user uses a terminal to input prompts that indicate the direction of the story. The input prompts are text data that includes the theme, characters, and setting of the story. Examples of inputs include "Pirate Adventure," "Captain Rose," and "Treasure Hunt." This input data is used for subsequent analysis.
[1425] Step 2:
[1426] The device uses an emotion engine to analyze the user's emotions as they input prompts. It uses facial recognition technology (e.g., OpenCV) and voice analysis technology (e.g., Google Cloud Speech-to-Text API) to capture the user's emotional state (e.g., joy, sadness, excitement, etc.). This emotional data influences the generation of the story and illustrations.
[1427] Step 3:
[1428] The device sends the prompt entered by the user and the analyzed emotion data to the server. This data is sent to the server as an HTTP request containing text data and emotion data. In order to send the input data (prompt and emotion data), the data is encoded and sent in this step.
[1429] Step 4:
[1430] The server parses the received prompt. It uses natural language processing algorithms (e.g., spaCy or NLTK) to extract important key phrases from the prompt. The extracted key phrases are used as the basis for narrative generation. This step involves text analysis and key phrase extraction.
[1431] Step 5:
[1432] The server launches a story generation engine based on the analysis results and emotional data. It uses a generative AI model (such as GPT-3) to generate story components (introduction, development, climax, and conclusion) step by step, and adjusts the content and tone of the story according to the emotional data. For example, if the user is excited about the prompt "Pirate adventure," a story with an emphasis on action and adventure will be generated.
[1433] Step 6:
[1434] The server runs an AI image generation algorithm (such as DeepArt or DALL-E) based on the generated story. The color and style of the generated illustration are adjusted according to the emotional data. For example, if the user is happy, the illustration will be colorful and bright. In this step, the process of generating visual content based on the story text is carried out.
[1435] Step 7:
[1436] The server combines the generated story and illustrations to create the pages of the digital picture book. Using software such as PIL (Python Imaging Library), the story and illustrations are laid out and beautiful page designs are created. In this step, the page layout is determined and content is integrated.
[1437] Step 8:
[1438] The server compresses the data of the completed digital picture book and sends it to the terminal as an HTTP response. The generated picture book data is provided in a compressed format (e.g., ZIP format) so that the user can download it. In this step, data compression and communication take place.
[1439] Step 9:
[1440] The device decompresses the received data and displays it to the user through a dedicated viewer application. The user can view the digital picture book using this viewer. It also provides a printing option if necessary. In this step, the data is decompressed and displayed.
[1441] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1442] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1443] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1444] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1445] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1446] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1447] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1448] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1449] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1450] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1451] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1452] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1453] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1454] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1455] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1456] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1457] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1458] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1459] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1460] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1461] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1462] The following is further disclosed regarding the above embodiment.
[1463] (Claim 1)
[1464] means for accepting narrative direction prompts from a user;
[1465] means for analyzing the prompt to extract key information;
[1466] means for generating a story based on the key information;
[1467] means for automatically generating illustrations corresponding to the story;
[1468] A means for combining the story and the illustrations to create a digital picture book;
[1469] A means for providing the digital picture book to a user;
[1470] A system including:
[1471] (Claim 2)
[1472] 10. The system of claim 1, wherein the digital storybook is personalized to suit the child's preferences and moods.
[1473] (Claim 3)
[1474] 10. The system of claim 1, wherein a natural language processing algorithm is used to analyze the prompt.
[1475] "Example 1"
[1476] (Claim 1)
[1477] means for accepting narrative direction prompts from a user;
[1478] means for analyzing the prompt to extract key information;
[1479] means for generating a story based on the key information;
[1480] means for automatically generating illustrations corresponding to the story;
[1481] A means for combining the story and the illustrations to create a digital picture book;
[1482] A means for providing the digital picture book to a user;
[1483] A means for compressing and transmitting digital picture book data to be provided to a user;
[1484] means for providing an interface that is displayed on a user's terminal;
[1485] A system including:
[1486] (Claim 2)
[1487] 10. The system of claim 1, wherein the digital storybook is personalized to suit the child's preferences and moods.
[1488] (Claim 3)
[1489] 10. The system of claim 1, wherein a natural language processing algorithm is used to analyze the prompt.
[1490] "Application Example 1"
[1491] (Claim 1)
[1492] means for accepting narrative direction prompts from a user;
[1493] means for analyzing the prompt to extract key information;
[1494] means for generating a story based on the key information;
[1495] means for automatically generating illustrations corresponding to the story;
[1496] A means for combining the story and the illustrations to form digital content;
[1497] means for providing said digital content to a user;
[1498] means for providing the digital content as an application installed on a smartphone, smart glasses, a head-mounted display, or a robot;
[1499] A system including:
[1500] (Claim 2)
[1501] 10. The system of claim 1, wherein the digital content is personalized to a user's tastes and moods.
[1502] (Claim 3)
[1503] 10. The system of claim 1, wherein a natural language processing algorithm is used to analyze the prompt.
[1504] "Example 2: Combining Emotion Engines"
[1505] (Claim 1)
[1506] means for accepting narrative direction prompts from a user;
[1507] means for using natural language processing algorithms to analyze the prompt and extract key information;
[1508] emotion recognition means for analyzing the emotion of the user;
[1509] means for generating a story based on the key information and the emotions;
[1510] means for automatically generating illustrations corresponding to the story in a color tone and style that corresponds to the user's emotions;
[1511] A means for combining the story and the illustrations to create a digital picture book;
[1512] A means for providing the digital picture book to a user;
[1513] A system including:
[1514] (Claim 2)
[1515] 10. The system of claim 1, wherein the digital picture book is personalized to suit the user's preferences and feelings.
[1516] (Claim 3)
[1517] 2. The system according to claim 1, wherein face recognition technology and voice analysis technology are included in analyzing the user's emotions.
[1518] "Application example 2 when combining emotion engines"
[1519] (Claim 1)
[1520] means for accepting narrative direction prompts from a user;
[1521] means for analyzing the prompt to extract key information;
[1522] means for generating a story based on the key information;
[1523] means for automatically generating illustrations corresponding to the story;
[1524] A means for combining the story and the illustrations to create a digital picture book;
[1525] A means for providing the digital picture book to a user;
[1526] means for analyzing user emotions;
[1527] a means for adjusting the content of the story and illustrations based on the analyzed emotion data;
[1528] A system including:
[1529] (Claim 2)
[1530] 10. The system of claim 1, wherein the digital storybook is personalized to suit the child's preferences and moods.
[1531] (Claim 3)
[1532] 10. The system of claim 1, wherein a natural language processing algorithm is used to analyze the prompt. [Explanation of symbols]
[1533] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for accepting narrative direction prompts from a user; means for analyzing the prompt to extract key information; means for generating a story based on the key information; means for automatically generating illustrations corresponding to the story; A means for combining the story and the illustrations to create a digital picture book; A means for providing the digital picture book to a user; A system including:
2. 10. The system of claim 1, wherein the digital storybook is personalized to suit a child's preferences and moods.
3. 10. The system of claim 1, wherein a natural language processing algorithm is used to analyze the prompt.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A