system

The system addresses the challenge of providing diverse and educational picture books by allowing users to select themes, generate original stories and images using AI, and output PDFs, offering a cost-effective and engaging solution for parents.

JP2026016179APending Publication Date: 2026-02-03SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024117269
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-22
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Parents face challenges in providing diverse and educational picture books to children without incurring high costs or spending time creating new stories, as existing picture books can become boring and are difficult to personalize.

Method used

A system that allows users to select a theme, generates an original story and images using AI, and outputs the content in PDF format for downloading or printing, optionally incorporating an emotion engine to tailor the content to the user's emotional state.

Benefits of technology

Enables parents to provide new stories and educational content to children daily, reducing financial burden and enhancing engagement through personalized and emotionally relevant picture books.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026016179000001_ABST
    Figure 2026016179000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for selecting a theme by a user; means for generating an original story based on the selected theme; means for generating an image based on the generated story; means for converting the generated story and image into a PDF format; and means for downloading or printing the picture book in the PDF format by the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Parents who read aloud to their children often face problems such as the high cost of picture books, not being able to frequently buy more, children getting bored with the same stories, and the difficulty of coming up with new stories on their own. The purpose of this invention is to solve these problems and provide a new story every night without spending time or money. [Means for solving the problem]

[0005] The present invention is a system that includes a means for a user to select a theme, a means for generating an original story based on the selected theme, a means for generating images based on the generated story, a means for converting the generated story and images into PDF format, and a means for the user to download or print the PDF picture book. Furthermore, if the theme includes an educational message or the generated PDF picture book includes a multi-page format, parents can provide a new story every night without spending time or money, satisfying their child's curiosity and providing an effective education.

[0006] "User" refers to an individual who uses the system or an entity that operates a terminal.

[0007] "Themes" refer to selectable categories or topics that serve as the basis for generating original stories and images.

[0008] A "story" refers to a series of sentences or narrative structures generated based on a theme.

[0009] "Image" refers to a visual illustration related to the story.

[0010] "PDF format" is an abbreviation for Portable Document Format, and refers to a file format that converts documents and images into a fixed layout.

[0011] "Download" refers to the process of getting a file from a server to a device.

[0012] "Printing" refers to the act of printing an electronic document onto physical paper or the like. [Brief explanation of the drawings]

[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2]1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0021] [First embodiment]

[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0034] This invention provides a system in which a user selects a theme for a picture book, AI generates an original story and images based on that theme, and the picture book can be downloaded or printed in PDF format. The program and processing of this system are explained in detail below.

[0035] System Overview

[0036] When a user accesses the "AI Storyteller" website using a device, a theme selection screen appears. The user selects the desired theme and presses the generate button. The AI ​​then generates an original story and images based on the selected theme. The generated story and images are then converted into PDF format, which the user can download or print.

[0037] Program processing flow

[0038] 1. When a user accesses a website using a device, the device sends a connection request to the server, and the server returns HTML content including a theme selection screen. The device displays the returned HTML content on the screen, allowing the user to select a theme.

[0039] 2. When the user selects a theme and clicks the generate button, the device sends the selected theme data as a POST request to the server. The server receives the POST request and generates an original story based on the selected theme using an AI model.

[0040] 3. The server uses an AI model to generate illustrations based on the generated story, and combines the generated story and illustrations to create a picture book in PDF format.

[0041] 4. The created PDF is saved as a temporary file by the server, and a response containing a link that the user can download is sent back to the device.

[0042] 5. When the user clicks the download link on their device, the PDF is saved to their device. The user can view, download, and print the saved PDF.

[0043] Specific examples

[0044] For example, if a user selects the theme "Dragon," the server generates a story based on this theme, such as "A brave boy makes friends with a dragon." The AI ​​then creates illustrations of the dragon and boy that fit the generated story. The server then combines these illustrations with the story to create a picture book in PDF format and provides the user with a download link. The user can click the link to save the PDF on their device and read it to their child before bedtime.

[0045] This system eliminates the need for users to frequently purchase expensive picture books, allowing them to provide their children with a new story every night. Furthermore, by setting themes with educational messages, it is possible to effectively educate children. In this way, the present invention alleviates parental concerns and provides an efficient means for supporting children's curiosity and learning.

[0046] The processing flow will be explained below.

[0047] Step 1:

[0048] A user accesses the "AI Storyteller" website using a device. The device sends a connection request to the web server, and the web server returns HTML content including a theme selection screen.

[0049] Step 2:

[0050] The device parses the HTML received from the server and displays the theme selection screen, allowing the user to see the theme selection form.

[0051] Step 3:

[0052] The user selects the desired theme from a drop-down menu, for example, one of the options "adventure," "dragon," or "spaceship."

[0053] Step 4:

[0054] When the user presses the "Generate" button, the form data is sent from the terminal to the server as a POST request, including the selected theme.

[0055] Step 5:

[0056] The server receives the POST request and retrieves the selected theme data from the request, which is then used for the next process.

[0057] Step 6:

[0058] The server uses an AI model to generate an original story based on the selected theme. The server calls an AI library (e.g., OpenAI GPT-3) and executes the story generation API.

[0059] Step 7:

[0060] The server generates images appropriate for the generated story. The server again uses an AI model (e.g., DALL-E) to execute an illustration generation API based on the content of the story.

[0061] Step 8:

[0062] The server creates a PDF book based on the generated story and images. It uses a PDF generation library (e.g., FPDF) to insert the story text and images into the PDF document.

[0063] Step 9:

[0064] The server saves the completed PDF as a temporary file and retains its path, which is later provided to the user.

[0065] Step 10:

[0066] The server sends a response back to the device containing the path to the generated PDF, along with a link to download the PDF.

[0067] Step 11:

[0068] The device parses the response received from the server and displays a PDF download link to the user, which the user can click to download the PDF.

[0069] Step 12:

[0070] When a user clicks the download link, the device saves the PDF file to local storage. By opening the saved PDF, the user can view, download, and print the generated original picture book.

[0071] Through this series of processes, users can easily obtain high-quality original picture books and provide their children with new stories every night.

[0072] Example 1

[0073] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0074] In modern child-rearing, parents want to provide their children with a variety of fresh picture books that meet their children's interests and educational needs. However, frequently purchasing new picture books is a significant financial burden, and commercially available picture books do not always contain the educational messages parents desire. Furthermore, busy parents find it difficult to quickly obtain personalized picture books tailored to their children's preferences. Efficient methods to solve these problems are needed.

[0075] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0076] In this invention, the server includes means for a user to select a theme, means for generating an original story using a generative AI model based on the selected theme, means for generating images using the generative AI model based on the content of the generated story, means for converting the generated story and images into a PDF format, and means for the user to download or print the PDF picture book. This allows a user to easily generate an original picture book based on a theme and download or print it.

[0077] A "theme" is a subject or topic selected by the user that will be the basis for the story or image that will be generated.

[0078] A "generative AI model" is an artificial intelligence algorithm or system that generates text or images based on a selected theme.

[0079] An "original story" is a unique narrative generated by a generative AI model based on a selected theme.

[0080] "Images" are visual illustrations or pictures generated by a generative AI model to match the content of the original story.

[0081] "PDF format" is an abbreviation for Portable Document Format, and is a file format for saving and displaying documents and images on a page-by-page basis.

[0082] A picture book is a book for children that contains stories and illustrations in written form.

[0083] "Downloading" is the act of saving a file from the Internet to your device.

[0084] "Printing" is the act of printing digital data onto physical media such as paper.

[0085] A "user" is someone who uses the system to select a theme, create, download, and print a picture book.

[0086] A "server" is a computer system that receives requests from users and generates, stores, and delivers appropriate content.

[0087] A "terminal" is a device operated by a user, and is a computer that communicates with a server and displays and saves results.

[0088] The present invention is a system that uses a generative AI model to generate original stories and images based on a theme selected by a user, and provides them as a PDF picture book. The system includes a means for a user to select a theme, a means for converting the generated stories and images into PDF format, and a means for the user to download or print the PDF picture book.

[0089] Hardware and Software

[0090] The system is implemented using a server, a terminal, a generative AI model, and various software. The server receives requests, generates stories and images, converts them into PDF format, and distributes them. The terminal acts as an interface for users to access the system, select a theme, and download the PDF.

[0091] The main hardware and software used are as follows:

[0092] Server: Processes the data and invokes the generative AI model.

[0093] Terminal: The device through which the user accesses the system, such as a regular PC, smartphone, or tablet.

[0094] Generative AI models: For example, ChatGPT and GPT-4 are used for story generation, while DALL-E and Stable Diffusion are used for image generation.

[0095] PDF generation library: Generate PDFs using Python's ReportLab, etc.

[0096] Data processing and calculation

[0097] 1. Theme selection: The user selects a theme using the terminal. For example, the user can select the theme "Dragon."

[0098] 2. Story generation: The server uses a generative AI model (such as ChatGPT or GPT-4) to generate an original story based on the selected theme. An example of a specific prompt is "Theme: Dragons. Please generate a story about a brave boy who becomes friends with a dragon."

[0099] 3. Image Generation: The server uses a generative AI model (such as DALL-E or Stable Diffusion) to generate images that match the generated story. The prompt might be something like, "Generate an illustration of a dragon and a boy."

[0100] 4. PDF generation: The server combines the generated story and images and generates a PDF picture book using Python's ReportLab or similar.

[0101] 5. PDF Delivery: The server saves the PDF in temporary storage and provides a download link that users can access, and they can download or print the PDF using their devices.

[0102] Specific examples

[0103] For example, if a user selects "Dragons" as the theme, the process goes as follows: The user selects a theme through the website and presses the generate button. The device then sends the theme data to the server. The server uses a generative AI model to generate a story such as "A brave boy befriends a dragon" and creates illustrations of the dragon and boy based on that story. The server then assembles this content into a PDF picture book using Python's ReportLab and provides the user with a download link. The user can click this link to download the PDF and read it to their child.

[0104] This system eliminates the need for users to frequently purchase expensive picture books, allowing them to provide their children with a new story every night. Furthermore, by setting themes with educational messages, it is possible to effectively educate children. In this way, the present invention alleviates parental concerns and provides an efficient means for supporting children's curiosity and learning.

[0105] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0106] Specific flow of system program processing

[0107] Step 1:

[0108] The user accesses the "AI Storyteller" website using a device.

[0109] Input: An HTTP request sent by a user typing a URL into a web browser.

[0110] Output: The device sends a connection request to the server, and the server returns a theme selection screen.

[0111] Specific behavior:

[0112] The user opens a browser, enters the specified URL, and presses enter.

[0113] The device sends this operation to the server as an HTTP request.

[0114] The server receives the request and generates HTML content including a theme selection screen.

[0115] The server returns the generated HTML content to the terminal as an HTTP response.

[0116] The terminal displays the received HTML content on the screen and allows the user to select a theme.

[0117] Step 2:

[0118] The user selects a theme and clicks the generate button.

[0119] Input: The user selects a theme and clicks the generate button.

[0120] Output: The device sends the selected theme data to the server as a POST request.

[0121] Specific behavior:

[0122] The user selects a desired theme from the displayed theme selection screen.

[0123] The user clicks the "Generate" button.

[0124] The device sends the user's selected theme data to the server as a POST request in JSON format.

[0125] Step 3:

[0126] The server receives the POST request and generates the story.

[0127] Input: A POST request containing theme data sent from the device.

[0128] Output: The original story generated using the generative AI model.

[0129] Specific behavior:

[0130] The server receives the POST request and extracts the theme data from the request body.

[0131] The server sends a prompt to the generative AI model (e.g., ChatGPT or GPT-4). Example prompt: "Theme: Dragons. Generate a story about a brave boy who becomes friends with a dragon."

[0132] The generative AI model generates a story based on the prompt sentence and returns the story to the server.

[0133] Step 4:

[0134] The server generates illustrations based on the story.

[0135] Input: Generated stories.

[0136] Output: An illustration generated using a generative AI model.

[0137] Specific behavior:

[0138] The server analyzes the content of the generated story and generates a prompt accordingly. Example prompt: "Please generate an illustration of a dragon and a boy."

[0139] The server sends a prompt to a generative AI model (e.g., DALL-E or Stable Diffusion).

[0140] The generative AI model generates an illustration based on the prompt sentence and returns the illustration to the server.

[0141] Step 5:

[0142] The server creates a picture book in PDF format.

[0143] Input: Generated story and illustrations.

[0144] Output: Picture book in PDF format.

[0145] Specific behavior:

[0146] The server combines the generated story and illustrations and creates a PDF picture book using Python's ReportLab etc.

[0147] The server saves the generated PDF file in temporary storage.

[0148] Step 6:

[0149] The server provides a download link for the PDF.

[0150] Input: PDF file stored on the server.

[0151] Output: A response containing a download link for the PDF.

[0152] Specific behavior:

[0153] The server generates a download link for the PDF file stored in temporary storage.

[0154] The server sends a response including the download link back to the device.

[0155] Step 7:

[0156] The user downloads or prints the PDF.

[0157] Input: Click on the download link displayed on your device.

[0158] Output: Downloaded PDF file.

[0159] Specific behavior:

[0160] The user clicks on the download link displayed in the device's browser.

[0161] The device downloads the PDF file from the server.

[0162] The user can view the PDF file stored on the device and print it if necessary.

[0163] Through the above specific processing steps, the system generates an original picture book based on the theme selected by the user, and makes it easy to download and print.

[0164] (Application example 1)

[0165] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0166] Currently, many families and educational institutions need to regularly purchase picture books for children. However, existing picture books have fixed content, which can make it difficult to keep children interested. Furthermore, because picture books are physical, collecting a large number of them requires storage space and is costly. Furthermore, parents and educators want to provide educational content that responds to children's development and interests. For these reasons, there is a need for a cost-effective system that can flexibly generate content.

[0167] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0168] In this invention, the server includes means for a user to select a theme, means for generating an original story based on the selected theme, means for generating visual materials based on the generated story, means for converting the generated story and visual materials into an electronic document format, means for the user to download or print the picture book in the electronic document format, and means for the user to select a theme and preview and check the generated picture book in real time using a touch panel terminal, thereby enabling users to generate and print personalized picture books tailored to their wishes on the spot in the store.

[0169] "User" refers to the person who uses this system to select a picture book theme, generate content, download it, and print it.

[0170] "Theme" refers to the subject that forms the basis of the content or subject matter of the picture book selected by the user.

[0171] "Story" refers to an original story generated by AI based on a selected theme.

[0172] "Visual materials" refer to illustrations and images created by AI based on the generated story.

[0173] "Electronic document format" refers to a digital document, such as a PDF format, that integrates and converts the generated narrative and visual materials.

[0174] "Downloading" refers to the process of saving the generated picture book in electronic document format to the user's terminal via the Internet.

[0175] "Printing" refers to the process of outputting the generated picture book in electronic document format onto physical paper media.

[0176] A "touch panel terminal" refers to a computer device equipped with a display that the user can operate by directly touching it.

[0177] "Real-time" refers to the timing at which data is generated and processed immediately.

[0178] "Preview" refers to a temporary display or sample that is displayed to confirm the final generated result.

[0179] "System" refers to the entire platform that provides a set of functions that allow users to generate, download, and print picture books.

[0180] This invention is a system in which a user selects a theme, an AI generates an original story and visual materials based on that theme, and provides a picture book in electronic document format. Specific embodiments for implementing this system will be described below.

[0181] System Configuration

[0182] The present invention is configured by a series of hardware and software components including a user terminal, a server, and a touch panel terminal.

[0183] 1. User Device:

[0184] A device that allows users to select a theme and download or print the generated picture book. Examples include computers, smartphones, and tablets.

[0185] 2. Server:

[0186] It is a central device that receives user requests and generates stories and visual materials using AI models. The server is installed with a language model (e.g., GPT-4) for story generation and an image model (e.g., DALL-E) for visual material generation.

[0187] 3. Touchscreen terminal:

[0188] This device is installed in educational institutions and brick-and-mortar stores, allowing users to select a theme and preview the generated picture book in real time.

[0189] Processing Overview

[0190] When a user accesses the system using a terminal, the touch panel terminal displays a theme selection screen. The user selects the desired theme and presses the generate button on the terminal, and the selected theme data is sent to the server. The server uses a generative AI model to generate an original story based on the selected theme. It then generates illustrations and images based on the generated story. The generated story and visual materials are converted into a digital document in PDF format, which the user can download or print.

[0191] Hardware and Software Use

[0192] Hardware: Touchscreen devices (e.g., Surface Pro), user devices (e.g., iPhone, Android tablet), servers (e.g., AWS EC2 instances)

[0193] Software: The server-side program is implemented in Python and uses web frameworks such as Flask, and also uses GPT-4 for story generation and DALL-E for visual material generation.

[0194] Data processing and calculation

[0195] The server inputs the topic information sent by the user into a natural language processing model and generates a story. Specifically, it uses the following prompt sentences:

[0196] Example prompt: "Generate a story based on the following theme: Dragons."

[0197] Based on the generated story, a request for illustration generation is sent to the visual material generation model. The generated visual material and story are integrated and converted into a PDF file. The PDF is temporarily stored on the server, and a download link is provided to the user.

[0198] Specific examples

[0199] For example, if a user selects the "Dragon" theme using a touchscreen device in a physical store, the server proceeds as follows:

[0200] 1. The user selects the "Dragons" theme.

[0201] 2. The server sends the prompt "Generate a story based on the following theme: Dragons" to the generative AI model.

[0202] 3. The generative AI model generates the story and sends it back to the server.

[0203] 4. The server inputs the story into a visual material generation model and generates illustrations related to dragons.

[0204] 5. Integrate the generated narrative and visual materials and convert them into a PDF file.

[0205] 6. The user downloads the generated PDF or prints it on an in-store printer.

[0206] This system allows users to flexibly create and use personalized picture books on the spot.

[0207] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0208] Step 1:

[0209] The user uses the device to access the theme selection screen on the touch panel device. The server returns HTML content including the theme selection screen, which the device displays. This stage mainly involves collecting user input. Input: None, Output: Display of theme selection screen.

[0210] Step 2:

[0211] The user selects a theme and clicks the generate button on the touch panel terminal. The terminal sends the selected theme data to the server as a POST request. At this stage, the theme selected by the user is sent to the server. Input: Theme selected by the user, Output: Theme data sent to the server.

[0212] Step 3:

[0213] The server processes the received POST request and sends a prompt to the generative AI model based on the selected theme. For example, it sends a prompt such as "Generate a story based on the following theme: Dragons." Input: Theme data, Output: Prompt to the generative AI model.

[0214] Step 4:

[0215] The generative AI model generates an original story based on the server's request and sends it back to the server. The server receives the generated story. Input: prompt sentence, output: generated story.

[0216] Step 5:

[0217] Based on the generated story, the server sends a request to the visual material generation model to generate visual materials (e.g., illustrations and images) appropriate for the story. Input: Generated story, Output: Request to the visual material generation model.

[0218] Step 6:

[0219] The visual material generation model generates visual material based on the server's request and sends it back to the server. The server receives the generated visual material. Input: request, Output: generated visual material.

[0220] Step 7:

[0221] The server combines the generated story and visual materials and converts them into an electronic document in PDF format. Input: Generated story and visual materials, Output: Electronic document in PDF format.

[0222] Step 8:

[0223] The server saves the generated PDF as a temporary file and returns a response containing a download link to the device. Input: PDF file, Output: Response containing a download link.

[0224] Step 9:

[0225] The user clicks the download link on their device to download or print the PDF. At this stage, the final product is obtained by the user. Input: Download link, Output: Save or print the PDF file.

[0226] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0227] This invention combines a system in which a user selects a theme for a picture book, AI generates an original story and images based on that theme, and the picture book can be downloaded or printed in PDF format, with an emotion engine that recognizes the user's emotions. The program and processing of this system are explained in detail below.

[0228] System Overview

[0229] When a user accesses the "AI Storyteller" website using a device, a theme selection screen is displayed. The user selects the desired theme and presses the generate button. The AI ​​then generates an original story and images based on the selected theme. The generated story and images are then converted into PDF format, which the user can download or print. The present invention also incorporates an emotion engine that recognizes the user's emotions, allowing it to recommend themes and adjust the story based on the user's emotions.

[0230] Program processing flow

[0231] 1. When a user accesses a website using a device, the device sends a connection request to the server, and the server returns HTML content including a theme selection screen. The device displays the returned HTML content on the screen, allowing the user to select a theme.

[0232] 2. The emotion engine installed in the device analyzes the user's facial expressions and voice to recognize their emotions. For example, it uses the device's camera and microphone to identify whether the user is happy, depressed, surprised, etc.

[0233] 3. The emotion engine sends the recognized emotion information to the server, which receives the emotion information and uses it to recommend themes and generate stories.

[0234] 4. When the user selects a theme and clicks the generate button, the device sends a POST request containing the selected theme data and emotion information to the server. The server receives the POST request and uses an AI model based on the selected theme and emotion information to generate an original story.

[0235] 5. The server generates images appropriate for the generated story. The server again uses the AI ​​model to execute an illustration generation API based on the content of the story. Emotional information is also taken into account here, and an illustration appropriate for the emotion is generated.

[0236] 6. The server creates a PDF book based on the generated story and images. It uses a PDF generation library to insert the story text and images into a PDF document.

[0237] 7. The created PDF is saved as a temporary file by the server and its path is saved. This temporary file is later provided to the user.

[0238] 8. The server sends a response back to the device containing the path to the generated PDF and a link to download the PDF.

[0239] 9. The device parses the response received from the server and displays a link to download the PDF to the user. When the user clicks the download link, the PDF is saved to the device.

[0240] Specific examples

[0241] For example, if a user selects the theme "Dragon" and the emotion engine recognizes that the user is feeling a little depressed, the server will generate a story based on this theme and emotion information, such as "A brave boy befriends a dragon and goes on a fun adventure together." The AI ​​then creates cheerful, hopeful illustrations of the dragon and boy that fit the generated story. The server then combines these illustrations with the story to create a picture book in PDF format and provides the user with a download link. The user can click the link to save the PDF to their device and read it to their child before bedtime.

[0242] This system allows users to easily obtain high-quality original picture books and provide their children with new stories every night. Furthermore, the introduction of an emotion engine allows the system to provide optimal themes and stories according to the user's emotional state, creating a more personalized educational experience.

[0243] The processing flow will be explained below.

[0244] Step 1:

[0245] A user accesses the "AI Storyteller" website using a device. The device sends a connection request to the web server, and the server returns HTML content including a theme selection screen.

[0246] Step 2:

[0247] The device parses the HTML received from the server and displays the theme selection screen, allowing the user to see the theme selection form.

[0248] Step 3:

[0249] The emotion engine installed in the device analyzes the user's facial expressions and voice to recognize their emotions. For example, it uses the device's camera and microphone to identify whether the user is happy, depressed, surprised, etc.

[0250] Step 4:

[0251] The emotion engine sends the emotional information it recognizes to the server, which then recommends themes appropriate for the user and adjusts story generation based on the emotional information received.

[0252] Step 5:

[0253] The user selects the desired theme from a drop-down menu, for example, one of the options "adventure," "dragon," or "spaceship."

[0254] Step 6:

[0255] When the user presses the "Generate" button, the form data is sent from the device to the server as a POST request, which includes the selected theme and emotion information.

[0256] Step 7:

[0257] The server receives the POST request and retrieves the selected theme and emotion information from the request, which is then used for the next process.

[0258] Step 8:

[0259] The server uses an AI model to generate an original story based on the selected theme and emotional information. The server calls an AI library (e.g., OpenAI GPT-3) and executes the story generation API.

[0260] Step 9:

[0261] The server generates images appropriate for the generated story. The server again uses an AI model (such as DALL-E) to execute an illustration generation API based on the story content and emotional information.

[0262] Step 10:

[0263] The server creates a PDF book based on the generated story and images. It uses a PDF generation library (such as FPDF) to insert the story text and images into the PDF document.

[0264] Step 11:

[0265] The server saves the completed PDF as a temporary file and retains its path, which is later provided to the user.

[0266] Step 12:

[0267] The server sends a response back to the device containing the path to the generated PDF, along with a link to download the PDF.

[0268] Step 13:

[0269] The device parses the response received from the server and displays a link to download the PDF to the user. When the user clicks the link, the PDF is saved on the device.

[0270] Step 14:

[0271] By opening the saved PDF, users can view, download, and print the original picture book they created. The emotional engine allows the system to provide the most appropriate theme and story based on the user's emotional state, enabling a personalized educational experience.

[0272] Through this series of steps, users can easily obtain high-quality original picture books and provide their children with new stories every night. The emotional engine generates themes and stories that best suit the user's emotional state, providing a more personalized educational experience.

[0273] Example 2

[0274] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0275] In conventional picture book generation systems, users simply select a theme, and the generated story and images are not optimized for the user's emotional state. As a result, picture books with content that does not match the user's emotions or mood are generated, resulting in low satisfaction. Furthermore, because emotion recognition technology is not incorporated, personalized educational experiences are not provided.

[0276] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0277] In this invention, the server includes: means for a user to select a theme; means for generating an original story based on the selected theme; means for generating images based on the generated story; means for converting the generated story and images into PDF format; means for the user to download or print the PDF picture book; and means including an emotion engine that recognizes the user's emotions and that recommends themes and adjusts the story based on the recognized emotions. This automatically generates an original picture book suited to the user's emotional state, enabling the user to have a highly satisfying and personalized educational experience.

[0278] A "user" is an individual who uses the system to create a picture book.

[0279] A "theme" is a topic or category that determines the content of the picture book selected by the user.

[0280] An "original story" is a unique narrative created by a generative AI model based on a selected theme.

[0281] "Images" are visual content that is automatically created based on the generated story.

[0282] "PDF format" is a format that integrates the generated story and images into a single document file and saves it in Portable Document Format (PDF).

[0283] An "emotion engine" is a software and hardware system that analyzes a user's facial expressions, voice, etc., and recognizes the user's emotional state.

[0284] A "server" is a computer system that receives user requests and generates stories and images and creates PDF files in response.

[0285] "Means of selection" refers to the interface (e.g., a screen or menu on a website) through which a user selects a theme.

[0286] "Generative means" refers to the AI ​​models or algorithms used to automatically create original stories and images based on selected themes and emotional information.

[0287] "Means for recommending themes and adjusting stories based on emotions" refers to a function that suggests appropriate themes and adjusts the content of stories based on the user's recognized emotional information.

[0288] This system allows users to select a theme, generates an original story and images based on that theme using a generative AI model, and provides a picture book in PDF format. Furthermore, it incorporates an emotion engine that recognizes the user's emotions, and recommends themes and adjusts the story based on this emotion information.

[0289] System configuration

[0290] This system consists of the following elements:

[0291] 1. User Device:

[0292] Browser: Software used to access websites.

[0293] Camera: Hardware that captures the user's facial expressions.

[0294] Microphone: Hardware that captures the user's voice.

[0295] Emotion engine: Software that recognizes user emotions by analyzing facial expressions and voice (e.g., OpenCV, Deep Learning models).

[0296] 2. Server:

[0297] Web server: A computer system that receives requests from users and returns appropriate responses.

[0298] Generative AI models: Language models that generate stories based on selected themes and sentiment information (e.g., GPT-3).

[0299] Illustration generation API: An API for generating images based on a generated story (e.g., DALL-E).

[0300] PDF Generation Library: A library that converts the generated stories and images into PDF format (e.g. FPDF, PDFBox).

[0301] System Operation

[0302] When a user accesses the "AI Storyteller" website using their device, the device sends a connection request to the server, which then returns HTML content including a theme selection screen. The device's built-in emotion engine then activates and analyzes the user's facial expressions and voice to recognize their emotions. For example, the device's camera and microphone can be used to determine whether the user is happy or depressed. The emotion information is then sent to the server, which uses it to recommend themes and generate stories.

[0303] When a user selects a theme and clicks the generate button, the device sends a POST request containing the selected theme data and emotional information to the server. The server receives the POST request and uses a generative AI model based on the selected theme and emotional information to generate an original story. Images appropriate for the generated story are also generated on the server, and further appropriate illustrations are selected based on the emotional information.

[0304] Once the story and images are ready, the server uses a PDF generation library to convert them into a PDF picture book. The server saves the created PDF as a temporary file and sends the path to the file back to the user. The user can click the link to download the PDF and save it on their device.

[0305] Specific examples

[0306] For example, if a user selects the theme "Dragon" and the emotion engine recognizes that the user is feeling a little depressed, the server will generate a story based on this theme and emotion information, such as "A brave boy befriends a dragon and goes on a fun adventure together." The AI ​​then creates bright and hopeful illustrations of the dragon and boy that fit the generated story. The server then combines these illustrations with the story to create a picture book in PDF format and provides the user with a download link. The user can click the link to save the PDF to their device and read it to their child before bedtime.

[0307] Example prompt sentence:

[0308] "Adventures with Dragons"

[0309] "Stories that make children brave"

[0310] "Make friends"

[0311] "Hopeful illustrations"

[0312] This system allows users to easily obtain high-quality original picture books and provide their children with new stories every night. Furthermore, the introduction of an emotion engine allows the system to provide the most appropriate themes and stories according to the user's emotional state, creating a more personalized educational experience.

[0313] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0314] Step 1:

[0315] When a user accesses a website using a device, the device sends an HTTP request to the server. The server receives the request, generates HTML content including a theme selection screen, and sends it back to the device as an HTTP response.

[0316] Input: URL access request from user

[0317] Data processing / calculation: The server analyzes the request and generates appropriate HTML content

[0318] Output: Response showing theme selection screen

[0319] Step 2:

[0320] The emotion engine installed on the device is activated and uses the camera and microphone to capture the user's facial expressions and voice in real time. The emotion engine then uses facial expression and voice analysis software to recognize the user's emotions and generate data.

[0321] Input: User's facial expressions and voice

[0322] Data processing / computation: Capture data using camera and microphone and perform sentiment analysis

[0323] Output: User emotion data

[0324] Step 3:

[0325] The device sends the recognized emotion information to the server, which receives the emotion data and stores it for later processing.

[0326] Input: Emotion data

[0327] Data processing / calculation: Data reception and storage

[0328] Output: Saved emotion data

[0329] Step 4:

[0330] The user selects the desired theme on the theme selection screen and clicks the Generate button. The device then sends a POST request to the server containing the selected theme and the previously recognized emotion data.

[0331] Input: User-selected theme, stored emotion data

[0332] Data processing / calculation: Combining theme selection data and emotion data, generating POST requests

[0333] Output: POST request sent

[0334] Step 5:

[0335] The server receives the POST request and generates an original story using a generative AI model based on the selected theme and emotion data. For example, a prompt sentence is input to a generative AI model (e.g., GPT-3) to generate the story as text.

[0336] Input: Theme data, emotion data

[0337] Data processing / calculation: Use of generative AI models, generation of story text

[0338] Output: Generated story text

[0339] Step 6:

[0340] The server uses an illustration generation API (e.g., DALL-E) based on the generated story to generate images that suit the story, taking into account emotional data to create visuals that are optimal for the user.

[0341] Input: Story text, emotion data

[0342] Data processing / calculation: Calling illustration generation API, image generation

[0343] Output: The generated image

[0344] Step 7:

[0345] The server converts the generated story and images into a PDF book using a PDF generation library (e.g., FPDF, PDFBox). The story and images are inserted into a template and a PDF file is generated.

[0346] Input: Story text, generated images

[0347] Data processing / calculation: Use of PDF generation library, creation of PDF files

[0348] Output: PDF picture book

[0349] Step 8:

[0350] The server saves the created PDF as a temporary file, retains its path, and returns a response including the path of the saved file to the terminal.

[0351] Input: Generated PDF

[0352] Data processing / calculation: file saving, path generation

[0353] Output: Response to device, download link

[0354] Step 9:

[0355] The device parses the response received from the server and displays a link to download the PDF to the user. When the user clicks the link, the PDF is saved on the device.

[0356] Input: Response from the server

[0357] Data processing / calculation: Response analysis, link display

[0358] Output: Download and save PDF

[0359] (Application example 2)

[0360] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0361] Conventional illustrated work instruction generation systems create formulaic instructions without considering the user's emotional state, which does not adequately improve worker motivation or provide emotional care. Furthermore, instructions that do not take emotions into consideration may affect the efficiency and effectiveness of work. Therefore, there is a need to provide customized work instructions and work plans that correspond to the emotional state of workers.

[0362] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for a user to select a theme, means for generating an original story based on the selected theme, means for generating images based on the generated story, means for converting the generated story and images into PDF format, means for the user to download or print the PDF format document, means for recognizing a user's emotion, means for adjusting the story and images based on the recognized emotion, means for recognizing a user's emotion in an industrial environment and adjusting work instructions and work plans, and means for converting the generated work instructions and work plans into PDF format. This makes it possible to provide work instructions and work plans customized according to the emotional state of the worker.

[0363] A "theme" is the content or subject matter that a user selects to create an original story.

[0364] A "story" is a narrative or story content that is generated based on a selected theme.

[0365] "Images" refers to visual content created based on the generated story, including visual information such as pictures and illustrations.

[0366] "PDF format" is an abbreviation for Portable Document Format, and refers to a fixed-layout file format suitable for storing and distributing electronic documents.

[0367] A "document" is a written document that contains textual information such as a story or instructions.

[0368] "Emotions" represent the user's psychological state or sensations, and include types such as joy, sadness, surprise, and anger.

[0369] An "emotion engine" is a software or hardware mechanism that uses sensors such as cameras and microphones to recognize and analyze a user's emotional state.

[0370] "Industrial environment" refers to the working environment at a production or manufacturing site or facility.

[0371] A "work instruction" is a document that describes the work content and procedures that workers should perform in an industrial environment.

[0372] A "work plan" is a document that includes detailed processes and schedules for efficiently carrying out a specific task.

[0373] A "server" is a computer system that provides services over a network, and refers to the equipment and software that processes and distributes data in response to requests.

[0374] A "generative AI model" is an algorithm or model that uses artificial intelligence technology to generate new content (stories or images) based on input data.

[0375] The present invention combines a system in which a user selects a theme, AI generates original stories and images based on that theme, and provides them in PDF format with an emotion engine that recognizes the user's emotions. This system can provide customized work instructions and work plans according to the emotional state of workers, especially in industrial environments.

[0376] Hardware and software used

[0377] Camera: A device that captures images for emotion recognition.

[0378] Microphone: A device that captures the user's voice emotions if necessary.

[0379] Speech recognition engine: A software component that analyzes the user's emotions.

[0380] EmotionRecognizer: A library that analyzes facial expression data and recognizes user emotions.

[0381] Generative AI Model: An AI algorithm for generating stories and images based on a theme (StoryGenerator, ImageGenerator).

[0382] PDFCreator: A library that converts generated stories and images into PDF format.

[0383] Natural language processing explanation

[0384] 1. User selects a theme:

[0385] A user uses a terminal to select a theme in an industrial environment (e.g., inspection work on line B). This theme indicates the work content and details of the workers.

[0386] 2. Emotion recognition:

[0387] The device's built-in camera and microphone are used to acquire the user's facial expression and voice data. The EmotionRecognizer library is used to recognize the user's emotions (e.g., joy, sadness, anger) from this data. This emotional information is sent to the server and used to generate stories and images.

[0388] 3. Story and image generation:

[0389] Based on the recognized emotional information and the selected theme, the server uses a Generative AI Model (Story Generator, Image Generator) to generate a customized story and associated images.

[0390] 4. Generate and serve PDF:

[0391] The generated stories and images are used to create PDF work instructions and work plans using the PDFCreator library, which are then saved on the server and a download link is provided to the user.

[0392] Specific examples

[0393] For example, if a user selects "Inspection work on Line B" as the theme and the emotion engine recognizes that the user is feeling a little down, the AI ​​can use this information to suggest "Perform equipment inspection on Line B. The theme will be adjusted to include simple and easy-to-understand steps and an encouraging message." This allows the user to receive customized instructions that reflect their emotional state that day, improving work efficiency.

[0394] Prompt Sentence Examples

[0395] "Today, the user is tasked with inspecting the machines on Line B. If the user is feeling down, generate a story that gives them simple instructions with an encouraging message."

[0396] "Emotion: Sad\nTask Details: Machine Inspection\nGenerate a positive and encouraging instruction set."

[0397] In this way, the system of the present invention can provide customized support according to the user's emotional state, contributing to improving work efficiency and maintaining worker motivation, particularly in industrial environments.

[0398] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0399] Step 1:

[0400] The user selects a theme on the terminal. From the theme list displayed on the terminal screen, the user selects a theme for a work instruction in an industrial environment (e.g., inspection work on line B). The input here is the theme selected by the user, and the output is detailed information about the selected theme. This information is passed to the subsequent processing step.

[0401] Step 2:

[0402] The system recognizes the user's emotions using the device's camera and microphone. The system uses the device's hardware (camera and microphone) to acquire the user's facial expression data and voice data. The system uses the EmotionRecognizer library to analyze the user's emotions from this data. The input is the acquired facial expression data and voice data, and the output is analyzed emotional information (e.g., joy, sadness, anger).

[0403] Step 3:

[0404] The recognized emotion information is sent to the server. The device then sends the analyzed emotion information to the server. At this time, the theme information selected by the user is also sent to the server. The emotion information and theme information are input, and the emotion information and theme information received by the server are output.

[0405] Step 4:

[0406] The server generates the story and images. Based on the theme and emotion information received on the server side, a generative AI model (Story Generator, Image Generator) is used to generate a customized story and images. The theme and emotion information are input here, and the generated story and images are output. Specifically, the AI ​​model creates a story while taking emotion information into account, and generates a prompt that generates an illustration based on the story.

[0407] Step 5:

[0408] The generated story and images are converted into PDF format. The server uses the generated story and images to convert them into a PDF document using the PDFCreator library. The input data is the generated story and images, and the output is a PDF document. Specifically, PDFCreator combines the story text and images to create a single PDF file.

[0409] Step 6:

[0410] The generated PDF is provided to the user. The server saves the generated PDF file as a temporary file and provides the link to the user. The input is the PDF file, and the output is a link to the PDF file that the user can access. The user can download or print the PDF from this link on their device.

[0411] Step 7:

[0412] The user downloads or prints a PDF. The device clicks on a link to download the PDF file and optionally prints it. The input to this step is the PDF link provided by the server, and the output is the downloaded PDF file or printed document. Specifically, the device opens the link, saves the file, and sends it to the printer to print the physical document.

[0413] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0414] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0415] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0416] [Second embodiment]

[0417] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0418] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0419] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0420] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0421] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0422] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0423] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0424] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0425] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0426] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0427] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0428] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0429] This invention provides a system in which a user selects a theme for a picture book, AI generates an original story and images based on that theme, and the picture book can be downloaded or printed in PDF format. The program and processing of this system are explained in detail below.

[0430] System Overview

[0431] When a user accesses the "AI Storyteller" website using a device, a theme selection screen appears. The user selects the desired theme and presses the generate button. The AI ​​then generates an original story and images based on the selected theme. The generated story and images are then converted into PDF format, which the user can download or print.

[0432] Program processing flow

[0433] 1. When a user accesses a website using a device, the device sends a connection request to the server, and the server returns HTML content including a theme selection screen. The device displays the returned HTML content on the screen, allowing the user to select a theme.

[0434] 2. When the user selects a theme and clicks the generate button, the device sends the selected theme data as a POST request to the server. The server receives the POST request and generates an original story based on the selected theme using an AI model.

[0435] 3. The server uses an AI model to generate illustrations based on the generated story, and combines the generated story and illustrations to create a picture book in PDF format.

[0436] 4. The created PDF is saved as a temporary file by the server, and a response containing a link that the user can download is sent back to the device.

[0437] 5. When the user clicks the download link on their device, the PDF is saved to their device. The user can view, download, and print the saved PDF.

[0438] Specific examples

[0439] For example, if a user selects the theme "Dragon," the server generates a story based on this theme, such as "A brave boy makes friends with a dragon." The AI ​​then creates illustrations of the dragon and boy that fit the generated story. The server then combines these illustrations with the story to create a picture book in PDF format and provides the user with a download link. The user can click the link to save the PDF on their device and read it to their child before bedtime.

[0440] This system eliminates the need for users to frequently purchase expensive picture books, allowing them to provide their children with a new story every night. Furthermore, by setting themes with educational messages, it is possible to effectively educate children. In this way, the present invention alleviates parental concerns and provides an efficient means for supporting children's curiosity and learning.

[0441] The processing flow will be explained below.

[0442] Step 1:

[0443] A user accesses the "AI Storyteller" website using a device. The device sends a connection request to the web server, and the web server returns HTML content including a theme selection screen.

[0444] Step 2:

[0445] The device parses the HTML received from the server and displays the theme selection screen, allowing the user to see the theme selection form.

[0446] Step 3:

[0447] The user selects the desired theme from a drop-down menu, for example, one of the options "adventure," "dragon," or "spaceship."

[0448] Step 4:

[0449] When the user presses the "Generate" button, the form data is sent from the terminal to the server as a POST request, including the selected theme.

[0450] Step 5:

[0451] The server receives the POST request and retrieves the selected theme data from the request, which is then used for the next process.

[0452] Step 6:

[0453] The server uses an AI model to generate an original story based on the selected theme. The server calls an AI library (e.g., OpenAI GPT-3) and executes the story generation API.

[0454] Step 7:

[0455] The server generates images appropriate for the generated story. The server again uses an AI model (e.g., DALL-E) to execute an illustration generation API based on the content of the story.

[0456] Step 8:

[0457] The server creates a PDF book based on the generated story and images. It uses a PDF generation library (e.g., FPDF) to insert the story text and images into the PDF document.

[0458] Step 9:

[0459] The server saves the completed PDF as a temporary file and retains its path, which is later provided to the user.

[0460] Step 10:

[0461] The server sends a response back to the device containing the path to the generated PDF, along with a link to download the PDF.

[0462] Step 11:

[0463] The device parses the response received from the server and displays a PDF download link to the user, which the user can click to download the PDF.

[0464] Step 12:

[0465] When a user clicks the download link, the device saves the PDF file to local storage. By opening the saved PDF, the user can view, download, and print the generated original picture book.

[0466] Through this series of processes, users can easily obtain high-quality original picture books and provide their children with new stories every night.

[0467] Example 1

[0468] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0469] In modern child-rearing, parents want to provide their children with a variety of fresh picture books that meet their children's interests and educational needs. However, frequently purchasing new picture books is a significant financial burden, and commercially available picture books do not always contain the educational messages parents desire. Furthermore, busy parents find it difficult to quickly obtain personalized picture books tailored to their children's preferences. Efficient methods to solve these problems are needed.

[0470] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0471] In this invention, the server includes means for a user to select a theme, means for generating an original story using a generative AI model based on the selected theme, means for generating images using the generative AI model based on the content of the generated story, means for converting the generated story and images into a PDF format, and means for the user to download or print the PDF picture book. This allows a user to easily generate an original picture book based on a theme and download or print it.

[0472] A "theme" is a subject or topic selected by the user that will be the basis for the story or image that will be generated.

[0473] A "generative AI model" is an artificial intelligence algorithm or system that generates text or images based on a selected theme.

[0474] An "original story" is a unique narrative generated by a generative AI model based on a selected theme.

[0475] "Images" are visual illustrations or pictures generated by a generative AI model to match the content of the original story.

[0476] "PDF format" is an abbreviation for Portable Document Format, and is a file format for saving and displaying documents and images on a page-by-page basis.

[0477] A picture book is a book for children that contains stories and illustrations in written form.

[0478] "Downloading" is the act of saving a file from the Internet to your device.

[0479] "Printing" is the act of printing digital data onto physical media such as paper.

[0480] A "user" is someone who uses the system to select a theme, create, download, and print a picture book.

[0481] A "server" is a computer system that receives requests from users and generates, stores, and delivers appropriate content.

[0482] A "terminal" is a device operated by a user, and is a computer that communicates with a server and displays and saves results.

[0483] The present invention is a system that uses a generative AI model to generate original stories and images based on a theme selected by a user, and provides them as a PDF picture book. The system includes a means for a user to select a theme, a means for converting the generated stories and images into PDF format, and a means for the user to download or print the PDF picture book.

[0484] Hardware and Software

[0485] The system is implemented using a server, a terminal, a generative AI model, and various software. The server receives requests, generates stories and images, converts them into PDF format, and distributes them. The terminal acts as an interface for users to access the system, select a theme, and download the PDF.

[0486] The main hardware and software used are as follows:

[0487] Server: Processes the data and invokes the generative AI model.

[0488] Terminal: The device through which the user accesses the system, such as a regular PC, smartphone, or tablet.

[0489] Generative AI models: For example, ChatGPT and GPT-4 are used for story generation, while DALL-E and Stable Diffusion are used for image generation.

[0490] PDF generation library: Generate PDFs using Python's ReportLab, etc.

[0491] Data processing and calculation

[0492] 1. Theme selection: The user selects a theme using the terminal. For example, the user can select the theme "Dragon."

[0493] 2. Story generation: The server uses a generative AI model (such as ChatGPT or GPT-4) to generate an original story based on the selected theme. An example of a specific prompt is "Theme: Dragons. Please generate a story about a brave boy who becomes friends with a dragon."

[0494] 3. Image Generation: The server uses a generative AI model (such as DALL-E or Stable Diffusion) to generate images that match the generated story. The prompt might be something like, "Generate an illustration of a dragon and a boy."

[0495] 4. PDF generation: The server combines the generated story and images and generates a PDF picture book using Python's ReportLab or similar.

[0496] 5. PDF Delivery: The server saves the PDF in temporary storage and provides a download link that users can access, and they can download or print the PDF using their devices.

[0497] Specific examples

[0498] For example, if a user selects "Dragons" as the theme, the process goes as follows: The user selects a theme through the website and presses the generate button. The device then sends the theme data to the server. The server uses a generative AI model to generate a story such as "A brave boy befriends a dragon" and creates illustrations of the dragon and boy based on that story. The server then assembles this content into a PDF picture book using Python's ReportLab and provides the user with a download link. The user can click this link to download the PDF and read it to their child.

[0499] This system eliminates the need for users to frequently purchase expensive picture books, allowing them to provide their children with a new story every night. Furthermore, by setting themes with educational messages, it is possible to effectively educate children. In this way, the present invention alleviates parental concerns and provides an efficient means for supporting children's curiosity and learning.

[0500] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0501] Specific flow of system program processing

[0502] Step 1:

[0503] The user accesses the "AI Storyteller" website using a device.

[0504] Input: An HTTP request sent by a user typing a URL into a web browser.

[0505] Output: The device sends a connection request to the server, and the server returns a theme selection screen.

[0506] Specific behavior:

[0507] The user opens a browser, enters the specified URL, and presses enter.

[0508] The device sends this operation to the server as an HTTP request.

[0509] The server receives the request and generates HTML content including a theme selection screen.

[0510] The server returns the generated HTML content to the terminal as an HTTP response.

[0511] The terminal displays the received HTML content on the screen and allows the user to select a theme.

[0512] Step 2:

[0513] The user selects a theme and clicks the generate button.

[0514] Input: The user selects a theme and clicks the generate button.

[0515] Output: The device sends the selected theme data to the server as a POST request.

[0516] Specific behavior:

[0517] The user selects a desired theme from the displayed theme selection screen.

[0518] The user clicks the "Generate" button.

[0519] The device sends the user's selected theme data to the server as a POST request in JSON format.

[0520] Step 3:

[0521] The server receives the POST request and generates the story.

[0522] Input: A POST request containing theme data sent from the device.

[0523] Output: The original story generated using the generative AI model.

[0524] Specific behavior:

[0525] The server receives the POST request and extracts the theme data from the request body.

[0526] The server sends a prompt to the generative AI model (e.g., ChatGPT or GPT-4). Example prompt: "Theme: Dragons. Generate a story about a brave boy who becomes friends with a dragon."

[0527] The generative AI model generates a story based on the prompt sentence and returns the story to the server.

[0528] Step 4:

[0529] The server generates illustrations based on the story.

[0530] Input: Generated stories.

[0531] Output: An illustration generated using a generative AI model.

[0532] Specific behavior:

[0533] The server analyzes the content of the generated story and generates a prompt accordingly. Example prompt: "Please generate an illustration of a dragon and a boy."

[0534] The server sends a prompt to a generative AI model (e.g., DALL-E or Stable Diffusion).

[0535] The generative AI model generates an illustration based on the prompt sentence and returns the illustration to the server.

[0536] Step 5:

[0537] The server creates a picture book in PDF format.

[0538] Input: Generated story and illustrations.

[0539] Output: Picture book in PDF format.

[0540] Specific behavior:

[0541] The server combines the generated story and illustrations and creates a PDF picture book using Python's ReportLab etc.

[0542] The server saves the generated PDF file in temporary storage.

[0543] Step 6:

[0544] The server provides a download link for the PDF.

[0545] Input: PDF file stored on the server.

[0546] Output: A response containing a download link for the PDF.

[0547] Specific behavior:

[0548] The server generates a download link for the PDF file stored in temporary storage.

[0549] The server sends a response including the download link back to the device.

[0550] Step 7:

[0551] The user downloads or prints the PDF.

[0552] Input: Click on the download link displayed on your device.

[0553] Output: Downloaded PDF file.

[0554] Specific behavior:

[0555] The user clicks on the download link displayed in the device's browser.

[0556] The device downloads the PDF file from the server.

[0557] The user can view the PDF file stored on the device and print it if necessary.

[0558] Through the above specific processing steps, the system generates an original picture book based on the theme selected by the user, and makes it easy to download and print.

[0559] (Application example 1)

[0560] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0561] Currently, many families and educational institutions need to regularly purchase picture books for children. However, existing picture books have fixed content, which can make it difficult to keep children interested. Furthermore, because picture books are physical, collecting a large number of them requires storage space and is costly. Furthermore, parents and educators want to provide educational content that responds to children's development and interests. For these reasons, there is a need for a cost-effective system that can flexibly generate content.

[0562] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0563] In this invention, the server includes means for a user to select a theme, means for generating an original story based on the selected theme, means for generating visual materials based on the generated story, means for converting the generated story and visual materials into an electronic document format, means for the user to download or print the picture book in the electronic document format, and means for the user to select a theme and preview and check the generated picture book in real time using a touch panel terminal, thereby enabling users to generate and print personalized picture books tailored to their wishes on the spot in the store.

[0564] "User" refers to the person who uses this system to select a picture book theme, generate content, download it, and print it.

[0565] "Theme" refers to the subject that forms the basis of the content or subject matter of the picture book selected by the user.

[0566] "Story" refers to an original story generated by AI based on a selected theme.

[0567] "Visual materials" refer to illustrations and images created by AI based on the generated story.

[0568] "Electronic document format" refers to a digital document, such as a PDF format, that integrates and converts the generated narrative and visual materials.

[0569] "Downloading" refers to the process of saving the generated picture book in electronic document format to the user's terminal via the Internet.

[0570] "Printing" refers to the process of outputting the generated picture book in electronic document format onto physical paper media.

[0571] A "touch panel terminal" refers to a computer device equipped with a display that the user can operate by directly touching it.

[0572] "Real-time" refers to the timing at which data is generated and processed immediately.

[0573] "Preview" refers to a temporary display or sample that is displayed to confirm the final generated result.

[0574] "System" refers to the entire platform that provides a set of functions that allow users to generate, download, and print picture books.

[0575] This invention is a system in which a user selects a theme, an AI generates an original story and visual materials based on that theme, and provides a picture book in electronic document format. Specific embodiments for implementing this system will be described below.

[0576] System Configuration

[0577] The present invention is configured by a series of hardware and software components including a user terminal, a server, and a touch panel terminal.

[0578] 1. User Device:

[0579] A device that allows users to select a theme and download or print the generated picture book. Examples include computers, smartphones, and tablets.

[0580] 2. Server:

[0581] It is a central device that receives user requests and generates stories and visual materials using AI models. The server is installed with a language model (e.g., GPT-4) for story generation and an image model (e.g., DALL-E) for visual material generation.

[0582] 3. Touchscreen terminal:

[0583] This device is installed in educational institutions and brick-and-mortar stores, allowing users to select a theme and preview the generated picture book in real time.

[0584] Processing Overview

[0585] When a user accesses the system using a terminal, the touch panel terminal displays a theme selection screen. The user selects the desired theme and presses the generate button on the terminal, and the selected theme data is sent to the server. The server uses a generative AI model to generate an original story based on the selected theme. It then generates illustrations and images based on the generated story. The generated story and visual materials are converted into a digital document in PDF format, which the user can download or print.

[0586] Hardware and Software Use

[0587] Hardware: Touchscreen devices (e.g., Surface Pro), user devices (e.g., iPhone, Android tablet), servers (e.g., AWS EC2 instances)

[0588] Software: The server-side program is implemented in Python and uses web frameworks such as Flask, and also uses GPT-4 for story generation and DALL-E for visual material generation.

[0589] Data processing and calculation

[0590] The server inputs the topic information sent by the user into a natural language processing model and generates a story. Specifically, it uses the following prompt sentences:

[0591] Example prompt: "Generate a story based on the following theme: Dragons."

[0592] Based on the generated story, a request for illustration generation is sent to the visual material generation model. The generated visual material and story are integrated and converted into a PDF file. The PDF is temporarily stored on the server, and a download link is provided to the user.

[0593] Specific examples

[0594] For example, if a user selects the "Dragon" theme using a touchscreen device in a physical store, the server proceeds as follows:

[0595] 1. The user selects the "Dragons" theme.

[0596] 2. The server sends the prompt "Generate a story based on the following theme: Dragons" to the generative AI model.

[0597] 3. The generative AI model generates the story and sends it back to the server.

[0598] 4. The server inputs the story into a visual material generation model and generates illustrations related to dragons.

[0599] 5. Integrate the generated narrative and visual materials and convert them into a PDF file.

[0600] 6. The user downloads the generated PDF or prints it on an in-store printer.

[0601] This system allows users to flexibly create and use personalized picture books on the spot.

[0602] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0603] Step 1:

[0604] The user uses the device to access the theme selection screen on the touch panel device. The server returns HTML content including the theme selection screen, which the device displays. This stage mainly involves collecting user input. Input: None, Output: Display of theme selection screen.

[0605] Step 2:

[0606] The user selects a theme and clicks the generate button on the touch panel terminal. The terminal sends the selected theme data to the server as a POST request. At this stage, the theme selected by the user is sent to the server. Input: Theme selected by the user, Output: Theme data sent to the server.

[0607] Step 3:

[0608] The server processes the received POST request and sends a prompt to the generative AI model based on the selected theme. For example, it sends a prompt such as "Generate a story based on the following theme: Dragons." Input: Theme data, Output: Prompt to the generative AI model.

[0609] Step 4:

[0610] The generative AI model generates an original story based on the server's request and sends it back to the server. The server receives the generated story. Input: prompt sentence, output: generated story.

[0611] Step 5:

[0612] Based on the generated story, the server sends a request to the visual material generation model to generate visual materials (e.g., illustrations and images) appropriate for the story. Input: Generated story, Output: Request to the visual material generation model.

[0613] Step 6:

[0614] The visual material generation model generates visual material based on the server's request and sends it back to the server. The server receives the generated visual material. Input: request, Output: generated visual material.

[0615] Step 7:

[0616] The server combines the generated story and visual materials and converts them into an electronic document in PDF format. Input: Generated story and visual materials, Output: Electronic document in PDF format.

[0617] Step 8:

[0618] The server saves the generated PDF as a temporary file and returns a response containing a download link to the device. Input: PDF file, Output: Response containing a download link.

[0619] Step 9:

[0620] The user clicks the download link on their device to download or print the PDF. At this stage, the final product is obtained by the user. Input: Download link, Output: Save or print the PDF file.

[0621] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0622] This invention combines a system in which a user selects a theme for a picture book, AI generates an original story and images based on that theme, and the picture book can be downloaded or printed in PDF format, with an emotion engine that recognizes the user's emotions. The program and processing of this system are explained in detail below.

[0623] System Overview

[0624] When a user accesses the "AI Storyteller" website using a device, a theme selection screen is displayed. The user selects the desired theme and presses the generate button. The AI ​​then generates an original story and images based on the selected theme. The generated story and images are then converted into PDF format, which the user can download or print. The present invention also incorporates an emotion engine that recognizes the user's emotions, allowing it to recommend themes and adjust the story based on the user's emotions.

[0625] Program processing flow

[0626] 1. When a user accesses a website using a device, the device sends a connection request to the server, and the server returns HTML content including a theme selection screen. The device displays the returned HTML content on the screen, allowing the user to select a theme.

[0627] 2. The emotion engine installed in the device analyzes the user's facial expressions and voice to recognize their emotions. For example, it uses the device's camera and microphone to identify whether the user is happy, depressed, surprised, etc.

[0628] 3. The emotion engine sends the recognized emotion information to the server, which receives the emotion information and uses it to recommend themes and generate stories.

[0629] 4. When the user selects a theme and clicks the generate button, the device sends a POST request containing the selected theme data and emotion information to the server. The server receives the POST request and uses an AI model based on the selected theme and emotion information to generate an original story.

[0630] 5. The server generates images appropriate for the generated story. The server again uses the AI ​​model to execute an illustration generation API based on the content of the story. Emotional information is also taken into account here, and an illustration appropriate for the emotion is generated.

[0631] 6. The server creates a PDF book based on the generated story and images. It uses a PDF generation library to insert the story text and images into a PDF document.

[0632] 7. The created PDF is saved as a temporary file by the server and its path is saved. This temporary file is later provided to the user.

[0633] 8. The server sends a response back to the device containing the path to the generated PDF and a link to download the PDF.

[0634] 9. The device parses the response received from the server and displays a link to download the PDF to the user. When the user clicks the download link, the PDF is saved to the device.

[0635] Specific examples

[0636] For example, if a user selects the theme "Dragon" and the emotion engine recognizes that the user is feeling a little depressed, the server will generate a story based on this theme and emotion information, such as "A brave boy befriends a dragon and goes on a fun adventure together." The AI ​​then creates cheerful, hopeful illustrations of the dragon and boy that fit the generated story. The server then combines these illustrations with the story to create a picture book in PDF format and provides the user with a download link. The user can click the link to save the PDF to their device and read it to their child before bedtime.

[0637] This system allows users to easily obtain high-quality original picture books and provide their children with new stories every night. Furthermore, the introduction of an emotion engine allows the system to provide optimal themes and stories according to the user's emotional state, creating a more personalized educational experience.

[0638] The processing flow will be explained below.

[0639] Step 1:

[0640] A user accesses the "AI Storyteller" website using a device. The device sends a connection request to the web server, and the server returns HTML content including a theme selection screen.

[0641] Step 2:

[0642] The device parses the HTML received from the server and displays the theme selection screen, allowing the user to see the theme selection form.

[0643] Step 3:

[0644] The emotion engine installed in the device analyzes the user's facial expressions and voice to recognize their emotions. For example, it uses the device's camera and microphone to identify whether the user is happy, depressed, surprised, etc.

[0645] Step 4:

[0646] The emotion engine sends the emotional information it recognizes to the server, which then recommends themes appropriate for the user and adjusts story generation based on the emotional information received.

[0647] Step 5:

[0648] The user selects the desired theme from a drop-down menu, for example, one of the options "adventure," "dragon," or "spaceship."

[0649] Step 6:

[0650] When the user presses the "Generate" button, the form data is sent from the device to the server as a POST request, which includes the selected theme and emotion information.

[0651] Step 7:

[0652] The server receives the POST request and retrieves the selected theme and emotion information from the request, which is then used for the next process.

[0653] Step 8:

[0654] The server uses an AI model to generate an original story based on the selected theme and emotional information. The server calls an AI library (e.g., OpenAI GPT-3) and executes the story generation API.

[0655] Step 9:

[0656] The server generates images appropriate for the generated story. The server again uses an AI model (such as DALL-E) to execute an illustration generation API based on the story content and emotional information.

[0657] Step 10:

[0658] The server creates a PDF book based on the generated story and images. It uses a PDF generation library (such as FPDF) to insert the story text and images into the PDF document.

[0659] Step 11:

[0660] The server saves the completed PDF as a temporary file and retains its path, which is later provided to the user.

[0661] Step 12:

[0662] The server sends a response back to the device containing the path to the generated PDF, along with a link to download the PDF.

[0663] Step 13:

[0664] The device parses the response received from the server and displays a link to download the PDF to the user. When the user clicks the link, the PDF is saved on the device.

[0665] Step 14:

[0666] By opening the saved PDF, users can view, download, and print the original picture book they created. The emotional engine allows the system to provide the most appropriate theme and story based on the user's emotional state, enabling a personalized educational experience.

[0667] Through this series of steps, users can easily obtain high-quality original picture books and provide their children with new stories every night. The emotional engine generates themes and stories that best suit the user's emotional state, providing a more personalized educational experience.

[0668] Example 2

[0669] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0670] In conventional picture book generation systems, users simply select a theme, and the generated story and images are not optimized for the user's emotional state. As a result, picture books with content that does not match the user's emotions or mood are generated, resulting in low satisfaction. Furthermore, because emotion recognition technology is not incorporated, personalized educational experiences are not provided.

[0671] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0672] In this invention, the server includes: means for a user to select a theme; means for generating an original story based on the selected theme; means for generating images based on the generated story; means for converting the generated story and images into PDF format; means for the user to download or print the PDF picture book; and means including an emotion engine that recognizes the user's emotions and that recommends themes and adjusts the story based on the recognized emotions. This automatically generates an original picture book suited to the user's emotional state, enabling the user to have a highly satisfying and personalized educational experience.

[0673] A "user" is an individual who uses the system to create a picture book.

[0674] A "theme" is a topic or category that determines the content of the picture book selected by the user.

[0675] An "original story" is a unique narrative created by a generative AI model based on a selected theme.

[0676] "Images" are visual content that is automatically created based on the generated story.

[0677] "PDF format" is a format that integrates the generated story and images into a single document file and saves it in Portable Document Format (PDF).

[0678] An "emotion engine" is a software and hardware system that analyzes a user's facial expressions, voice, etc., and recognizes the user's emotional state.

[0679] A "server" is a computer system that receives user requests and generates stories and images and creates PDF files in response.

[0680] "Means of selection" refers to the interface (e.g., a screen or menu on a website) through which a user selects a theme.

[0681] "Generative means" refers to the AI ​​models or algorithms used to automatically create original stories and images based on selected themes and emotional information.

[0682] "Means for recommending themes and adjusting stories based on emotions" refers to a function that suggests appropriate themes and adjusts the content of stories based on the user's recognized emotional information.

[0683] This system allows users to select a theme, generates an original story and images based on that theme using a generative AI model, and provides a picture book in PDF format. Furthermore, it incorporates an emotion engine that recognizes the user's emotions, and recommends themes and adjusts the story based on this emotion information.

[0684] System configuration

[0685] This system consists of the following elements:

[0686] 1. User Device:

[0687] Browser: Software used to access websites.

[0688] Camera: Hardware that captures the user's facial expressions.

[0689] Microphone: Hardware that captures the user's voice.

[0690] Emotion engine: Software that recognizes user emotions by analyzing facial expressions and voice (e.g., OpenCV, Deep Learning models).

[0691] 2. Server:

[0692] Web server: A computer system that receives requests from users and returns appropriate responses.

[0693] Generative AI models: Language models that generate stories based on selected themes and sentiment information (e.g., GPT-3).

[0694] Illustration generation API: An API for generating images based on a generated story (e.g., DALL-E).

[0695] PDF Generation Library: A library that converts the generated stories and images into PDF format (e.g. FPDF, PDFBox).

[0696] System Operation

[0697] When a user accesses the "AI Storyteller" website using their device, the device sends a connection request to the server, which then returns HTML content including a theme selection screen. The device's built-in emotion engine then activates and analyzes the user's facial expressions and voice to recognize their emotions. For example, the device's camera and microphone can be used to determine whether the user is happy or depressed. The emotion information is then sent to the server, which uses it to recommend themes and generate stories.

[0698] When a user selects a theme and clicks the generate button, the device sends a POST request containing the selected theme data and emotional information to the server. The server receives the POST request and uses a generative AI model based on the selected theme and emotional information to generate an original story. Images appropriate for the generated story are also generated on the server, and further appropriate illustrations are selected based on the emotional information.

[0699] Once the story and images are ready, the server uses a PDF generation library to convert them into a PDF picture book. The server saves the created PDF as a temporary file and sends the path to the file back to the user. The user can click the link to download the PDF and save it on their device.

[0700] Specific examples

[0701] For example, if a user selects the theme "Dragon" and the emotion engine recognizes that the user is feeling a little depressed, the server will generate a story based on this theme and emotion information, such as "A brave boy befriends a dragon and goes on a fun adventure together." The AI ​​then creates bright and hopeful illustrations of the dragon and boy that fit the generated story. The server then combines these illustrations with the story to create a picture book in PDF format and provides the user with a download link. The user can click the link to save the PDF to their device and read it to their child before bedtime.

[0702] Example prompt sentence:

[0703] "Adventures with Dragons"

[0704] "Stories that make children brave"

[0705] "Make friends"

[0706] "Hopeful illustrations"

[0707] This system allows users to easily obtain high-quality original picture books and provide their children with new stories every night. Furthermore, the introduction of an emotion engine allows the system to provide the most appropriate themes and stories according to the user's emotional state, creating a more personalized educational experience.

[0708] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0709] Step 1:

[0710] When a user accesses a website using a device, the device sends an HTTP request to the server. The server receives the request, generates HTML content including a theme selection screen, and sends it back to the device as an HTTP response.

[0711] Input: URL access request from user

[0712] Data processing / calculation: The server analyzes the request and generates appropriate HTML content

[0713] Output: Response showing theme selection screen

[0714] Step 2:

[0715] The emotion engine installed on the device is activated and uses the camera and microphone to capture the user's facial expressions and voice in real time. The emotion engine then uses facial expression and voice analysis software to recognize the user's emotions and generate data.

[0716] Input: User's facial expressions and voice

[0717] Data processing / computation: Capture data using camera and microphone and perform sentiment analysis

[0718] Output: User emotion data

[0719] Step 3:

[0720] The device sends the recognized emotion information to the server, which receives the emotion data and stores it for later processing.

[0721] Input: Emotion data

[0722] Data processing / calculation: Data reception and storage

[0723] Output: Saved emotion data

[0724] Step 4:

[0725] The user selects the desired theme on the theme selection screen and clicks the Generate button. The device then sends a POST request to the server containing the selected theme and the previously recognized emotion data.

[0726] Input: User-selected theme, stored emotion data

[0727] Data processing / calculation: Combining theme selection data and emotion data, generating POST requests

[0728] Output: POST request sent

[0729] Step 5:

[0730] The server receives the POST request and generates an original story using a generative AI model based on the selected theme and emotion data. For example, a prompt sentence is input to a generative AI model (e.g., GPT-3) to generate the story as text.

[0731] Input: Theme data, emotion data

[0732] Data processing / calculation: Use of generative AI models, generation of story text

[0733] Output: Generated story text

[0734] Step 6:

[0735] The server uses an illustration generation API (e.g., DALL-E) based on the generated story to generate images that suit the story, taking into account emotional data to create visuals that are optimal for the user.

[0736] Input: Story text, emotion data

[0737] Data processing / calculation: Calling illustration generation API, image generation

[0738] Output: The generated image

[0739] Step 7:

[0740] The server converts the generated story and images into a PDF book using a PDF generation library (e.g., FPDF, PDFBox). The story and images are inserted into a template and a PDF file is generated.

[0741] Input: Story text, generated images

[0742] Data processing / calculation: Use of PDF generation library, creation of PDF files

[0743] Output: PDF picture book

[0744] Step 8:

[0745] The server saves the created PDF as a temporary file, retains its path, and returns a response including the path of the saved file to the terminal.

[0746] Input: Generated PDF

[0747] Data processing / calculation: file saving, path generation

[0748] Output: Response to device, download link

[0749] Step 9:

[0750] The device parses the response received from the server and displays a link to download the PDF to the user. When the user clicks the link, the PDF is saved on the device.

[0751] Input: Response from the server

[0752] Data processing / calculation: Response analysis, link display

[0753] Output: Download and save PDF

[0754] (Application example 2)

[0755] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0756] Conventional illustrated work instruction generation systems create formulaic instructions without considering the user's emotional state, which does not adequately improve worker motivation or provide emotional care. Furthermore, instructions that do not take emotions into consideration may affect the efficiency and effectiveness of work. Therefore, there is a need to provide customized work instructions and work plans that correspond to the emotional state of workers.

[0757] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for a user to select a theme, means for generating an original story based on the selected theme, means for generating images based on the generated story, means for converting the generated story and images into PDF format, means for the user to download or print the PDF format document, means for recognizing a user's emotion, means for adjusting the story and images based on the recognized emotion, means for recognizing a user's emotion in an industrial environment and adjusting work instructions and work plans, and means for converting the generated work instructions and work plans into PDF format. This makes it possible to provide work instructions and work plans customized according to the emotional state of the worker.

[0758] A "theme" is the content or subject matter that a user selects to create an original story.

[0759] A "story" is a narrative or story content that is generated based on a selected theme.

[0760] "Images" refers to visual content created based on the generated story, including visual information such as pictures and illustrations.

[0761] "PDF format" is an abbreviation for Portable Document Format, and refers to a fixed-layout file format suitable for storing and distributing electronic documents.

[0762] A "document" is a written document that contains textual information such as a story or instructions.

[0763] "Emotions" represent the user's psychological state or sensations, and include types such as joy, sadness, surprise, and anger.

[0764] An "emotion engine" is a software or hardware mechanism that uses sensors such as cameras and microphones to recognize and analyze a user's emotional state.

[0765] "Industrial environment" refers to the working environment at a production or manufacturing site or facility.

[0766] A "work instruction" is a document that describes the work content and procedures that workers should perform in an industrial environment.

[0767] A "work plan" is a document that includes detailed processes and schedules for efficiently carrying out a specific task.

[0768] A "server" is a computer system that provides services over a network, and refers to the equipment and software that processes and distributes data in response to requests.

[0769] A "generative AI model" is an algorithm or model that uses artificial intelligence technology to generate new content (stories or images) based on input data.

[0770] The present invention combines a system in which a user selects a theme, AI generates original stories and images based on that theme, and provides them in PDF format with an emotion engine that recognizes the user's emotions. This system can provide customized work instructions and work plans according to the emotional state of workers, especially in industrial environments.

[0771] Hardware and software used

[0772] Camera: A device that captures images for emotion recognition.

[0773] Microphone: A device that captures the user's voice emotions if necessary.

[0774] Speech recognition engine: A software component that analyzes the user's emotions.

[0775] EmotionRecognizer: A library that analyzes facial expression data and recognizes user emotions.

[0776] Generative AI Model: An AI algorithm for generating stories and images based on a theme (StoryGenerator, ImageGenerator).

[0777] PDFCreator: A library that converts generated stories and images into PDF format.

[0778] Natural language processing explanation

[0779] 1. User selects a theme:

[0780] A user uses a terminal to select a theme in an industrial environment (e.g., inspection work on line B). This theme indicates the work content and details of the workers.

[0781] 2. Emotion recognition:

[0782] The device's built-in camera and microphone are used to acquire the user's facial expression and voice data. The EmotionRecognizer library is used to recognize the user's emotions (e.g., joy, sadness, anger) from this data. This emotional information is sent to the server and used to generate stories and images.

[0783] 3. Story and image generation:

[0784] Based on the recognized emotional information and the selected theme, the server uses a Generative AI Model (Story Generator, Image Generator) to generate a customized story and associated images.

[0785] 4. Generate and serve PDF:

[0786] The generated stories and images are used to create PDF work instructions and work plans using the PDFCreator library, which are then saved on the server and a download link is provided to the user.

[0787] Specific examples

[0788] For example, if a user selects "Inspection work on Line B" as the theme and the emotion engine recognizes that the user is feeling a little down, the AI ​​can use this information to suggest "Perform equipment inspection on Line B. The theme will be adjusted to include simple and easy-to-understand instructions and encouraging messages." This allows the user to receive customized instructions that reflect their emotional state that day, improving work efficiency.

[0789] Prompt Sentence Examples

[0790] "Today, the user is tasked with inspecting the machines on Line B. If the user is feeling down, generate a story that gives them simple instructions with an encouraging message."

[0791] "Emotion: Sad\nTask Details: Machine Inspection\nGenerate a positive and encouraging instruction set."

[0792] In this way, the system of the present invention can provide customized support according to the user's emotional state, contributing to improving work efficiency and maintaining worker motivation, particularly in industrial environments.

[0793] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0794] Step 1:

[0795] The user selects a theme on the terminal. From the theme list displayed on the terminal screen, the user selects a theme for a work instruction in an industrial environment (e.g., inspection work on line B). The input here is the theme selected by the user, and the output is detailed information about the selected theme. This information is passed to the subsequent processing step.

[0796] Step 2:

[0797] The system recognizes the user's emotions using the device's camera and microphone. The system uses the device's hardware (camera and microphone) to acquire the user's facial expression data and voice data. The system uses the EmotionRecognizer library to analyze the user's emotions from this data. The input is the acquired facial expression data and voice data, and the output is analyzed emotional information (e.g., joy, sadness, anger).

[0798] Step 3:

[0799] The recognized emotion information is sent to the server. The device then sends the analyzed emotion information to the server. At this time, the theme information selected by the user is also sent to the server. The emotion information and theme information are input, and the emotion information and theme information received by the server are output.

[0800] Step 4:

[0801] The server generates the story and images. Based on the theme and emotion information received on the server side, a generative AI model (Story Generator, Image Generator) is used to generate a customized story and images. The theme and emotion information are input here, and the generated story and images are output. Specifically, the AI ​​model creates a story while taking emotion information into account, and generates a prompt that generates an illustration based on the story.

[0802] Step 5:

[0803] The generated story and images are converted into PDF format. The server uses the generated story and images to convert them into a PDF document using the PDFCreator library. The input data is the generated story and images, and the output is a PDF document. Specifically, PDFCreator combines the story text and images to create a single PDF file.

[0804] Step 6:

[0805] The generated PDF is provided to the user. The server saves the generated PDF file as a temporary file and provides the link to the user. The input is the PDF file, and the output is a link to the PDF file that the user can access. The user can download or print the PDF from this link on their device.

[0806] Step 7:

[0807] The user downloads or prints a PDF. The device clicks on a link to download the PDF file and optionally prints it. The input to this step is the PDF link provided by the server, and the output is the downloaded PDF file or printed document. Specifically, the device opens the link, saves the file, and sends it to the printer to print the physical document.

[0808] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0809] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0810] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0811] [Third embodiment]

[0812] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0813] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0814] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0815] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0816] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0817] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0818] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0819] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0820] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0821] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0822] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0823] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0824] This invention provides a system in which a user selects a theme for a picture book, AI generates an original story and images based on that theme, and the picture book can be downloaded or printed in PDF format. The program and processing of this system are explained in detail below.

[0825] System Overview

[0826] When a user accesses the "AI Storyteller" website using a device, a theme selection screen appears. The user selects the desired theme and presses the generate button. The AI ​​then generates an original story and images based on the selected theme. The generated story and images are then converted into PDF format, which the user can download or print.

[0827] Program processing flow

[0828] 1. When a user accesses a website using a device, the device sends a connection request to the server, and the server returns HTML content including a theme selection screen. The device displays the returned HTML content on the screen, allowing the user to select a theme.

[0829] 2. When the user selects a theme and clicks the generate button, the device sends the selected theme data as a POST request to the server. The server receives the POST request and generates an original story based on the selected theme using an AI model.

[0830] 3. The server uses an AI model to generate illustrations based on the generated story, and combines the generated story and illustrations to create a picture book in PDF format.

[0831] 4. The created PDF is saved as a temporary file by the server, and a response containing a link that the user can download is sent back to the device.

[0832] 5. When the user clicks the download link on their device, the PDF is saved to their device. The user can view, download, and print the saved PDF.

[0833] Specific examples

[0834] For example, if a user selects the theme "Dragon," the server generates a story based on this theme, such as "A brave boy makes friends with a dragon." The AI ​​then creates illustrations of the dragon and boy that fit the generated story. The server then combines these illustrations with the story to create a picture book in PDF format and provides the user with a download link. The user can click the link to save the PDF on their device and read it to their child before bedtime.

[0835] This system eliminates the need for users to frequently purchase expensive picture books, allowing them to provide their children with a new story every night. Furthermore, by setting themes with educational messages, it is possible to effectively educate children. In this way, the present invention alleviates parental concerns and provides an efficient means for supporting children's curiosity and learning.

[0836] The processing flow will be explained below.

[0837] Step 1:

[0838] A user accesses the "AI Storyteller" website using a device. The device sends a connection request to the web server, and the web server returns HTML content including a theme selection screen.

[0839] Step 2:

[0840] The device parses the HTML received from the server and displays the theme selection screen, allowing the user to see the theme selection form.

[0841] Step 3:

[0842] The user selects the desired theme from a drop-down menu, for example, one of the options "adventure," "dragon," or "spaceship."

[0843] Step 4:

[0844] When the user presses the "Generate" button, the form data is sent from the terminal to the server as a POST request, including the selected theme.

[0845] Step 5:

[0846] The server receives the POST request and retrieves the selected theme data from the request, which is then used for the next process.

[0847] Step 6:

[0848] The server uses an AI model to generate an original story based on the selected theme. The server calls an AI library (e.g., OpenAI GPT-3) and executes the story generation API.

[0849] Step 7:

[0850] The server generates images appropriate for the generated story. The server again uses an AI model (e.g., DALL-E) to execute an illustration generation API based on the content of the story.

[0851] Step 8:

[0852] The server creates a PDF book based on the generated story and images. It uses a PDF generation library (e.g., FPDF) to insert the story text and images into the PDF document.

[0853] Step 9:

[0854] The server saves the completed PDF as a temporary file and retains its path, which is later provided to the user.

[0855] Step 10:

[0856] The server sends a response back to the device containing the path to the generated PDF, along with a link to download the PDF.

[0857] Step 11:

[0858] The device parses the response received from the server and displays a PDF download link to the user, which the user can click to download the PDF.

[0859] Step 12:

[0860] When a user clicks the download link, the device saves the PDF file to local storage. By opening the saved PDF, the user can view, download, and print the generated original picture book.

[0861] Through this series of processes, users can easily obtain high-quality original picture books and provide their children with new stories every night.

[0862] Example 1

[0863] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0864] In modern child-rearing, parents want to provide their children with a variety of fresh picture books that meet their children's interests and educational needs. However, frequently purchasing new picture books is a significant financial burden, and commercially available picture books do not always contain the educational messages parents desire. Furthermore, busy parents find it difficult to quickly obtain personalized picture books tailored to their children's preferences. Efficient methods to solve these problems are needed.

[0865] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0866] In this invention, the server includes means for a user to select a theme, means for generating an original story using a generative AI model based on the selected theme, means for generating images using the generative AI model based on the content of the generated story, means for converting the generated story and images into a PDF format, and means for the user to download or print the PDF picture book. This allows a user to easily generate an original picture book based on a theme and download or print it.

[0867] A "theme" is a subject or topic selected by the user that will be the basis for the story or image that will be generated.

[0868] A "generative AI model" is an artificial intelligence algorithm or system that generates text or images based on a selected theme.

[0869] An "original story" is a unique narrative generated by a generative AI model based on a selected theme.

[0870] "Images" are visual illustrations or pictures generated by a generative AI model to match the content of the original story.

[0871] "PDF format" is an abbreviation for Portable Document Format, and is a file format for saving and displaying documents and images on a page-by-page basis.

[0872] A picture book is a book for children that contains stories and illustrations in written form.

[0873] "Downloading" is the act of saving a file from the Internet to your device.

[0874] "Printing" is the act of printing digital data onto physical media such as paper.

[0875] A "user" is someone who uses the system to select a theme, create, download, and print a picture book.

[0876] A "server" is a computer system that receives requests from users and generates, stores, and delivers appropriate content.

[0877] A "terminal" is a device operated by a user, and is a computer that communicates with a server and displays and saves results.

[0878] The present invention is a system that uses a generative AI model to generate original stories and images based on a theme selected by a user, and provides them as a PDF picture book. The system includes a means for a user to select a theme, a means for converting the generated stories and images into PDF format, and a means for the user to download or print the PDF picture book.

[0879] Hardware and Software

[0880] The system is implemented using a server, a terminal, a generative AI model, and various software. The server receives requests, generates stories and images, converts them into PDF format, and distributes them. The terminal acts as an interface for users to access the system, select a theme, and download the PDF.

[0881] The main hardware and software used are as follows:

[0882] Server: Processes the data and invokes the generative AI model.

[0883] Terminal: The device through which the user accesses the system, such as a regular PC, smartphone, or tablet.

[0884] Generative AI models: For example, ChatGPT and GPT-4 are used for story generation, while DALL-E and Stable Diffusion are used for image generation.

[0885] PDF generation library: Generate PDFs using Python's ReportLab, etc.

[0886] Data processing and calculation

[0887] 1. Theme selection: The user selects a theme using the terminal. For example, the user can select the theme "Dragon."

[0888] 2. Story generation: The server uses a generative AI model (such as ChatGPT or GPT-4) to generate an original story based on the selected theme. An example of a specific prompt is "Theme: Dragons. Please generate a story about a brave boy who becomes friends with a dragon."

[0889] 3. Image Generation: The server uses a generative AI model (such as DALL-E or Stable Diffusion) to generate images that match the generated story. The prompt might be something like, "Generate an illustration of a dragon and a boy."

[0890] 4. PDF generation: The server combines the generated story and images and generates a PDF picture book using Python's ReportLab or similar.

[0891] 5. PDF Delivery: The server saves the PDF in temporary storage and provides a download link that users can access, and they can download or print the PDF using their devices.

[0892] Specific examples

[0893] For example, if a user selects "Dragons" as the theme, the process goes as follows: The user selects a theme through the website and presses the generate button. The device then sends the theme data to the server. The server uses a generative AI model to generate a story such as "A brave boy befriends a dragon" and creates illustrations of the dragon and boy based on that story. The server then assembles this content into a PDF picture book using Python's ReportLab and provides the user with a download link. The user can click this link to download the PDF and read it to their child.

[0894] This system eliminates the need for users to frequently purchase expensive picture books, allowing them to provide their children with a new story every night. Furthermore, by setting themes with educational messages, it is possible to effectively educate children. In this way, the present invention alleviates parental concerns and provides an efficient means for supporting children's curiosity and learning.

[0895] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0896] Specific flow of system program processing

[0897] Step 1:

[0898] The user accesses the "AI Storyteller" website using a device.

[0899] Input: An HTTP request sent by a user typing a URL into a web browser.

[0900] Output: The device sends a connection request to the server, and the server returns a theme selection screen.

[0901] Specific behavior:

[0902] The user opens a browser, enters the specified URL, and presses enter.

[0903] The device sends this operation to the server as an HTTP request.

[0904] The server receives the request and generates HTML content including a theme selection screen.

[0905] The server returns the generated HTML content to the terminal as an HTTP response.

[0906] The terminal displays the received HTML content on the screen and allows the user to select a theme.

[0907] Step 2:

[0908] The user selects a theme and clicks the generate button.

[0909] Input: The user selects a theme and clicks the generate button.

[0910] Output: The device sends the selected theme data to the server as a POST request.

[0911] Specific behavior:

[0912] The user selects a desired theme from the displayed theme selection screen.

[0913] The user clicks the "Generate" button.

[0914] The device sends the user's selected theme data to the server as a POST request in JSON format.

[0915] Step 3:

[0916] The server receives the POST request and generates the story.

[0917] Input: A POST request containing theme data sent from the device.

[0918] Output: The original story generated using the generative AI model.

[0919] Specific behavior:

[0920] The server receives the POST request and extracts the theme data from the request body.

[0921] The server sends a prompt to the generative AI model (e.g., ChatGPT or GPT-4). Example prompt: "Theme: Dragons. Generate a story about a brave boy who becomes friends with a dragon."

[0922] The generative AI model generates a story based on the prompt sentence and returns the story to the server.

[0923] Step 4:

[0924] The server generates illustrations based on the story.

[0925] Input: Generated stories.

[0926] Output: An illustration generated using a generative AI model.

[0927] Specific behavior:

[0928] The server analyzes the content of the generated story and generates a prompt accordingly. Example prompt: "Please generate an illustration of a dragon and a boy."

[0929] The server sends a prompt to a generative AI model (e.g., DALL-E or Stable Diffusion).

[0930] The generative AI model generates an illustration based on the prompt sentence and returns the illustration to the server.

[0931] Step 5:

[0932] The server creates a picture book in PDF format.

[0933] Input: Generated story and illustrations.

[0934] Output: Picture book in PDF format.

[0935] Specific behavior:

[0936] The server combines the generated story and illustrations and creates a PDF picture book using Python's ReportLab etc.

[0937] The server saves the generated PDF file in temporary storage.

[0938] Step 6:

[0939] The server provides a download link for the PDF.

[0940] Input: PDF file stored on the server.

[0941] Output: A response containing a download link for the PDF.

[0942] Specific behavior:

[0943] The server generates a download link for the PDF file stored in temporary storage.

[0944] The server sends a response including the download link back to the device.

[0945] Step 7:

[0946] The user downloads or prints the PDF.

[0947] Input: Click on the download link displayed on your device.

[0948] Output: Downloaded PDF file.

[0949] Specific behavior:

[0950] The user clicks on the download link displayed in the device's browser.

[0951] The device downloads the PDF file from the server.

[0952] The user can view the PDF file stored on the device and print it if necessary.

[0953] Through the above specific processing steps, the system generates an original picture book based on the theme selected by the user, and makes it easy to download and print.

[0954] (Application example 1)

[0955] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0956] Currently, many families and educational institutions need to regularly purchase picture books for children. However, existing picture books have fixed content, which can make it difficult to keep children interested. Furthermore, because picture books are physical, collecting a large number of them requires storage space and is costly. Furthermore, parents and educators want to provide educational content that responds to children's development and interests. For these reasons, there is a need for a cost-effective system that can flexibly generate content.

[0957] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0958] In this invention, the server includes means for a user to select a theme, means for generating an original story based on the selected theme, means for generating visual materials based on the generated story, means for converting the generated story and visual materials into an electronic document format, means for the user to download or print the picture book in the electronic document format, and means for the user to select a theme and preview and check the generated picture book in real time using a touch panel terminal, thereby enabling users to generate and print personalized picture books tailored to their wishes on the spot in the store.

[0959] "User" refers to the person who uses this system to select a picture book theme, generate content, download it, and print it.

[0960] "Theme" refers to the subject that forms the basis of the content or subject matter of the picture book selected by the user.

[0961] "Story" refers to an original story generated by AI based on a selected theme.

[0962] "Visual materials" refer to illustrations and images created by AI based on the generated story.

[0963] "Electronic document format" refers to a digital document, such as a PDF format, that integrates and converts the generated narrative and visual materials.

[0964] "Downloading" refers to the process of saving the generated picture book in electronic document format to the user's terminal via the Internet.

[0965] "Printing" refers to the process of outputting the generated picture book in electronic document format onto physical paper media.

[0966] A "touch panel terminal" refers to a computer device equipped with a display that the user can operate by directly touching it.

[0967] "Real-time" refers to the timing at which data is generated and processed immediately.

[0968] "Preview" refers to a temporary display or sample that is displayed to confirm the final generated result.

[0969] "System" refers to the entire platform that provides a set of functions that allow users to generate, download, and print picture books.

[0970] This invention is a system in which a user selects a theme, an AI generates an original story and visual materials based on that theme, and provides a picture book in electronic document format. Specific embodiments for implementing this system will be described below.

[0971] System Configuration

[0972] The present invention is configured by a series of hardware and software components including a user terminal, a server, and a touch panel terminal.

[0973] 1. User Device:

[0974] A device that allows users to select a theme and download or print the generated picture book. Examples include computers, smartphones, and tablets.

[0975] 2. Server:

[0976] It is a central device that receives user requests and generates stories and visual materials using AI models. The server is installed with a language model (e.g., GPT-4) for story generation and an image model (e.g., DALL-E) for visual material generation.

[0977] 3. Touchscreen terminal:

[0978] This device is installed in educational institutions and brick-and-mortar stores, allowing users to select a theme and preview the generated picture book in real time.

[0979] Processing Overview

[0980] When a user accesses the system using a terminal, the touch panel terminal displays a theme selection screen. The user selects the desired theme and presses the generate button on the terminal, and the selected theme data is sent to the server. The server uses a generative AI model to generate an original story based on the selected theme. It then generates illustrations and images based on the generated story. The generated story and visual materials are converted into a digital document in PDF format, which the user can download or print.

[0981] Hardware and Software Use

[0982] Hardware: Touchscreen devices (e.g., Surface Pro), user devices (e.g., iPhone, Android tablet), servers (e.g., AWS EC2 instances)

[0983] Software: The server-side program is implemented in Python and uses web frameworks such as Flask, and also uses GPT-4 for story generation and DALL-E for visual material generation.

[0984] Data processing and calculation

[0985] The server inputs the topic information sent by the user into a natural language processing model and generates a story. Specifically, it uses the following prompt sentences:

[0986] Example prompt: "Generate a story based on the following theme: Dragons."

[0987] Based on the generated story, a request for illustration generation is sent to the visual material generation model. The generated visual material and story are integrated and converted into a PDF file. The PDF is temporarily stored on the server, and a download link is provided to the user.

[0988] Specific examples

[0989] For example, if a user selects the "Dragon" theme using a touchscreen device in a physical store, the server proceeds as follows:

[0990] 1. The user selects the "Dragons" theme.

[0991] 2. The server sends the prompt "Generate a story based on the following theme: Dragons" to the generative AI model.

[0992] 3. The generative AI model generates the story and sends it back to the server.

[0993] 4. The server inputs the story into a visual material generation model and generates illustrations related to dragons.

[0994] 5. Integrate the generated narrative and visual materials and convert them into a PDF file.

[0995] 6. The user downloads the generated PDF or prints it on an in-store printer.

[0996] This system allows users to flexibly create and use personalized picture books on the spot.

[0997] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0998] Step 1:

[0999] The user uses the device to access the theme selection screen on the touch panel device. The server returns HTML content including the theme selection screen, which the device displays. This stage mainly involves collecting user input. Input: None, Output: Display of theme selection screen.

[1000] Step 2:

[1001] The user selects a theme and clicks the generate button on the touch panel terminal. The terminal sends the selected theme data to the server as a POST request. At this stage, the theme selected by the user is sent to the server. Input: Theme selected by the user, Output: Theme data sent to the server.

[1002] Step 3:

[1003] The server processes the received POST request and sends a prompt to the generative AI model based on the selected theme. For example, it sends a prompt such as "Generate a story based on the following theme: Dragons." Input: Theme data, Output: Prompt to the generative AI model.

[1004] Step 4:

[1005] The generative AI model generates an original story based on the server's request and sends it back to the server. The server receives the generated story. Input: prompt sentence, output: generated story.

[1006] Step 5:

[1007] Based on the generated story, the server sends a request to the visual material generation model to generate visual materials (e.g., illustrations and images) appropriate for the story. Input: Generated story, Output: Request to the visual material generation model.

[1008] Step 6:

[1009] The visual material generation model generates visual material based on the server's request and sends it back to the server. The server receives the generated visual material. Input: request, Output: generated visual material.

[1010] Step 7:

[1011] The server combines the generated story and visual materials and converts them into an electronic document in PDF format. Input: Generated story and visual materials, Output: Electronic document in PDF format.

[1012] Step 8:

[1013] The server saves the generated PDF as a temporary file and returns a response containing a download link to the device. Input: PDF file, Output: Response containing a download link.

[1014] Step 9:

[1015] The user clicks the download link on their device to download or print the PDF. At this stage, the final product is obtained by the user. Input: Download link, Output: Save or print the PDF file.

[1016] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1017] This invention combines a system in which a user selects a theme for a picture book, AI generates an original story and images based on that theme, and the picture book can be downloaded or printed in PDF format, with an emotion engine that recognizes the user's emotions. The program and processing of this system are explained in detail below.

[1018] System Overview

[1019] When a user accesses the "AI Storyteller" website using a device, a theme selection screen is displayed. The user selects the desired theme and presses the generate button. The AI ​​then generates an original story and images based on the selected theme. The generated story and images are then converted into PDF format, which the user can download or print. The present invention also incorporates an emotion engine that recognizes the user's emotions, allowing it to recommend themes and adjust the story based on the user's emotions.

[1020] Program processing flow

[1021] 1. When a user accesses a website using a device, the device sends a connection request to the server, and the server returns HTML content including a theme selection screen. The device displays the returned HTML content on the screen, allowing the user to select a theme.

[1022] 2. The emotion engine installed in the device analyzes the user's facial expressions and voice to recognize their emotions. For example, it uses the device's camera and microphone to identify whether the user is happy, depressed, surprised, etc.

[1023] 3. The emotion engine sends the recognized emotion information to the server, which receives the emotion information and uses it to recommend themes and generate stories.

[1024] 4. When the user selects a theme and clicks the generate button, the device sends a POST request containing the selected theme data and emotion information to the server. The server receives the POST request and uses an AI model based on the selected theme and emotion information to generate an original story.

[1025] 5. The server generates images appropriate for the generated story. The server again uses the AI ​​model to execute an illustration generation API based on the content of the story. Emotional information is also taken into account here, and an illustration appropriate for the emotion is generated.

[1026] 6. The server creates a PDF book based on the generated story and images. It uses a PDF generation library to insert the story text and images into a PDF document.

[1027] 7. The created PDF is saved as a temporary file by the server and its path is saved. This temporary file is later provided to the user.

[1028] 8. The server sends a response back to the device containing the path to the generated PDF and a link to download the PDF.

[1029] 9. The device parses the response received from the server and displays a link to download the PDF to the user. When the user clicks the download link, the PDF is saved to the device.

[1030] Specific examples

[1031] For example, if a user selects the theme "Dragon" and the emotion engine recognizes that the user is feeling a little depressed, the server will generate a story based on this theme and emotion information, such as "A brave boy befriends a dragon and goes on a fun adventure together." The AI ​​then creates cheerful, hopeful illustrations of the dragon and boy that fit the generated story. The server then combines these illustrations with the story to create a picture book in PDF format and provides the user with a download link. The user can click the link to save the PDF to their device and read it to their child before bedtime.

[1032] This system allows users to easily obtain high-quality original picture books and provide their children with new stories every night. Furthermore, the introduction of an emotion engine allows the system to provide optimal themes and stories according to the user's emotional state, creating a more personalized educational experience.

[1033] The processing flow will be explained below.

[1034] Step 1:

[1035] A user accesses the "AI Storyteller" website using a device. The device sends a connection request to the web server, and the server returns HTML content including a theme selection screen.

[1036] Step 2:

[1037] The device parses the HTML received from the server and displays the theme selection screen, allowing the user to see the theme selection form.

[1038] Step 3:

[1039] The emotion engine installed in the device analyzes the user's facial expressions and voice to recognize their emotions. For example, it uses the device's camera and microphone to identify whether the user is happy, depressed, surprised, etc.

[1040] Step 4:

[1041] The emotion engine sends the emotional information it recognizes to the server, which then recommends themes appropriate for the user and adjusts story generation based on the emotional information received.

[1042] Step 5:

[1043] The user selects the desired theme from a drop-down menu, for example, one of the options "adventure," "dragon," or "spaceship."

[1044] Step 6:

[1045] When the user presses the "Generate" button, the form data is sent from the device to the server as a POST request, which includes the selected theme and emotion information.

[1046] Step 7:

[1047] The server receives the POST request and retrieves the selected theme and emotion information from the request, which is then used for the next process.

[1048] Step 8:

[1049] The server uses an AI model to generate an original story based on the selected theme and emotional information. The server calls an AI library (e.g., OpenAI GPT-3) and executes the story generation API.

[1050] Step 9:

[1051] The server generates images appropriate for the generated story. The server again uses an AI model (such as DALL-E) to execute an illustration generation API based on the story content and emotional information.

[1052] Step 10:

[1053] The server creates a PDF book based on the generated story and images. It uses a PDF generation library (such as FPDF) to insert the story text and images into the PDF document.

[1054] Step 11:

[1055] The server saves the completed PDF as a temporary file and retains its path, which is later provided to the user.

[1056] Step 12:

[1057] The server sends a response back to the device containing the path to the generated PDF, along with a link to download the PDF.

[1058] Step 13:

[1059] The device parses the response received from the server and displays a link to download the PDF to the user. When the user clicks the link, the PDF is saved on the device.

[1060] Step 14:

[1061] By opening the saved PDF, users can view, download, and print the original picture book they created. The emotional engine allows the system to provide the most appropriate theme and story based on the user's emotional state, enabling a personalized educational experience.

[1062] Through this series of steps, users can easily obtain high-quality original picture books and provide their children with new stories every night. The emotional engine generates themes and stories that best suit the user's emotional state, providing a more personalized educational experience.

[1063] Example 2

[1064] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1065] In conventional picture book generation systems, users simply select a theme, and the generated story and images are not optimized for the user's emotional state. As a result, picture books with content that does not match the user's emotions or mood are generated, resulting in low satisfaction. Furthermore, because emotion recognition technology is not incorporated, personalized educational experiences are not provided.

[1066] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1067] In this invention, the server includes: means for a user to select a theme; means for generating an original story based on the selected theme; means for generating images based on the generated story; means for converting the generated story and images into PDF format; means for the user to download or print the PDF picture book; and means including an emotion engine that recognizes the user's emotions and that recommends themes and adjusts the story based on the recognized emotions. This automatically generates an original picture book suited to the user's emotional state, enabling the user to have a highly satisfying and personalized educational experience.

[1068] A "user" is an individual who uses the system to create a picture book.

[1069] A "theme" is a topic or category that determines the content of the picture book selected by the user.

[1070] An "original story" is a unique narrative created by a generative AI model based on a selected theme.

[1071] "Images" are visual content that is automatically created based on the generated story.

[1072] "PDF format" is a format that integrates the generated story and images into a single document file and saves it in Portable Document Format (PDF).

[1073] An "emotion engine" is a software and hardware system that analyzes a user's facial expressions, voice, etc., and recognizes the user's emotional state.

[1074] A "server" is a computer system that receives user requests and generates stories and images and creates PDF files in response.

[1075] "Means of selection" refers to the interface (e.g., a screen or menu on a website) through which a user selects a theme.

[1076] "Generative means" refers to the AI ​​models or algorithms used to automatically create original stories and images based on selected themes and emotional information.

[1077] "Means for recommending themes and adjusting stories based on emotions" refers to a function that suggests appropriate themes and adjusts the content of stories based on the user's recognized emotional information.

[1078] This system allows users to select a theme, generates an original story and images based on that theme using a generative AI model, and provides a picture book in PDF format. Furthermore, it incorporates an emotion engine that recognizes the user's emotions, and recommends themes and adjusts the story based on this emotion information.

[1079] System configuration

[1080] This system consists of the following elements:

[1081] 1. User Device:

[1082] Browser: Software used to access websites.

[1083] Camera: Hardware that captures the user's facial expressions.

[1084] Microphone: Hardware that captures the user's voice.

[1085] Emotion engine: Software that recognizes user emotions by analyzing facial expressions and voice (e.g., OpenCV, Deep Learning models).

[1086] 2. Server:

[1087] Web server: A computer system that receives requests from users and returns appropriate responses.

[1088] Generative AI models: Language models that generate stories based on selected themes and sentiment information (e.g., GPT-3).

[1089] Illustration generation API: An API for generating images based on a generated story (e.g., DALL-E).

[1090] PDF Generation Library: A library that converts the generated stories and images into PDF format (e.g. FPDF, PDFBox).

[1091] System Operation

[1092] When a user accesses the "AI Storyteller" website using their device, the device sends a connection request to the server, which then returns HTML content including a theme selection screen. The device's built-in emotion engine then activates and analyzes the user's facial expressions and voice to recognize their emotions. For example, the device's camera and microphone can be used to determine whether the user is happy or depressed. The emotion information is then sent to the server, which uses it to recommend themes and generate stories.

[1093] When a user selects a theme and clicks the generate button, the device sends a POST request containing the selected theme data and emotional information to the server. The server receives the POST request and uses a generative AI model based on the selected theme and emotional information to generate an original story. Images appropriate for the generated story are also generated on the server, and further appropriate illustrations are selected based on the emotional information.

[1094] Once the story and images are ready, the server uses a PDF generation library to convert them into a PDF picture book. The server saves the created PDF as a temporary file and sends the path to the file back to the user. The user can click the link to download the PDF and save it on their device.

[1095] Specific examples

[1096] For example, if a user selects the theme "Dragon" and the emotion engine recognizes that the user is feeling a little depressed, the server will generate a story based on this theme and emotion information, such as "A brave boy befriends a dragon and goes on a fun adventure together." The AI ​​then creates bright and hopeful illustrations of the dragon and boy that fit the generated story. The server then combines these illustrations with the story to create a picture book in PDF format and provides the user with a download link. The user can click the link to save the PDF to their device and read it to their child before bedtime.

[1097] Example prompt sentence:

[1098] "Adventures with Dragons"

[1099] "Stories that make children brave"

[1100] "Make friends"

[1101] "Hopeful illustrations"

[1102] This system allows users to easily obtain high-quality original picture books and provide their children with new stories every night. Furthermore, the introduction of an emotion engine allows the system to provide the most appropriate themes and stories according to the user's emotional state, creating a more personalized educational experience.

[1103] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1104] Step 1:

[1105] When a user accesses a website using a device, the device sends an HTTP request to the server. The server receives the request, generates HTML content including a theme selection screen, and sends it back to the device as an HTTP response.

[1106] Input: URL access request from user

[1107] Data processing / calculation: The server analyzes the request and generates appropriate HTML content

[1108] Output: Response showing theme selection screen

[1109] Step 2:

[1110] The emotion engine installed on the device is activated and uses the camera and microphone to capture the user's facial expressions and voice in real time. The emotion engine then uses facial expression and voice analysis software to recognize the user's emotions and generate data.

[1111] Input: User's facial expressions and voice

[1112] Data processing / computation: Capture data using camera and microphone and perform sentiment analysis

[1113] Output: User emotion data

[1114] Step 3:

[1115] The device sends the recognized emotion information to the server, which receives the emotion data and stores it for later processing.

[1116] Input: Emotion data

[1117] Data processing / calculation: Data reception and storage

[1118] Output: Saved emotion data

[1119] Step 4:

[1120] The user selects the desired theme on the theme selection screen and clicks the Generate button. The device then sends a POST request to the server containing the selected theme and the previously recognized emotion data.

[1121] Input: User-selected theme, stored emotion data

[1122] Data processing / calculation: Combining theme selection data and emotion data, generating POST requests

[1123] Output: POST request sent

[1124] Step 5:

[1125] The server receives the POST request and generates an original story using a generative AI model based on the selected theme and emotion data. For example, a prompt sentence is input to a generative AI model (e.g., GPT-3) to generate the story as text.

[1126] Input: Theme data, emotion data

[1127] Data processing / calculation: Use of generative AI models, generation of story text

[1128] Output: Generated story text

[1129] Step 6:

[1130] The server uses an illustration generation API (e.g., DALL-E) based on the generated story to generate images that suit the story, taking into account emotional data to create visuals that are optimal for the user.

[1131] Input: Story text, emotion data

[1132] Data processing / calculation: Calling illustration generation API, image generation

[1133] Output: The generated image

[1134] Step 7:

[1135] The server converts the generated story and images into a PDF book using a PDF generation library (e.g., FPDF, PDFBox). The story and images are inserted into a template and a PDF file is generated.

[1136] Input: Story text, generated images

[1137] Data processing / calculation: Use of PDF generation library, creation of PDF files

[1138] Output: PDF picture book

[1139] Step 8:

[1140] The server saves the created PDF as a temporary file, retains its path, and returns a response including the path of the saved file to the terminal.

[1141] Input: Generated PDF

[1142] Data processing / calculation: file saving, path generation

[1143] Output: Response to device, download link

[1144] Step 9:

[1145] The device parses the response received from the server and displays a link to download the PDF to the user. When the user clicks the link, the PDF is saved on the device.

[1146] Input: Response from the server

[1147] Data processing / calculation: Response analysis, link display

[1148] Output: Download and save PDF

[1149] (Application example 2)

[1150] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1151] Conventional illustrated work instruction generation systems create formulaic instructions without considering the user's emotional state, which does not adequately improve worker motivation or provide emotional care. Furthermore, instructions that do not take emotions into consideration may affect the efficiency and effectiveness of work. Therefore, there is a need to provide customized work instructions and work plans that correspond to the emotional state of workers.

[1152] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for a user to select a theme, means for generating an original story based on the selected theme, means for generating images based on the generated story, means for converting the generated story and images into PDF format, means for the user to download or print the PDF format document, means for recognizing a user's emotion, means for adjusting the story and images based on the recognized emotion, means for recognizing a user's emotion in an industrial environment and adjusting work instructions and work plans, and means for converting the generated work instructions and work plans into PDF format. This makes it possible to provide work instructions and work plans customized according to the emotional state of the worker.

[1153] A "theme" is the content or subject matter that a user selects to create an original story.

[1154] A "story" is a narrative or story content that is generated based on a selected theme.

[1155] "Images" refers to visual content created based on the generated story, including visual information such as pictures and illustrations.

[1156] "PDF format" is an abbreviation for Portable Document Format, and refers to a fixed-layout file format suitable for storing and distributing electronic documents.

[1157] A "document" is a written document that contains textual information such as a story or instructions.

[1158] "Emotions" represent the user's psychological state or sensations, and include types such as joy, sadness, surprise, and anger.

[1159] An "emotion engine" is a software or hardware mechanism that uses sensors such as cameras and microphones to recognize and analyze a user's emotional state.

[1160] "Industrial environment" refers to the working environment at a production or manufacturing site or facility.

[1161] A "work instruction" is a document that describes the work content and procedures that workers should perform in an industrial environment.

[1162] A "work plan" is a document that includes detailed processes and schedules for efficiently carrying out a specific task.

[1163] A "server" is a computer system that provides services over a network, and refers to the equipment and software that processes and distributes data in response to requests.

[1164] A "generative AI model" is an algorithm or model that uses artificial intelligence technology to generate new content (stories or images) based on input data.

[1165] The present invention combines a system in which a user selects a theme, AI generates original stories and images based on that theme, and provides them in PDF format with an emotion engine that recognizes the user's emotions. This system can provide customized work instructions and work plans according to the emotional state of workers, especially in industrial environments.

[1166] Hardware and software used

[1167] Camera: A device that captures images for emotion recognition.

[1168] Microphone: A device that captures the user's voice emotions if necessary.

[1169] Speech recognition engine: A software component that analyzes the user's emotions.

[1170] EmotionRecognizer: A library that analyzes facial expression data and recognizes user emotions.

[1171] Generative AI Model: An AI algorithm for generating stories and images based on a theme (StoryGenerator, ImageGenerator).

[1172] PDFCreator: A library that converts generated stories and images into PDF format.

[1173] Natural language processing explanation

[1174] 1. User selects a theme:

[1175] A user uses a terminal to select a theme in an industrial environment (e.g., inspection work on line B). This theme indicates the work content and details of the workers.

[1176] 2. Emotion recognition:

[1177] The device's built-in camera and microphone are used to acquire the user's facial expression and voice data. The EmotionRecognizer library is used to recognize the user's emotions (e.g., joy, sadness, anger) from this data. This emotional information is sent to the server and used to generate stories and images.

[1178] 3. Story and image generation:

[1179] Based on the recognized emotional information and the selected theme, the server uses a Generative AI Model (Story Generator, Image Generator) to generate a customized story and associated images.

[1180] 4. Generate and serve PDF:

[1181] The generated stories and images are used to create PDF work instructions and work plans using the PDFCreator library, which are then saved on the server and a download link is provided to the user.

[1182] Specific examples

[1183] For example, if a user selects "Inspection work on Line B" as the theme and the emotion engine recognizes that the user is feeling a little down, the AI ​​can use this information to suggest "Perform equipment inspection on Line B. The theme will be adjusted to include simple and easy-to-understand steps and an encouraging message." This allows the user to receive customized instructions that reflect their emotional state that day, improving work efficiency.

[1184] Prompt Sentence Examples

[1185] "Today, the user is tasked with inspecting the machines on Line B. If the user is feeling down, generate a story that gives them simple instructions with an encouraging message."

[1186] "Emotion: Sad\nTask Details: Machine Inspection\nGenerate a positive and encouraging instruction set."

[1187] In this way, the system of the present invention can provide customized support according to the user's emotional state, contributing to improving work efficiency and maintaining worker motivation, particularly in industrial environments.

[1188] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1189] Step 1:

[1190] The user selects a theme on the terminal. From the theme list displayed on the terminal screen, the user selects a theme for a work instruction in an industrial environment (e.g., inspection work on line B). The input here is the theme selected by the user, and the output is detailed information about the selected theme. This information is passed to the subsequent processing step.

[1191] Step 2:

[1192] The system recognizes the user's emotions using the device's camera and microphone. The system uses the device's hardware (camera and microphone) to acquire the user's facial expression data and voice data. The system uses the EmotionRecognizer library to analyze the user's emotions from this data. The input is the acquired facial expression data and voice data, and the output is analyzed emotional information (e.g., joy, sadness, anger).

[1193] Step 3:

[1194] The recognized emotion information is sent to the server. The device then sends the analyzed emotion information to the server. At this time, the theme information selected by the user is also sent to the server. The emotion information and theme information are input, and the emotion information and theme information received by the server are output.

[1195] Step 4:

[1196] The server generates the story and images. Based on the theme and emotion information received on the server side, a generative AI model (Story Generator, Image Generator) is used to generate a customized story and images. The theme and emotion information are input here, and the generated story and images are output. Specifically, the AI ​​model creates a story while taking emotion information into account, and generates a prompt that generates an illustration based on the story.

[1197] Step 5:

[1198] The generated story and images are converted into PDF format. The server uses the generated story and images to convert them into a PDF document using the PDFCreator library. The input data is the generated story and images, and the output is a PDF document. Specifically, PDFCreator combines the story text and images to create a single PDF file.

[1199] Step 6:

[1200] The generated PDF is provided to the user. The server saves the generated PDF file as a temporary file and provides the link to the user. The input is the PDF file, and the output is a link to the PDF file that the user can access. The user can download or print the PDF from this link on their device.

[1201] Step 7:

[1202] The user downloads or prints a PDF. The device clicks on a link to download the PDF file and optionally prints it. The input to this step is the PDF link provided by the server, and the output is the downloaded PDF file or printed document. Specifically, the device opens the link, saves the file, and sends it to the printer to print the physical document.

[1203] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1204] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1205] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1206] [Fourth embodiment]

[1207] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1208] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1209] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1210] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1211] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1212] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1213] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1214] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1215] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1216] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1217] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1218] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1219] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1220] This invention provides a system in which a user selects a theme for a picture book, AI generates an original story and images based on that theme, and the picture book can be downloaded or printed in PDF format. The program and processing of this system are explained in detail below.

[1221] System Overview

[1222] When a user accesses the "AI Storyteller" website using a device, a theme selection screen appears. The user selects the desired theme and presses the generate button. The AI ​​then generates an original story and images based on the selected theme. The generated story and images are then converted into PDF format, which the user can download or print.

[1223] Program processing flow

[1224] 1. When a user accesses a website using a device, the device sends a connection request to the server, and the server returns HTML content including a theme selection screen. The device displays the returned HTML content on the screen, allowing the user to select a theme.

[1225] 2. When the user selects a theme and clicks the generate button, the device sends the selected theme data as a POST request to the server. The server receives the POST request and generates an original story based on the selected theme using an AI model.

[1226] 3. The server uses an AI model to generate illustrations based on the generated story, and combines the generated story and illustrations to create a picture book in PDF format.

[1227] 4. The created PDF is saved as a temporary file by the server, and a response containing a link that the user can download is sent back to the device.

[1228] 5. When the user clicks the download link on their device, the PDF is saved to their device. The user can view, download, and print the saved PDF.

[1229] Specific examples

[1230] For example, if a user selects the theme "Dragon," the server generates a story based on this theme, such as "A brave boy makes friends with a dragon." The AI ​​then creates illustrations of the dragon and boy that fit the generated story. The server then combines these illustrations with the story to create a picture book in PDF format and provides the user with a download link. The user can click the link to save the PDF on their device and read it to their child before bedtime.

[1231] This system eliminates the need for users to frequently purchase expensive picture books, allowing them to provide their children with a new story every night. Furthermore, by setting themes with educational messages, it is possible to effectively educate children. In this way, the present invention alleviates parental concerns and provides an efficient means for supporting children's curiosity and learning.

[1232] The processing flow will be explained below.

[1233] Step 1:

[1234] A user accesses the "AI Storyteller" website using a device. The device sends a connection request to the web server, and the web server returns HTML content including a theme selection screen.

[1235] Step 2:

[1236] The device parses the HTML received from the server and displays the theme selection screen, allowing the user to see the theme selection form.

[1237] Step 3:

[1238] The user selects the desired theme from a drop-down menu, for example, one of the options "adventure," "dragon," or "spaceship."

[1239] Step 4:

[1240] When the user presses the "Generate" button, the form data is sent from the terminal to the server as a POST request, including the selected theme.

[1241] Step 5:

[1242] The server receives the POST request and retrieves the selected theme data from the request, which is then used for the next process.

[1243] Step 6:

[1244] The server uses an AI model to generate an original story based on the selected theme. The server calls an AI library (e.g., OpenAI GPT-3) and executes the story generation API.

[1245] Step 7:

[1246] The server generates images appropriate for the generated story. The server again uses an AI model (e.g., DALL-E) to execute an illustration generation API based on the content of the story.

[1247] Step 8:

[1248] The server creates a PDF book based on the generated story and images. It uses a PDF generation library (e.g., FPDF) to insert the story text and images into the PDF document.

[1249] Step 9:

[1250] The server saves the completed PDF as a temporary file and retains its path, which is later provided to the user.

[1251] Step 10:

[1252] The server sends a response back to the device containing the path to the generated PDF, along with a link to download the PDF.

[1253] Step 11:

[1254] The device parses the response received from the server and displays a PDF download link to the user, which the user can click to download the PDF.

[1255] Step 12:

[1256] When a user clicks the download link, the device saves the PDF file to local storage. By opening the saved PDF, the user can view, download, and print the generated original picture book.

[1257] Through this series of processes, users can easily obtain high-quality original picture books and provide their children with new stories every night.

[1258] Example 1

[1259] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1260] In modern child-rearing, parents want to provide their children with a variety of fresh picture books that meet their children's interests and educational needs. However, frequently purchasing new picture books is a significant financial burden, and commercially available picture books do not always contain the educational messages parents desire. Furthermore, busy parents find it difficult to quickly obtain personalized picture books tailored to their children's preferences. Efficient methods to solve these problems are needed.

[1261] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1262] In this invention, the server includes means for a user to select a theme, means for generating an original story using a generative AI model based on the selected theme, means for generating images using the generative AI model based on the content of the generated story, means for converting the generated story and images into a PDF format, and means for the user to download or print the PDF picture book. This allows a user to easily generate an original picture book based on a theme and download or print it.

[1263] A "theme" is a subject or topic selected by the user that will be the basis for the story or image that will be generated.

[1264] A "generative AI model" is an artificial intelligence algorithm or system that generates text or images based on a selected theme.

[1265] An "original story" is a unique narrative generated by a generative AI model based on a selected theme.

[1266] "Images" are visual illustrations or pictures generated by a generative AI model to match the content of the original story.

[1267] "PDF format" is an abbreviation for Portable Document Format, and is a file format for saving and displaying documents and images on a page-by-page basis.

[1268] A picture book is a book for children that contains stories and illustrations in written form.

[1269] "Downloading" is the act of saving a file from the Internet to your device.

[1270] "Printing" is the act of printing digital data onto physical media such as paper.

[1271] A "user" is someone who uses the system to select a theme, create, download, and print a picture book.

[1272] A "server" is a computer system that receives requests from users and generates, stores, and delivers appropriate content.

[1273] A "terminal" is a device operated by a user, and is a computer that communicates with a server and displays and saves results.

[1274] The present invention is a system that uses a generative AI model to generate original stories and images based on a theme selected by a user, and provides them as a PDF picture book. The system includes a means for a user to select a theme, a means for converting the generated stories and images into PDF format, and a means for the user to download or print the PDF picture book.

[1275] Hardware and Software

[1276] The system is implemented using a server, a terminal, a generative AI model, and various software. The server receives requests, generates stories and images, converts them into PDF format, and distributes them. The terminal acts as an interface for users to access the system, select a theme, and download the PDF.

[1277] The main hardware and software used are as follows:

[1278] Server: Processes the data and invokes the generative AI model.

[1279] Terminal: The device through which the user accesses the system, such as a regular PC, smartphone, or tablet.

[1280] Generative AI models: For example, ChatGPT and GPT-4 are used for story generation, while DALL-E and Stable Diffusion are used for image generation.

[1281] PDF generation library: Generate PDFs using Python's ReportLab, etc.

[1282] Data processing and calculation

[1283] 1. Theme selection: The user selects a theme using the terminal. For example, the user can select the theme "Dragon."

[1284] 2. Story generation: The server uses a generative AI model (such as ChatGPT or GPT-4) to generate an original story based on the selected theme. An example of a specific prompt is "Theme: Dragons. Please generate a story about a brave boy who becomes friends with a dragon."

[1285] 3. Image Generation: The server uses a generative AI model (such as DALL-E or Stable Diffusion) to generate images that match the generated story. The prompt might be something like, "Generate an illustration of a dragon and a boy."

[1286] 4. PDF generation: The server combines the generated story and images and generates a PDF picture book using Python's ReportLab or similar.

[1287] 5. PDF Delivery: The server saves the PDF in temporary storage and provides a download link that users can access, and they can download or print the PDF using their devices.

[1288] Specific examples

[1289] For example, if a user selects "Dragons" as the theme, the process goes as follows: The user selects a theme through the website and presses the generate button. The device then sends the theme data to the server. The server uses a generative AI model to generate a story such as "A brave boy befriends a dragon" and creates illustrations of the dragon and boy based on that story. The server then assembles this content into a PDF picture book using Python's ReportLab and provides the user with a download link. The user can click this link to download the PDF and read it to their child.

[1290] This system eliminates the need for users to frequently purchase expensive picture books, allowing them to provide their children with a new story every night. Furthermore, by setting themes with educational messages, it is possible to effectively educate children. In this way, the present invention alleviates parental concerns and provides an efficient means for supporting children's curiosity and learning.

[1291] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1292] Specific flow of system program processing

[1293] Step 1:

[1294] The user accesses the "AI Storyteller" website using a device.

[1295] Input: An HTTP request sent by a user typing a URL into a web browser.

[1296] Output: The device sends a connection request to the server, and the server returns a theme selection screen.

[1297] Specific behavior:

[1298] The user opens a browser, enters the specified URL, and presses enter.

[1299] The device sends this operation to the server as an HTTP request.

[1300] The server receives the request and generates HTML content including a theme selection screen.

[1301] The server returns the generated HTML content to the terminal as an HTTP response.

[1302] The terminal displays the received HTML content on the screen and allows the user to select a theme.

[1303] Step 2:

[1304] The user selects a theme and clicks the generate button.

[1305] Input: The user selects a theme and clicks the generate button.

[1306] Output: The device sends the selected theme data to the server as a POST request.

[1307] Specific behavior:

[1308] The user selects a desired theme from the displayed theme selection screen.

[1309] The user clicks the "Generate" button.

[1310] The device sends the user's selected theme data to the server as a POST request in JSON format.

[1311] Step 3:

[1312] The server receives the POST request and generates the story.

[1313] Input: A POST request containing theme data sent from the device.

[1314] Output: The original story generated using the generative AI model.

[1315] Specific behavior:

[1316] The server receives the POST request and extracts the theme data from the request body.

[1317] The server sends a prompt to the generative AI model (e.g., ChatGPT or GPT-4). Example prompt: "Theme: Dragons. Generate a story about a brave boy who becomes friends with a dragon."

[1318] The generative AI model generates a story based on the prompt sentence and returns the story to the server.

[1319] Step 4:

[1320] The server generates illustrations based on the story.

[1321] Input: Generated stories.

[1322] Output: An illustration generated using a generative AI model.

[1323] Specific behavior:

[1324] The server analyzes the content of the generated story and generates a prompt accordingly. Example prompt: "Please generate an illustration of a dragon and a boy."

[1325] The server sends a prompt to a generative AI model (e.g., DALL-E or Stable Diffusion).

[1326] The generative AI model generates an illustration based on the prompt sentence and returns the illustration to the server.

[1327] Step 5:

[1328] The server creates a picture book in PDF format.

[1329] Input: Generated story and illustrations.

[1330] Output: Picture book in PDF format.

[1331] Specific behavior:

[1332] The server combines the generated story and illustrations and creates a PDF picture book using Python's ReportLab etc.

[1333] The server saves the generated PDF file in temporary storage.

[1334] Step 6:

[1335] The server provides a download link for the PDF.

[1336] Input: PDF file stored on the server.

[1337] Output: A response containing a download link for the PDF.

[1338] Specific behavior:

[1339] The server generates a download link for the PDF file stored in temporary storage.

[1340] The server sends a response including the download link back to the device.

[1341] Step 7:

[1342] The user downloads or prints the PDF.

[1343] Input: Click on the download link displayed on your device.

[1344] Output: Downloaded PDF file.

[1345] Specific behavior:

[1346] The user clicks on the download link displayed in the device's browser.

[1347] The device downloads the PDF file from the server.

[1348] The user can view the PDF file stored on the device and print it if necessary.

[1349] Through the above specific processing steps, the system generates an original picture book based on the theme selected by the user, and makes it easy to download and print.

[1350] (Application example 1)

[1351] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1352] Currently, many families and educational institutions need to regularly purchase picture books for children. However, existing picture books have fixed content, which can make it difficult to keep children interested. Furthermore, because picture books are physical, collecting a large number of them requires storage space and is costly. Furthermore, parents and educators want to provide educational content that responds to children's development and interests. For these reasons, there is a need for a cost-effective system that can flexibly generate content.

[1353] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1354] In this invention, the server includes means for a user to select a theme, means for generating an original story based on the selected theme, means for generating visual materials based on the generated story, means for converting the generated story and visual materials into an electronic document format, means for the user to download or print the picture book in the electronic document format, and means for the user to select a theme and preview and check the generated picture book in real time using a touch panel terminal, thereby enabling users to generate and print personalized picture books tailored to their wishes on the spot in the store.

[1355] "User" refers to the person who uses this system to select a picture book theme, generate content, download it, and print it.

[1356] "Theme" refers to the subject that forms the basis of the content or subject matter of the picture book selected by the user.

[1357] "Story" refers to an original story generated by AI based on a selected theme.

[1358] "Visual materials" refer to illustrations and images created by AI based on the generated story.

[1359] "Electronic document format" refers to a digital document, such as a PDF format, that integrates and converts the generated narrative and visual materials.

[1360] "Downloading" refers to the process of saving the generated picture book in electronic document format to the user's terminal via the Internet.

[1361] "Printing" refers to the process of outputting the generated picture book in electronic document format onto physical paper media.

[1362] A "touch panel terminal" refers to a computer device equipped with a display that the user can operate by directly touching it.

[1363] "Real-time" refers to the timing at which data is generated and processed immediately.

[1364] "Preview" refers to a temporary display or sample that is displayed to confirm the final generated result.

[1365] "System" refers to the entire platform that provides a set of functions that allow users to generate, download, and print picture books.

[1366] This invention is a system in which a user selects a theme, an AI generates an original story and visual materials based on that theme, and provides a picture book in electronic document format. Specific embodiments for implementing this system will be described below.

[1367] System Configuration

[1368] The present invention is configured by a series of hardware and software components including a user terminal, a server, and a touch panel terminal.

[1369] 1. User Device:

[1370] A device that allows users to select a theme and download or print the generated picture book. Examples include computers, smartphones, and tablets.

[1371] 2. Server:

[1372] It is a central device that receives user requests and generates stories and visual materials using AI models. The server is installed with a language model (e.g., GPT-4) for story generation and an image model (e.g., DALL-E) for visual material generation.

[1373] 3. Touchscreen terminal:

[1374] This device is installed in educational institutions and brick-and-mortar stores, allowing users to select a theme and preview the generated picture book in real time.

[1375] Processing Overview

[1376] When a user accesses the system using a terminal, the touch panel terminal displays a theme selection screen. The user selects the desired theme and presses the generate button on the terminal, and the selected theme data is sent to the server. The server uses a generative AI model to generate an original story based on the selected theme. It then generates illustrations and images based on the generated story. The generated story and visual materials are converted into a digital document in PDF format, which the user can download or print.

[1377] Hardware and Software Use

[1378] Hardware: Touchscreen devices (e.g., Surface Pro), user devices (e.g., iPhone, Android tablet), servers (e.g., AWS EC2 instances)

[1379] Software: The server-side program is implemented in Python and uses web frameworks such as Flask, and also uses GPT-4 for story generation and DALL-E for visual material generation.

[1380] Data processing and calculation

[1381] The server inputs the topic information sent by the user into a natural language processing model and generates a story. Specifically, it uses the following prompt sentences:

[1382] Example prompt: "Generate a story based on the following theme: Dragons."

[1383] Based on the generated story, a request for illustration generation is sent to the visual material generation model. The generated visual material and story are integrated and converted into a PDF file. The PDF is temporarily stored on the server, and a download link is provided to the user.

[1384] Specific examples

[1385] For example, if a user selects the "Dragon" theme using a touchscreen device in a physical store, the server proceeds as follows:

[1386] 1. The user selects the "Dragons" theme.

[1387] 2. The server sends the prompt "Generate a story based on the following theme: Dragons" to the generative AI model.

[1388] 3. The generative AI model generates the story and sends it back to the server.

[1389] 4. The server inputs the story into a visual material generation model and generates illustrations related to dragons.

[1390] 5. Integrate the generated narrative and visual materials and convert them into a PDF file.

[1391] 6. The user downloads the generated PDF or prints it on an in-store printer.

[1392] This system allows users to flexibly create and use personalized picture books on the spot.

[1393] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1394] Step 1:

[1395] The user uses the device to access the theme selection screen on the touch panel device. The server returns HTML content including the theme selection screen, which the device displays. This stage mainly involves collecting user input. Input: None, Output: Display of theme selection screen.

[1396] Step 2:

[1397] The user selects a theme and clicks the generate button on the touch panel terminal. The terminal sends the selected theme data to the server as a POST request. At this stage, the theme selected by the user is sent to the server. Input: Theme selected by the user, Output: Theme data sent to the server.

[1398] Step 3:

[1399] The server processes the received POST request and sends a prompt to the generative AI model based on the selected theme. For example, it sends a prompt such as "Generate a story based on the following theme: Dragons." Input: Theme data, Output: Prompt to the generative AI model.

[1400] Step 4:

[1401] The generative AI model generates an original story based on the server's request and sends it back to the server. The server receives the generated story. Input: prompt sentence, output: generated story.

[1402] Step 5:

[1403] Based on the generated story, the server sends a request to the visual material generation model to generate visual materials (e.g., illustrations and images) appropriate for the story. Input: Generated story, Output: Request to the visual material generation model.

[1404] Step 6:

[1405] The visual material generation model generates visual material based on the server's request and sends it back to the server. The server receives the generated visual material. Input: request, Output: generated visual material.

[1406] Step 7:

[1407] The server combines the generated story and visual materials and converts them into an electronic document in PDF format. Input: Generated story and visual materials, Output: Electronic document in PDF format.

[1408] Step 8:

[1409] The server saves the generated PDF as a temporary file and returns a response containing a download link to the device. Input: PDF file, Output: Response containing a download link.

[1410] Step 9:

[1411] The user clicks the download link on their device to download or print the PDF. At this stage, the final product is obtained by the user. Input: Download link, Output: Save or print the PDF file.

[1412] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1413] This invention combines a system in which a user selects a theme for a picture book, AI generates an original story and images based on that theme, and the picture book can be downloaded or printed in PDF format, with an emotion engine that recognizes the user's emotions. The program and processing of this system are explained in detail below.

[1414] System Overview

[1415] When a user accesses the "AI Storyteller" website using a device, a theme selection screen is displayed. The user selects the desired theme and presses the generate button. The AI ​​then generates an original story and images based on the selected theme. The generated story and images are then converted into PDF format, which the user can download or print. The present invention also incorporates an emotion engine that recognizes the user's emotions, allowing it to recommend themes and adjust the story based on the user's emotions.

[1416] Program processing flow

[1417] 1. When a user accesses a website using a device, the device sends a connection request to the server, and the server returns HTML content including a theme selection screen. The device displays the returned HTML content on the screen, allowing the user to select a theme.

[1418] 2. The emotion engine installed in the device analyzes the user's facial expressions and voice to recognize their emotions. For example, it uses the device's camera and microphone to identify whether the user is happy, depressed, surprised, etc.

[1419] 3. The emotion engine sends the recognized emotion information to the server, which receives the emotion information and uses it to recommend themes and generate stories.

[1420] 4. When the user selects a theme and clicks the generate button, the device sends a POST request containing the selected theme data and emotion information to the server. The server receives the POST request and uses an AI model based on the selected theme and emotion information to generate an original story.

[1421] 5. The server generates images appropriate for the generated story. The server again uses the AI ​​model to execute an illustration generation API based on the content of the story. Emotional information is also taken into account here, and an illustration appropriate for the emotion is generated.

[1422] 6. The server creates a PDF book based on the generated story and images. It uses a PDF generation library to insert the story text and images into a PDF document.

[1423] 7. The created PDF is saved as a temporary file by the server and its path is saved. This temporary file is later provided to the user.

[1424] 8. The server sends a response back to the device containing the path to the generated PDF and a link to download the PDF.

[1425] 9. The device parses the response received from the server and displays a link to download the PDF to the user. When the user clicks the download link, the PDF is saved to the device.

[1426] Specific examples

[1427] For example, if a user selects the theme "Dragon" and the emotion engine recognizes that the user is feeling a little depressed, the server will generate a story based on this theme and emotion information, such as "A brave boy befriends a dragon and goes on a fun adventure together." The AI ​​then creates cheerful, hopeful illustrations of the dragon and boy that fit the generated story. The server then combines these illustrations with the story to create a picture book in PDF format and provides the user with a download link. The user can click the link to save the PDF to their device and read it to their child before bedtime.

[1428] This system allows users to easily obtain high-quality original picture books and provide their children with new stories every night. Furthermore, the introduction of an emotion engine allows the system to provide optimal themes and stories according to the user's emotional state, creating a more personalized educational experience.

[1429] The processing flow will be explained below.

[1430] Step 1:

[1431] A user accesses the "AI Storyteller" website using a device. The device sends a connection request to the web server, and the server returns HTML content including a theme selection screen.

[1432] Step 2:

[1433] The device parses the HTML received from the server and displays the theme selection screen, allowing the user to see the theme selection form.

[1434] Step 3:

[1435] The emotion engine installed in the device analyzes the user's facial expressions and voice to recognize their emotions. For example, it uses the device's camera and microphone to identify whether the user is happy, depressed, surprised, etc.

[1436] Step 4:

[1437] The emotion engine sends the emotional information it recognizes to the server, which then recommends themes appropriate for the user and adjusts story generation based on the emotional information received.

[1438] Step 5:

[1439] The user selects the desired theme from a drop-down menu, for example, one of the options "adventure," "dragon," or "spaceship."

[1440] Step 6:

[1441] When the user presses the "Generate" button, the form data is sent from the device to the server as a POST request, which includes the selected theme and emotion information.

[1442] Step 7:

[1443] The server receives the POST request and retrieves the selected theme and emotion information from the request, which is then used for the next process.

[1444] Step 8:

[1445] The server uses an AI model to generate an original story based on the selected theme and emotional information. The server calls an AI library (e.g., OpenAI GPT-3) and executes the story generation API.

[1446] Step 9:

[1447] The server generates images appropriate for the generated story. The server again uses an AI model (such as DALL-E) to execute an illustration generation API based on the story content and emotional information.

[1448] Step 10:

[1449] The server creates a PDF book based on the generated story and images. It uses a PDF generation library (such as FPDF) to insert the story text and images into the PDF document.

[1450] Step 11:

[1451] The server saves the completed PDF as a temporary file and retains its path, which is later provided to the user.

[1452] Step 12:

[1453] The server sends a response back to the device containing the path to the generated PDF, along with a link to download the PDF.

[1454] Step 13:

[1455] The device parses the response received from the server and displays a link to download the PDF to the user. When the user clicks the link, the PDF is saved on the device.

[1456] Step 14:

[1457] By opening the saved PDF, users can view, download, and print the original picture book they created. The emotional engine allows the system to provide the most appropriate theme and story based on the user's emotional state, enabling a personalized educational experience.

[1458] Through this series of steps, users can easily obtain high-quality original picture books and provide their children with new stories every night. The emotional engine generates themes and stories that best suit the user's emotional state, providing a more personalized educational experience.

[1459] Example 2

[1460] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1461] In conventional picture book generation systems, users simply select a theme, and the generated story and images are not optimized for the user's emotional state. As a result, picture books with content that does not match the user's emotions or mood are generated, resulting in low satisfaction. Furthermore, because emotion recognition technology is not incorporated, personalized educational experiences are not provided.

[1462] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1463] In this invention, the server includes: means for a user to select a theme; means for generating an original story based on the selected theme; means for generating images based on the generated story; means for converting the generated story and images into PDF format; means for the user to download or print the PDF picture book; and means including an emotion engine that recognizes the user's emotions and that recommends themes and adjusts the story based on the recognized emotions. This automatically generates an original picture book suited to the user's emotional state, enabling the user to have a highly satisfying and personalized educational experience.

[1464] A "user" is an individual who uses the system to create a picture book.

[1465] A "theme" is a topic or category that determines the content of the picture book selected by the user.

[1466] An "original story" is a unique narrative created by a generative AI model based on a selected theme.

[1467] "Images" are visual content that is automatically created based on the generated story.

[1468] "PDF format" is a format that integrates the generated story and images into a single document file and saves it in Portable Document Format (PDF).

[1469] An "emotion engine" is a software and hardware system that analyzes a user's facial expressions, voice, etc., and recognizes the user's emotional state.

[1470] A "server" is a computer system that receives user requests and generates stories and images and creates PDF files in response.

[1471] "Means of selection" refers to the interface (e.g., a screen or menu on a website) through which a user selects a theme.

[1472] "Generative means" refers to the AI ​​models or algorithms used to automatically create original stories and images based on selected themes and emotional information.

[1473] "Means for recommending themes and adjusting stories based on emotions" refers to a function that suggests appropriate themes and adjusts the content of stories based on the user's recognized emotional information.

[1474] This system allows users to select a theme, generates an original story and images based on that theme using a generative AI model, and provides a picture book in PDF format. Furthermore, it incorporates an emotion engine that recognizes the user's emotions, and recommends themes and adjusts the story based on this emotion information.

[1475] System configuration

[1476] This system consists of the following elements:

[1477] 1. User Device:

[1478] Browser: Software used to access websites.

[1479] Camera: Hardware that captures the user's facial expressions.

[1480] Microphone: Hardware that captures the user's voice.

[1481] Emotion engine: Software that recognizes user emotions by analyzing facial expressions and voice (e.g., OpenCV, Deep Learning models).

[1482] 2. Server:

[1483] Web server: A computer system that receives requests from users and returns appropriate responses.

[1484] Generative AI models: Language models that generate stories based on selected themes and sentiment information (e.g., GPT-3).

[1485] Illustration generation API: An API for generating images based on a generated story (e.g., DALL-E).

[1486] PDF Generation Library: A library that converts the generated stories and images into PDF format (e.g. FPDF, PDFBox).

[1487] System Operation

[1488] When a user accesses the "AI Storyteller" website using their device, the device sends a connection request to the server, which then returns HTML content including a theme selection screen. The device's built-in emotion engine then activates and analyzes the user's facial expressions and voice to recognize their emotions. For example, the device's camera and microphone can be used to determine whether the user is happy or depressed. The emotion information is then sent to the server, which uses it to recommend themes and generate stories.

[1489] When a user selects a theme and clicks the generate button, the device sends a POST request containing the selected theme data and emotional information to the server. The server receives the POST request and uses a generative AI model based on the selected theme and emotional information to generate an original story. Images appropriate for the generated story are also generated on the server, and further appropriate illustrations are selected based on the emotional information.

[1490] Once the story and images are ready, the server uses a PDF generation library to convert them into a PDF picture book. The server saves the created PDF as a temporary file and sends the path to the file back to the user. The user can click the link to download the PDF and save it on their device.

[1491] Specific examples

[1492] For example, if a user selects the theme "Dragon" and the emotion engine recognizes that the user is feeling a little depressed, the server will generate a story based on this theme and emotion information, such as "A brave boy befriends a dragon and goes on a fun adventure together." The AI ​​then creates bright and hopeful illustrations of the dragon and boy that fit the generated story. The server then combines these illustrations with the story to create a picture book in PDF format and provides the user with a download link. The user can click the link to save the PDF to their device and read it to their child before bedtime.

[1493] Example prompt sentence:

[1494] "Adventures with Dragons"

[1495] "Stories that make children brave"

[1496] "Make friends"

[1497] "Hopeful illustrations"

[1498] This system allows users to easily obtain high-quality original picture books and provide their children with new stories every night. Furthermore, the introduction of an emotion engine allows the system to provide the most appropriate themes and stories according to the user's emotional state, creating a more personalized educational experience.

[1499] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1500] Step 1:

[1501] When a user accesses a website using a device, the device sends an HTTP request to the server. The server receives the request, generates HTML content including a theme selection screen, and sends it back to the device as an HTTP response.

[1502] Input: URL access request from user

[1503] Data processing / calculation: The server analyzes the request and generates appropriate HTML content

[1504] Output: Response showing theme selection screen

[1505] Step 2:

[1506] The emotion engine installed on the device is activated and uses the camera and microphone to capture the user's facial expressions and voice in real time. The emotion engine then uses facial expression and voice analysis software to recognize the user's emotions and generate data.

[1507] Input: User's facial expressions and voice

[1508] Data processing / computation: Capture data using camera and microphone and perform sentiment analysis

[1509] Output: User emotion data

[1510] Step 3:

[1511] The device sends the recognized emotion information to the server, which receives the emotion data and stores it for later processing.

[1512] Input: Emotion data

[1513] Data processing / calculation: Data reception and storage

[1514] Output: Saved emotion data

[1515] Step 4:

[1516] The user selects the desired theme on the theme selection screen and clicks the Generate button. The device then sends a POST request to the server containing the selected theme and the previously recognized emotion data.

[1517] Input: User-selected theme, stored emotion data

[1518] Data processing / calculation: Combining theme selection data and emotion data, generating POST requests

[1519] Output: POST request sent

[1520] Step 5:

[1521] The server receives the POST request and generates an original story using a generative AI model based on the selected theme and emotion data. For example, a prompt sentence is input to a generative AI model (e.g., GPT-3) to generate the story as text.

[1522] Input: Theme data, emotion data

[1523] Data processing / calculation: Use of generative AI models, generation of story text

[1524] Output: Generated story text

[1525] Step 6:

[1526] The server uses an illustration generation API (e.g., DALL-E) based on the generated story to generate images that suit the story, taking into account emotional data to create visuals that are optimal for the user.

[1527] Input: Story text, emotion data

[1528] Data processing / calculation: Calling illustration generation API, image generation

[1529] Output: The generated image

[1530] Step 7:

[1531] The server converts the generated story and images into a PDF book using a PDF generation library (e.g., FPDF, PDFBox). The story and images are inserted into a template and a PDF file is generated.

[1532] Input: Story text, generated images

[1533] Data processing / calculation: Use of PDF generation library, creation of PDF files

[1534] Output: PDF picture book

[1535] Step 8:

[1536] The server saves the created PDF as a temporary file, retains its path, and returns a response including the path of the saved file to the terminal.

[1537] Input: Generated PDF

[1538] Data processing / calculation: file saving, path generation

[1539] Output: Response to device, download link

[1540] Step 9:

[1541] The device parses the response received from the server and displays a link to download the PDF to the user. When the user clicks the link, the PDF is saved on the device.

[1542] Input: Response from the server

[1543] Data processing / calculation: Response analysis, link display

[1544] Output: Download and save PDF

[1545] (Application example 2)

[1546] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1547] Conventional illustrated work instruction generation systems create formulaic instructions without considering the user's emotional state, which does not adequately improve worker motivation or provide emotional care. Furthermore, instructions that do not take emotions into consideration may affect the efficiency and effectiveness of work. Therefore, there is a need to provide customized work instructions and work plans that correspond to the emotional state of workers.

[1548] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for a user to select a theme, means for generating an original story based on the selected theme, means for generating images based on the generated story, means for converting the generated story and images into PDF format, means for the user to download or print the PDF format document, means for recognizing a user's emotion, means for adjusting the story and images based on the recognized emotion, means for recognizing a user's emotion in an industrial environment and adjusting work instructions and work plans, and means for converting the generated work instructions and work plans into PDF format. This makes it possible to provide work instructions and work plans customized according to the emotional state of the worker.

[1549] A "theme" is the content or subject matter that a user selects to create an original story.

[1550] A "story" is a narrative or story content that is generated based on a selected theme.

[1551] "Images" refers to visual content created based on the generated story, including visual information such as pictures and illustrations.

[1552] "PDF format" is an abbreviation for Portable Document Format, and refers to a fixed-layout file format suitable for storing and distributing electronic documents.

[1553] A "document" is a written document that contains textual information such as a story or instructions.

[1554] "Emotions" represent the user's psychological state or sensations, and include types such as joy, sadness, surprise, and anger.

[1555] An "emotion engine" is a software or hardware mechanism that uses sensors such as cameras and microphones to recognize and analyze a user's emotional state.

[1556] "Industrial environment" refers to the working environment at a production or manufacturing site or facility.

[1557] A "work instruction" is a document that describes the work content and procedures that workers should perform in an industrial environment.

[1558] A "work plan" is a document that includes detailed processes and schedules for efficiently carrying out a specific task.

[1559] A "server" is a computer system that provides services over a network, and refers to the equipment and software that processes and distributes data in response to requests.

[1560] A "generative AI model" is an algorithm or model that uses artificial intelligence technology to generate new content (stories or images) based on input data.

[1561] The present invention combines a system in which a user selects a theme, AI generates original stories and images based on that theme, and provides them in PDF format with an emotion engine that recognizes the user's emotions. This system can provide customized work instructions and work plans according to the emotional state of workers, especially in industrial environments.

[1562] Hardware and software used

[1563] Camera: A device that captures images for emotion recognition.

[1564] Microphone: A device that captures the user's voice emotions if necessary.

[1565] Speech recognition engine: A software component that analyzes the user's emotions.

[1566] EmotionRecognizer: A library that analyzes facial expression data and recognizes user emotions.

[1567] Generative AI Model: An AI algorithm for generating stories and images based on a theme (StoryGenerator, ImageGenerator).

[1568] PDFCreator: A library that converts generated stories and images into PDF format.

[1569] Natural language processing explanation

[1570] 1. User selects a theme:

[1571] A user uses a terminal to select a theme in an industrial environment (e.g., inspection work on line B). This theme indicates the work content and details of the workers.

[1572] 2. Emotion recognition:

[1573] The device's built-in camera and microphone are used to acquire the user's facial expression and voice data. The EmotionRecognizer library is used to recognize the user's emotions (e.g., joy, sadness, anger) from this data. This emotional information is sent to the server and used to generate stories and images.

[1574] 3. Story and image generation:

[1575] Based on the recognized emotional information and the selected theme, the server uses a Generative AI Model (Story Generator, Image Generator) to generate a customized story and associated images.

[1576] 4. Generate and serve PDF:

[1577] The generated stories and images are used to create PDF work instructions and work plans using the PDFCreator library, which are then saved on the server and a download link is provided to the user.

[1578] Specific examples

[1579] For example, if a user selects "Inspection work on Line B" as the theme and the emotion engine recognizes that the user is feeling a little down, the AI ​​can use this information to suggest "Perform equipment inspection on Line B. The theme will be adjusted to include simple and easy-to-understand steps and an encouraging message." This allows the user to receive customized instructions that reflect their emotional state that day, improving work efficiency.

[1580] Prompt Sentence Examples

[1581] "Today, the user is tasked with inspecting the machines on Line B. If the user is feeling down, generate a story that gives them simple instructions with an encouraging message."

[1582] "Emotion: Sad\nTask Details: Machine Inspection\nGenerate a positive and encouraging instruction set."

[1583] In this way, the system of the present invention can provide customized support according to the user's emotional state, contributing to improving work efficiency and maintaining worker motivation, particularly in industrial environments.

[1584] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1585] Step 1:

[1586] The user selects a theme on the terminal. From the theme list displayed on the terminal screen, the user selects a theme for a work instruction in an industrial environment (e.g., inspection work on line B). The input here is the theme selected by the user, and the output is detailed information about the selected theme. This information is passed to the subsequent processing step.

[1587] Step 2:

[1588] The system recognizes the user's emotions using the device's camera and microphone. The system uses the device's hardware (camera and microphone) to acquire the user's facial expression data and voice data. The system uses the EmotionRecognizer library to analyze the user's emotions from this data. The input is the acquired facial expression data and voice data, and the output is analyzed emotional information (e.g., joy, sadness, anger).

[1589] Step 3:

[1590] The recognized emotion information is sent to the server. The device then sends the analyzed emotion information to the server. At this time, the theme information selected by the user is also sent to the server. The emotion information and theme information are input, and the emotion information and theme information received by the server are output.

[1591] Step 4:

[1592] The server generates the story and images. Based on the theme and emotion information received on the server side, a generative AI model (Story Generator, Image Generator) is used to generate a customized story and images. The theme and emotion information are input here, and the generated story and images are output. Specifically, the AI ​​model creates a story while taking emotion information into account, and generates a prompt that generates an illustration based on the story.

[1593] Step 5:

[1594] The generated story and images are converted into PDF format. The server uses the generated story and images to convert them into a PDF document using the PDFCreator library. The input data is the generated story and images, and the output is a PDF document. Specifically, PDFCreator combines the story text and images to create a single PDF file.

[1595] Step 6:

[1596] The generated PDF is provided to the user. The server saves the generated PDF file as a temporary file and provides the link to the user. The input is the PDF file, and the output is a link to the PDF file that the user can access. The user can download or print the PDF from this link on their device.

[1597] Step 7:

[1598] The user downloads or prints a PDF. The device clicks on a link to download the PDF file and optionally prints it. The input to this step is the PDF link provided by the server, and the output is the downloaded PDF file or printed document. Specifically, the device opens the link, saves the file, and sends it to the printer to print the physical document.

[1599] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1600] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1601] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1602] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1603] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1604] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1605] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1606] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1607] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1608] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1609] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1610] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1611] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1612] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1613] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1614] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1615] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1616] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1617] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1618] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, in order to avoid confusion and to facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1619] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1620] The following is further disclosed regarding the above embodiment.

[1621] (Claim 1)

[1622] a means for a user to select a theme;

[1623] a means for generating an original story based on a selected theme;

[1624] means for generating images based on the generated story;

[1625] A means to convert the generated stories and images into PDF format;

[1626] A means for users to download or print the book in PDF format;

[1627] A system including:

[1628] (Claim 2)

[1629] 10. The system of claim 1, wherein the theme includes an educational message.

[1630] (Claim 3)

[1631] 10. The system of claim 1, wherein the generated PDF picture book includes a multi-page format.

[1632] "Example 1"

[1633] (Claim 1)

[1634] a means for a user to select a theme;

[1635] a means for generating original stories using a generative AI model based on the selected theme; and

[1636] a means for generating images using a generative AI model based on the content of the generated story;

[1637] A means to convert the generated stories and images into PDF format;

[1638] A means for users to download or print the book in PDF format;

[1639] A system including:

[1640] (Claim 2)

[1641] 10. The system of claim 1, wherein the theme includes an educational message.

[1642] (Claim 3)

[1643] 10. The system of claim 1, wherein the generated PDF picture book includes a multi-page format.

[1644] "Application Example 1"

[1645] (Claim 1)

[1646] a means for a user to select a theme;

[1647] a means for generating an original story based on a selected theme;

[1648] a means for generating visual material based on the generated narrative;

[1649] a means of converting the generated narrative and visual material into an electronic document format;

[1650] means for a user to download or print the picture book in electronic document format;

[1651] The system includes a means for users to select a theme using a touch panel terminal and preview and check the generated picture book in real time.

[1652] (Claim 2)

[1653] 10. The system of claim 1, wherein the theme includes an educational message.

[1654] (Claim 3)

[1655] 10. The system of claim 1, wherein the generated picture book in electronic document format includes a multi-page format.

[1656] "Example 2: Combining Emotion Engines"

[1657] (Claim 1)

[1658] a means for a user to select a theme;

[1659] a means for generating an original story based on a selected theme;

[1660] means for generating images based on the generated story;

[1661] A means to convert the generated stories and images into PDF format;

[1662] A means for users to download or print the book in PDF format;

[1663] a means for recommending themes and adjusting stories based on the recognized emotions, the means including an emotion engine for recognizing the emotions of the user;

[1664] A system including:

[1665] (Claim 2)

[1666] 10. The system of claim 1, wherein the theme includes an educational message.

[1667] (Claim 3)

[1668] 10. The system of claim 1, wherein the generated PDF picture book includes a multi-page format.

[1669] "Application example 2 when combining emotion engines"

[1670] (Claim 1)

[1671] a means for a user to select a theme;

[1672] a means for generating an original story based on a selected theme;

[1673] means for generating images based on the generated story;

[1674] A means to convert the generated stories and images into PDF format;

[1675] A means for users to download or print documents in PDF format;

[1676] means for recognizing a user's emotion;

[1677] A means to tailor stories and images based on perceived emotions;

[1678] A system including:

[1679] (Claim 2)

[1680] 10. The system of claim 1, wherein the theme includes an educational message.

[1681] (Claim 3)

[1682] 10. The system of claim 1, wherein the generated PDF format document comprises a multi-page format.

[1683] (Claim 4)

[1684] 10. The system of claim 1, further comprising: means for recognizing a user's emotion in an industrial environment and adjusting a work instruction or work plan; and means for converting the generated work instruction or work plan into a PDF format. [Explanation of symbols]

[1685] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for a user to select a theme; a means for generating an original story based on a selected theme; means for generating images based on the generated story; A means to convert the generated stories and images into PDF format; A means for users to download or print the book in PDF format; A system including:

2. 10. The system of claim 1, wherein the theme includes an educational message.

3. The system of claim 1 , wherein the generated PDF picture book includes a multi-page format.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A