System
A system using generative AI generates personalized video picture books with educational themes, addressing the challenge of parental time constraints and content availability, facilitating effective educational storytelling for young children.
Patent Information
- Application Number
- JP2024117270
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-22
- Publication Date
- 2026-02-03
AI Technical Summary
Parents of young children aged 0 to 6 face challenges in dedicating time to read picture books daily and finding content that conveys specific educational messages, making it difficult to support their children's educational growth effectively.
A system that allows users to select a theme for a video picture book, using a generative AI model to generate original stories and illustrations, integrate a child's preferred voice, and include educational messages, enabling easy creation and display on digital devices.
Reduces parental burden by allowing easy generation of original video picture books with educational content, supporting children's learning and growth through personalized storytelling.
Smart Images

Figure 2026016180000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In today's busy lifestyles, it is difficult for parents of young children aged 0 to 6 to devote much time and effort to reading picture books to their children. The task of preparing a different picture book each day and reading it with emotion is particularly arduous. It is also difficult to find content in ready-made picture books that includes the specific educational message parents want to convey. There is a need for a method that solves these problems, reduces the burden on parents, and effectively conveys educational messages to children. [Means for solving the problem]
[0005] The present invention is a system that allows users to select a theme for a video picture book, and then uses a generative AI model to generate and integrate an original story and illustrations to create a video picture book. Furthermore, users can select their child's favorite voice, use that voice to synthesize speech, and integrate it into the generated video picture book. This system allows parents to easily generate original video picture books and display them on their devices. Furthermore, by including a means to sample the parent's voice and generate deepfake voice, reading aloud in the parent's voice can be realized. Additionally, a means is provided to select a theme containing an educational message, allowing specific learning to be effectively conveyed to children. This reduces the burden on parents and supports their children's educational growth.
[0006] An animated picture book is a digital picture book in which the user selects a theme and generates a story, illustrations, and audio that are integrated.
[0007] A "theme" is a concept or subject that forms the basis of the content of an animated picture book, and includes educational messages such as "adventure," "dragons," and "eat all your food."
[0008] "User" refers to a person who uses this system to create animated picture books, typically a parent or guardian of a young child.
[0009] A "generative AI model" is an artificial intelligence model that automatically creates original stories and illustrations based on a theme selected by the user.
[0010] "Story" refers to the text and narrative content that unfolds within the animated picture book, and is generated based on the selected theme.
[0011] "Illustrations" are colorful, child-friendly diagrams and pictures that are generated to match the story content.
[0012] "Speech synthesis" is the process of providing audio for a story based on the audio format selected by the user.
[0013] "Deepfake voice" is a voice synthesized using AI technology based on a sampled parent's voice.
[0014] "Digital device" refers to an electronic device, such as an iPad, that displays the generated animated picture book.
[0015] "Educational messages" refer to specific lessons or morals included to promote children's growth and learning. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] This invention is a system in which a generative AI model creates a story and illustrations based on a theme selected by the user, and then integrates them to generate a video picture book. This system allows the user to set the desired audio and integrate the synthesized audio into the video picture book, allowing the system to display an original video picture book for young children on a device.
[0038] 1. Select a theme for your video book
[0039] The user selects a theme for the animated picture book. For example, they can choose a theme that includes an educational message such as "dragon," "adventure," or "eat all your food." The device receives the selected theme and sends the theme information to the server.
[0040] 2. Creating original stories and illustrations
[0041] Once the server receives the theme information, it uses a generative AI model to create an original story. The story can include episodes and characters that fit the theme. The server also uses the generative AI model to generate colorful, child-friendly illustrations to accompany the story. These stories and illustrations form the basis of the animated picture book.
[0042] 3. Voice settings and synthesis
[0043] The user selects the voice their child prefers, such as a deepfake voice sampled from a parent's voice or the voice of an animated character. The device then sends the selected voice settings to the server, which uses speech synthesis technology to generate the voice that matches the video storybook in the selected audio format.
[0044] 4. Integration and display of animated picture books
[0045] The server integrates the generated story, illustrations, and synthesized voice to create a complete animated picture book. The terminal receives the completed animated picture book and displays it on a digital device such as an iPad. This allows parents to easily read original animated picture books to their children using their digital devices.
[0046] Specific examples
[0047] For example, if the user selects the "Dragon" theme and requests the "Parent Voice" audio setting, the process will proceed as follows:
[0048] The user selects the "Dragon" theme and the device sends the theme to the server.
[0049] The server uses a generative AI model to generate original stories and illustrations based on the theme of "dragons."
[0050] The user selects "Parent's Voice" as the audio setting, and the device sends the setting to the server.
[0051] The server synthesizes voice using the parent's voice and integrates the voice into the animated picture book.
[0052] The completed video picture book is sent to the terminal and displayed on a digital device such as an iPad.
[0053] In this way, the system of the present invention allows parents to easily create animated picture books containing original educational messages and read them to their young children, thereby reducing the burden on parents and supporting their children's growth and learning.
[0054] The processing flow will be explained below.
[0055] Step 1:
[0056] The user selects the "theme of the animated picture book." For example, they can choose themes such as "Dragons" or "Eat all your food."
[0057] Step 2:
[0058] The terminal receives the user's selected theme information and transmits the information to the server.
[0059] Step 3:
[0060] The server analyzes the received theme information.
[0061] Step 4:
[0062] Based on the selected theme, the server uses a generative AI model to generate an original story, which includes episodes and characters that fit the theme.
[0063] Step 5:
[0064] Similarly, the server uses generative AI models to generate illustrations that fit the story, with colorful, child-friendly designs.
[0065] Step 6:
[0066] The server integrates the generated story and illustrations to create an animated picture book.
[0067] Step 7:
[0068] Users can choose the voice their child prefers, such as a deepfake voice sampled from a parent's voice or the voice of an animated character.
[0069] Step 8:
[0070] The device sends the user's selected audio settings to the server.
[0071] Step 9:
[0072] The server synthesizes the voice based on the selected voice format, either using the sampled parent voice to generate the deepfake voice or using the specified character voice.
[0073] Step 10:
[0074] The server integrates the synthesized voice into the generated animated picture book.
[0075] Step 11:
[0076] The server transmits the completed animated picture book data to the terminal.
[0077] Step 12:
[0078] The terminal displays the received video picture book on a digital device (such as an iPad), allowing parents to read original video picture books to their children using the digital device.
[0079] Through the above steps, users can easily create original animated picture books containing educational messages and easily read them to their children.
[0080] Example 1
[0081] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0082] In conventional video picture book generation systems, the process of generating original stories and illustrations based on themes and audio selected by the user and integrating them into videos is complicated. Furthermore, the use of voice synthesis technology is limited, making it difficult to accommodate individual audio settings desired by users. Therefore, there is a need for a simple and efficient way to generate original video picture books containing educational messages.
[0083] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0084] In this invention, the server includes means for a user to select a theme for the animated picture book, means for transmitting the selected theme information to the server, means for generating an original story based on the selected theme using a generative AI model, means for generating illustrations appropriate for the theme using the generative AI model, means for transmitting the generated story and illustrations to a terminal, means for a user to select audio settings, means for transmitting audio setting information to the server, means for generating audio in a selected audio format using speech synthesis technology, means for integrating the generated audio with the story and illustrations to create the animated picture book, and means for outputting the animated picture book to a display device. This allows a user to simply select a theme and set the audio, automatically generating an original animated picture book and efficiently providing high-quality content including educational messages.
[0085] "User" refers to an individual or organization that uses the system to create animated picture books.
[0086] "Theme" refers to the basic concept or subject matter of the content of the animated picture book.
[0087] "Terminal" refers to a computer or digital device operated by a user.
[0088] "Server" refers to the central processing unit that receives thematic information and audio setting information and generates stories and illustrations using generative AI models.
[0089] A "generative AI model" refers to an algorithm that uses artificial intelligence technology to generate text and images.
[0090] A "prompt sentence" is text data that is input into a generative AI model to guide the generated content based on a specific theme or setting.
[0091] "Original story" refers to a unique narrative generated by a generative AI model based on a selected theme.
[0092] "Illustrations" refer to visual images generated by a generative AI model based on a story.
[0093] "Audio Settings" refers to the particular setting options a user selects regarding the audio for an animated picture book.
[0094] "Speech synthesis technology" refers to technology for converting text data into speech.
[0095] An "animated picture book" is a digital picture book that integrates an original story, illustrations, and synthesized audio.
[0096] "Display device" refers to a digital device for displaying the generated animated picture book.
[0097] This invention is a system in which a generative AI model creates a story and illustrations based on a theme selected by the user, and then integrates them to generate a video picture book. This system allows the user to set the desired audio and integrate the synthesized audio into the video picture book, allowing the system to display an original video picture book for young children on a device.
[0098] Hardware and software used
[0099] Device: A digital device operated by a user (e.g., tablet, smartphone)
[0100] Server: Central processing unit (e.g., cloud server)
[0101] Generative AI models: Natural language processing and image generation techniques (e.g., GPT-4, DALL-E 2)
[0102] Speech synthesis technology: Technology that converts text into speech (e.g., Google Text-to-Speech API)
[0103] How it works
[0104] 1. Choose a theme
[0105] The user operates the device to select a theme for the animated picture book from a list of themes provided, such as "Dragons" or "Eat all your food."
[0106] The terminal transmits the selected theme information to the server.
[0107] 2. Story and illustration generation
[0108] The server generates an original story based on the received theme information using a generative AI model (e.g., GPT-4), and also generates colorful, child-friendly illustrations that fit the story using a generative AI model (e.g., DALL-E 2).
[0109] 3. Voice settings and synthesis
[0110] The user selects the voice their child prefers on the device. For example, they can choose the voice of a parent or an anime character. The device then sends the selected voice setting information to the server. The server then uses speech synthesis technology (e.g., Google Text-to-Speech API) to generate speech in the selected voice format.
[0111] 4. Video picture book integration
[0112] The server completes the animated picture book by integrating the generated story, illustrations, and synthesized voice.
[0113] 5. Outputting a video picture book
[0114] The completed animated picture book data is sent from the server to the terminal and displayed on the user's device.
[0115] Specific examples
[0116] If the user selects the "Dragon" theme and requests the "Parent Voice" audio setting, the process will proceed as follows:
[0117] The user selects the "Dragon" theme and the device sends the theme to the server.
[0118] The server generates an original story based on the theme of "dragon" using a generative AI model (e.g., GPT-4) and generates illustrations using a generative AI model (e.g., DALL-E 2).
[0119] The user selects "Parent's Voice" as the audio setting, and the device sends the setting information to the server.
[0120] The server uses the parent's voice to synthesize the text and integrate it into the story and illustrations.
[0121] The completed animated picture book is sent to the terminal and displayed on the device.
[0122] Examples of prompt statements
[0123] "Create a dragon adventure story. The main character is a brave child who befriends a dragon. The illustrations should be colorful and child-friendly."
[0124] By inputting this prompt into a generative AI model, the system generates a story and illustrations based on the specified theme and setting.
[0125] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0126] Step 1:
[0127] The user selects the theme of the animated picture book.
[0128] The user operates the device to select a theme for the animated picture book (e.g., "Dragon") from a list of themes provided. The input data is the theme selected by the user, and the output data is this theme information. Specifically, the user taps the theme name using the device's touchscreen.
[0129] Step 2:
[0130] The terminal transmits the theme information to the server.
[0131] The terminal sends the selected theme information to the server as an HTTP POST request. The input data is the theme information, and the output data is the request sent to the server. In concrete terms, the terminal sends the data to the server via the network.
[0132] Step 3:
[0133] The server generates the original story.
[0134] The server generates an original story using a generative AI model (e.g., GPT-4) based on the received theme information. The input data is the theme information, and the output data is the generated story. Specifically, based on the "dragon theme," the server sends a prompt to the generative AI model and receives the generated story.
[0135] Step 4:
[0136] The server generates the illustration.
[0137] The server uses a generative AI model (e.g., DALL-E 2) to generate illustrations appropriate for the generated story. The input data is the generated story, and the output data is the generated illustration. Specifically, the server uses the generated story as a prompt, sends it to the generative AI model, and receives the illustration.
[0138] Step 5:
[0139] The server sends the generated story and illustrations to the device.
[0140] The server sends the generated story and illustration data to the terminal as an HTTP response. The input data is the generated story and illustration, and the output data is the request sent to the terminal. In concrete terms, the server sends the data to the terminal via the network.
[0141] Step 6:
[0142] The user selects an audio setting.
[0143] The user selects the desired voice (e.g., "Parent's Voice") from the voice setting options provided on the device. The input data is the user's voice setting selection, and the output data is the selected voice setting. Specifically, the user taps the voice option using the device's touchscreen.
[0144] Step 7:
[0145] The terminal transmits audio setting information to the server.
[0146] The terminal sends the audio setting information selected by the user to the server as an HTTP POST request. The input data is the audio setting information, and the output data is the request sent to the server. Specifically, the terminal sends the data to the server via the network.
[0147] Step 8:
[0148] The server generates the audio and integrates it into the story and illustrations.
[0149] The server uses speech synthesis technology (e.g., Google Text-to-Speech API) to generate audio in the selected audio format and integrates the generated audio into the story and illustrations. The input data is audio setting information and the generated story and illustrations, and the output data is a complete animated picture book with the integrated audio. Specifically, the server synthesizes audio using the "parent's voice" and integrates it into the pre-generated story and illustrations.
[0150] Step 9:
[0151] The server sends the completed animated picture book to the terminal.
[0152] The server sends the completed animated picture book data to the terminal as an HTTP response. The input data is the completed animated picture book, and the output data is the request sent to the terminal. In concrete terms, the server sends the data to the terminal via the network.
[0153] Step 10:
[0154] The device displays the animated picture book.
[0155] The terminal displays the received animated picture book on a display device (e.g., a tablet or smartphone). The input data is the completed animated picture book, and the output data is the displayed animated picture book. The specific operation is to launch the terminal's video playback application and play the animated picture book.
[0156] The above are the specific processing steps of the system.
[0157] (Application example 1)
[0158] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0159] While existing video picture book generation systems can easily generate original stories and illustrations based on a user-selected theme, integrating the generated content into a single video picture book and playing it on the user's device requires significant effort. Especially when applied to content distribution services, an effective approach is needed to ensure easy user access and use. Furthermore, it is difficult to integrate diverse functions, such as voice synthesis using deepfake technology and the selection of themes that include educational messages.
[0160] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0161] In this invention, the server includes: a means for a user to select a theme for the animated picture book; a means for generating an original story based on the selected theme using a generative AI model; a means for generating illustrations appropriate to the theme using the generative AI model; a means for integrating the generated story and illustrations to create an animated picture book; a means for selecting a user's preferred audio; a means for synthesizing the selected audio; a means for integrating the synthesized audio into the animated picture book; a means for outputting the animated picture book to a display device; a means for providing the content distribution service as an application for a smartphone or tablet; a means for inputting a prompt to the generative AI model; a means for integrating the generated story and illustrations into a single animated picture book; and a means for playing the animated picture book on a display device. This allows users to easily create original animated picture books and play them on digital devices. Furthermore, a multifunctional animated picture book generation system that can be used in a user-friendly environment can also be provided for content distribution services.
[0162] An animated picture book is a digital picture book that combines illustrations and a story created based on a theme selected by the user, and is played back along with audio.
[0163] A "generative AI model" is a type of artificial intelligence that automatically generates original stories and illustrations based on a given theme and prompt.
[0164] A "theme" is a user-selected subject or theme that forms the basis of the story and illustrations in the animated picture book.
[0165] "Speech synthesis" is a technology that creates narration and character voices corresponding to a generated story based on user-selected voice settings.
[0166] "Content distribution service" refers to a service that provides digital content to users via the Internet.
[0167] A "prompt" refers to a sentence or keyword that is input to a generative AI model to instruct it to produce a specific output.
[0168] A "display device" is a device (e.g., smartphone, tablet, smart TV) for visually playing the generated animated picture book.
[0169] "Deepfake voice" is a synthetic voice that is generated by sampling the voice of a specific individual and using artificial intelligence technology.
[0170] "User preferred voice" refers to the voice settings of narration and characters selected by the user for a specific purpose.
[0171] A system for implementing this invention is configured as follows: First, a user launches an application on a display device such as a smartphone, tablet, smart TV, etc. The user selects a theme for the animated picture book, and the theme information is sent from the terminal to a server.
[0172] The server uses a generative AI model (e.g., GPT-4, DALL-E) to generate an original story and illustrations based on the received theme information. This generative AI model generates appropriate output by inputting a prompt. For example, if the theme is "adventure," the following prompt is used:
[0173] Example story generation prompt: "The main character is a brave boy named Tarrou. He and his dragon friend explore a magical land. They..."
[0174] Example prompt for generating an illustration: "Draw a dragon flying through the sky, looking at the stars. Make sure the colors are colorful and child-friendly."
[0175] The generated story and illustrations are integrated on the server side to create a single animated picture book. The user selects their preferred audio, and the audio settings are sent from the device to the server. Using speech synthesis technology (e.g., Neural TTS), audio is generated in the selected audio format, and this audio is integrated into the animated picture book.
[0176] Finally, the completed animated picture book is sent to the user's device and played through the application, allowing users to easily create original animated picture books and view them on the spot.
[0177] The system uses display devices such as smartphones, tablets, and smart TVs, as well as cloud services that utilize a serverless architecture (e.g., AWS Lambda and Google Cloud Functions). Data is sent and received using API requests and responses in JSON format. Specific generative AI models used include GPT-4, DALL-E, and Stable Diffusion.
[0178] For example, if the user selects the "Dragon" theme and requests speech synthesis using the parent's voice, the following process occurs:
[0179] 1. The user selects the "Dragon" theme and the device sends the theme to the server.
[0180] 2. The server generates original stories and illustrations using a generative AI model based on the theme of "dragons."
[0181] 3. The user selects "Parent's Voice" as the audio setting, and the device sends the setting to the server.
[0182] 4. The server uses the parent's voice to synthesize the audio and integrate it into the animated picture book.
[0183] 5. The completed animated picture book is sent to the terminal and played on the display device through the application.
[0184] In this way, users can easily create original animated picture books and provide them to children as educational content.
[0185] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0186] Step 1:
[0187] The user selects a theme for the animated picture book. The user launches the application and selects one from the list of themes provided. The theme name selected by the user is input here, and the device sends this theme information to the server. The request data including the theme information is output.
[0188] Step 2:
[0189] The server receives the theme information. The server receives the theme information (e.g., "Adventure" or "Dragon") sent from the device and generates a prompt sentence to be input to the generative AI model based on this. The input is the theme information, and the output is the prompt sentence to be passed to the generative AI model.
[0190] Step 3:
[0191] The server generates a story using a generative AI model. Based on the generated prompt, it asks the generative AI model (e.g., GPT-4) to generate a story. The input is the prompt, and the output is the original story. For example, a story like this might be generated: "The main character is a brave boy named Taro. Taro and his friend, the dragon, go on an adventure in a magical land. They..."
[0192] Step 4:
[0193] The server generates illustrations using a generative AI model. Based on the generated story, a prompt is input into the illustration generation AI model (e.g., DALL-E, Stable Diffusion) to request the generation of an illustration. The input is a prompt related to the story text, and the output is an original illustration image. For example, a prompt might be, "Draw a dragon flying through the sky, looking at the shining stars in the night sky. Please make it colorful and suitable for children."
[0194] Step 5:
[0195] The server integrates the generated story and illustrations. The server interactively integrates the story and illustrations and compiles them into a single animated picture book. The input is the generated story text and illustration images, and the output is the constituent data of the animated picture book.
[0196] Step 6:
[0197] The user selects the voice they prefer. The user selects one of the voice options provided within the application (e.g., parent's voice, anime character's voice), and the setting information is sent from the device to the server. The input is the voice setting information selected by the user, and the output is the request data to the server.
[0198] Step 7:
[0199] The server generates the voice using speech synthesis technology. Based on the selected voice settings, it generates the story narration using speech synthesis technology (e.g., Neural TTS). The input is the voice setting information and the story text, and the output is the synthesized voice data.
[0200] Step 8:
[0201] The server integrates the synthesized audio into the animated picture book. The generated audio is added to the existing animated picture book's constituent data, and is finally integrated into an animated picture book with audio. The input is the synthesized audio data and the animated picture book's constituent data, and the output is the completed animated picture book data.
[0202] Step 9:
[0203] The server sends the completed animated picture book to the terminal. The completed animated picture book data is sent to the terminal so that it can be played in the application. The completed animated picture book data is input, and the result of data transfer to the terminal is output. The terminal plays the received animated picture book on a display device.
[0204] Through the above processing steps, users can easily create original animated picture books and view them on the spot.
[0205] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0206] This invention is a system that generates animated picture books by using a generative AI model to create and integrate stories and illustrations based on a theme selected by the user. This system allows users to set their desired audio and integrate the synthesized audio into the animated picture book. In addition, by combining it with an emotion engine that recognizes the user's emotions, it is possible to provide more personalized animated picture books.
[0207] 1. Select a theme for your video book
[0208] The user selects the "animated picture book theme." For example, they can choose themes such as "Dragons" or "Eat all your food." The device receives the selected theme and sends the theme information to the server.
[0209] 2. Emotion Recognition by Emotion Engine
[0210] An emotion engine installed on the device or server recognizes the user's emotions from their facial expressions and voice. This emotion data is used for processing in later steps.
[0211] 3. Creating original stories and illustrations
[0212] Once the server receives the theme information and emotion data, it uses a generative AI model to create an original story. Based on the emotion data, the character's actions and dialogue within the story are adjusted. Similarly, the generative AI model is used to generate illustrations that fit the theme and emotion, resulting in facial expressions and scenes that correspond to the emotion.
[0213] 4. Voice settings and synthesis
[0214] The user selects the voice their child prefers. For example, they can choose a deepfake voice sampled from a parent's voice or the voice of an animated character. The device then sends the selected voice settings to the server. The server then uses voice synthesis technology to generate a voice that matches the video book in the selected audio format. The tone and emotional expression of the voice are then adjusted based on the emotional data.
[0215] 5. Integration and display of animated picture books
[0216] The server integrates the generated story, illustrations, and synthesized voice to create a complete animated picture book. The terminal receives the completed animated picture book and displays it on a digital device such as an iPad. This allows users to easily read original animated picture books to their children using their digital devices.
[0217] Specific examples
[0218] For example, if the user selects the "Dragon" theme and requests the "Parent Voice" audio setting, the process will proceed as follows:
[0219] The user selects the "Dragon" theme and the device sends the theme to the server.
[0220] The emotion engine on the terminal or server recognizes the user's emotion and transmits the data to the server.
[0221] Based on the theme of "dragon," the server uses a generative AI model to generate original stories and illustrations, taking emotional data into consideration.
[0222] The user selects "Parent's Voice" as the audio setting, and the device sends the setting to the server.
[0223] The server synthesizes voice using the parent's voice, adjusts the tone and emotional expression of the voice based on the emotional data, and integrates the voice into the animated picture book.
[0224] The completed video picture book is sent to the terminal and displayed on a digital device such as an iPad.
[0225] In this way, the system of the present invention allows parents to easily create animated picture books containing original educational messages tailored to their children's emotions and interests, enabling them to read aloud to their young children more effectively, thereby reducing the burden on parents and providing individual support for their children's growth and learning.
[0226] The processing flow will be explained below.
[0227] Step 1:
[0228] The user selects the "theme of the animated picture book." For example, they can choose themes such as "Dragons" or "Eat all your food."
[0229] Step 2:
[0230] The terminal receives the user's selected theme information and transmits the information to the server.
[0231] Step 3:
[0232] The emotion engine uses the device's camera and microphone to recognize the user's emotions, collecting and analyzing the user's facial expressions and voice data to extract emotional data.
[0233] Step 4:
[0234] The device transmits the collected emotion data to a server.
[0235] Step 5:
[0236] The server analyzes the received theme information and emotion data.
[0237] Step 6:
[0238] The server generates an original story based on the theme information using a generative AI model, adjusting the characters' actions and dialogue based on the emotional data.
[0239] Step 7:
[0240] The server then uses a generative AI model to generate illustrations that fit the story, adjusting the character's facial expressions and the mood of the scene based on the emotional data.
[0241] Step 8:
[0242] The server integrates the generated story and illustrations to create an animated picture book.
[0243] Step 9:
[0244] Users can choose the voice their child prefers, such as a deepfake voice sampled from a parent's voice or the voice of an animated character.
[0245] Step 10:
[0246] The device sends the user's selected audio settings to the server.
[0247] Step 11:
[0248] The server synthesizes speech based on the selected voice format, adjusting the tone and emotional expression of the speech based on the emotional data.
[0249] Step 12:
[0250] The server integrates the synthesized audio into the generated animated picture book.
[0251] Step 13:
[0252] The server transmits the completed animated picture book data to the terminal.
[0253] Step 14:
[0254] The video picture book received by the terminal is displayed on a digital device (such as an iPad), allowing users to easily read original video picture books to their children using their digital device.
[0255] As a specific example, when the user selects "dragon" as the theme and "parent's voice" as the voice, the processing flow is as follows.
[0256] The user selects the "Dragon" theme and the terminal transmits the theme information to the server.
[0257] The device's emotion engine analyzes the user's facial expressions and voice and sends the emotion data to the server.
[0258] The server uses a generative AI model to generate original stories and illustrations based on thematic information and emotional data.
[0259] The user selects "Parent's Voice" as the audio setting, and the device sends the setting information to the server.
[0260] The server generates synthetic speech using the parent's voice and integrates the speech, which reflects emotional data, into the animated picture book.
[0261] The completed animated picture book is sent to the terminal and displayed on the user's digital device.
[0262] This allows parents to create original animated picture books that are personalized according to their emotions, enabling them to read to their children effectively.
[0263] Example 2
[0264] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0265] Conventional video picture book creation systems have had problems with personalization based on the user's emotions and interests, and lack of natural emotional expression in their voice synthesis. Furthermore, generating original stories and illustrations based on a theme and integrating them is not easy, placing a heavy burden on the user. Therefore, there is a need for a system that can recognize the user's emotions and easily create more personalized original video picture books.
[0266] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0267] In this invention, the server includes means for analyzing the user's facial expressions and voice to acquire emotional data, means for generating an original story based on the theme information and emotional data using a generative AI model, and means for adjusting the tone of the voice and emotional expression based on the emotional data using voice synthesis technology. This makes it possible to generate a story and illustrations that reflect the user's emotions and to perform voice synthesis that incorporates natural emotional expressions.
[0268] A "user" is an entity that uses the system to select the theme and audio settings for an animated picture book.
[0269] A "terminal" is a digital device that is operated by a user and has the role of transmitting theme information and audio setting information to a server.
[0270] The "server" is a central device that receives thematic information and emotional data selected by the user, and generates original stories and illustrations and synthesizes voices based on that information.
[0271] A "generative AI model" is an artificial intelligence algorithm that generates original stories and illustrations based on thematic information and emotional data.
[0272] "Theme information" is data indicating the content selected by the user as the subject of the animated picture book.
[0273] An "emotion engine" is a combination of software and hardware that analyzes a user's facial expressions and voice to recognize emotional data.
[0274] "Emotion data" is data that indicates the emotional state of the user as recognized by the emotion engine.
[0275] A "story" is a narrative plot that a generative AI model generates based on thematic information and emotional data.
[0276] "Illustrations" are visual images associated with a story that are generated by a generative AI model.
[0277] "Speech synthesis" is the process of using speech synthesis technology to generate narration or character voices in a user-defined voice format.
[0278] "Voice setting information" is data indicating the voice synthesis settings selected by the user.
[0279] An animated picture book is a digital picture book created by integrating generated stories, illustrations, and audio.
[0280] This invention is a system that generates animated picture books by using a generative AI model to create and integrate stories and illustrations based on a theme selected by the user. This system allows users to set their desired audio and integrate the synthesized audio into the animated picture book. In addition, by combining it with an emotion engine that recognizes the user's emotions, it is possible to provide more personalized animated picture books.
[0281] Specifically, the user first uses the device to select a theme for the animated picture book, such as "Dragons" or "Eat all your food." The device then sends the selected theme information to the server. The server is equipped with an emotion engine that analyzes the user's facial expressions and voice to obtain emotion data. This emotion data is then used for subsequent processing.
[0282] The server inputs the received theme information and emotional data into a generative AI model. The generative AI model generates an original story based on this information. Furthermore, the behavior and dialogue of the characters in the story are adjusted based on the emotional data. Similarly, illustrations appropriate for the theme and emotion are generated. This allows expressions and scenes to be depicted according to the emotion.
[0283] Next, the user uses the device to select the desired voice, such as "parent's voice" or "animated character's voice." The device then sends the selected voice setting to the server. The server then uses voice synthesis technology to generate voice that matches the selected voice format for the animated picture book. At this time, the tone and emotional expression of the voice are adjusted based on the emotional data.
[0284] Finally, the server integrates the generated story, illustrations, and synthesized voice to create a complete animated picture book. The completed animated picture book is sent to the terminal and displayed on a digital device such as an iPad. This allows users to easily read original animated picture books to their children using their digital devices.
[0285] For example, if the user selects the "Dragon" theme and uses "Parent Voice", the following happens:
[0286] 1. The user selects the "Dragon" theme, and the device sends the theme information to the server.
[0287] 2. The server's emotion engine analyzes the user's facial expressions and voice to obtain emotional data such as "it looks fun."
[0288] 3. The server inputs the theme information "dragon" and the emotion data "looks fun" into the AI model to generate an original story and illustrations.
[0289] 4. The user selects "Parent's voice" and the device sends the information to the server.
[0290] 5. The server uses speech synthesis technology to generate audio in the parent's voice and adjusts the tone and emotional expression based on the emotional data.
[0291] 6. The server combines the story, illustrations, and audio to complete the animated picture book and sends it to the device.
[0292] 7. The device will display the completed video picture book on a digital device such as an iPad.
[0293] An example prompt is, "The theme is dragons. Use your parent's voice to make the book character happy based on the emotion data."
[0294] In this way, the system of the present invention makes it possible to easily create original animated picture books that are in line with the user's emotions and interests, and can also include educational messages, thereby reducing the burden on the user and providing individual support for children's growth and learning.
[0295] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0296] Step 1: Select a theme
[0297] The user operates the device to select a theme for the animated picture book. The input is the theme selection information. Using this information as input, the device sends the theme information to the server. Specifically, the user selects a theme such as "dragon," and the device records this selection and sends it to the server. The output is the selected theme information.
[0298] Step 2: Emotion Recognition
[0299] The server's emotion engine analyzes the user's facial expressions and voice to obtain emotion data. The input is the user's facial expressions and voice information. The server analyzes this and outputs emotion data such as "happy" or "sad." Specifically, the emotion engine analyzes data captured by a camera or microphone.
[0300] Step 3: Generate an original story
[0301] The server uses a generative AI model to generate an original story based on thematic information and emotional data. The input is thematic information and emotional data. Based on this, the generative AI model generates a story, which is then output. Specifically, it generates a "fun story about a dragon's adventure."
[0302] Step 4: Creating an illustration
[0303] The server uses the same generative AI model to generate suitable illustrations based on thematic information and emotional data. The input is thematic information and emotional data. The output is the generated illustration. For example, an illustration of a "smiling dragon" is generated. Here, too, the generative AI model creates the illustration based on visual elements.
[0304] Step 5: Select your audio settings
[0305] The user selects the desired voice using the terminal. The input is the voice setting information selected by the user. The terminal sends this information to the server, and the voice setting information is output. Specifically, the user selects "parent's voice."
[0306] Step 6: Text-to-Speech
[0307] The server uses voice synthesis technology based on the voice setting information to generate voice in the voice format set by the user. The input is the voice setting information and emotional data. The server generates voice based on this, and outputs voice with the tone and emotional expression adjusted based on the emotional data. For example, a voice with a "parent's voice in a happy tone" may be generated.
[0308] Step 7: Integrating the video book
[0309] The server integrates the generated story, illustrations, and audio to create a moving picture book. The inputs are the story, illustrations, and audio. These are integrated and output as a single moving picture book. Specifically, the generated content is appropriately edited and compiled into a moving picture format.
[0310] Step 8: Output and display the animated picture book
[0311] The server sends the completed animated picture book to the terminal. The input is the integrated animated picture book file. The terminal receives this data and displays it on a digital device such as an iPad. The output is the animated picture book displayed on the digital device. Specifically, the terminal plays the data and allows the user to view the animated picture book.
[0312] (Application example 2)
[0313] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0314] Conventional animated picture book systems have difficulty personalizing content based on user emotions and preferences, resulting in generic content. Furthermore, they lack an interactive experience that utilizes the user's wearable device, leaving room for further improvement in user engagement. This invention aims to solve these problems by providing a personalization function that includes user emotion recognition and an experience that utilizes a wearable device.
[0315] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0316] In this invention, the server includes: means for a user to select a theme for the animated picture book; means for generating an original story based on the selected theme using a generative AI model; means for generating illustrations appropriate for the theme using the generative AI model; means for integrating the generated story and illustrations to create an animated picture book; means for selecting a voice preferred by the user; means for performing voice synthesis using the selected voice; means for integrating the synthesized voice into the animated picture book; means for outputting the animated picture book to a display device; means for recognizing emotions from the user's facial expressions and voice and generating a story and illustrations and adjusting the tone of the voice based on the emotions; and means for recognizing emotions and displaying the animated picture book using a wearable device such as smart glasses. This allows a personalized animated picture book to be generated based on the user's emotions and to be experienced interactively via the wearable device.
[0317] An "animated picture book" is a digital book that combines an original story and illustrations generated by a generative AI model based on a theme selected by the user, and displays them with synthesized audio.
[0318] A "means of selection" is a feature of an interface or application that allows a user to select a particular theme or voice.
[0319] A "generative AI model" is an artificial intelligence algorithm that automatically generates stories and illustrations based on given themes and conditions.
[0320] "Means for recognizing emotions" refers to technologies and systems for reading emotions from a user's facial expressions, voice, etc., and processing that emotional data.
[0321] "Voice synthesis" is a technology for generating specific voices and tones, and involves synthesizing sampled voices using deepfake technology, etc.
[0322] A "wearable device" is a digital device (e.g., smart glasses) that can be worn by the user and has functions such as displaying animated picture books and recognizing emotions.
[0323] A "display device" is a digital screen or projection device for displaying the generated animated picture book.
[0324] This invention is a system that generates animated picture books using a generative AI model by having the user select a theme, and displays them through a wearable device. Specific embodiments of each step are described below.
[0325] System Program
[0326] The server includes means for a user to select a theme for the animated picture book, means for generating an original story based on the selected theme using a generative AI model, means for generating illustrations suitable for the theme using a generative AI model, means for integrating the generated story and illustrations to create an animated picture book, means for selecting an audio preferred by the user, means for performing voice synthesis using the selected audio, means for integrating the synthesized audio into the animated picture book, means for outputting the animated picture book to a display device, means for recognizing emotions from the user's facial expressions and voice, and generating a story and illustrations and adjusting the tone of the voice based thereon, and means for recognizing emotions and displaying the animated picture book using a wearable device such as smart glasses.
[0327] Hardware and software used
[0328] Hardware: smart glasses, camera, microphone
[0329] Software: OpenCV (for facial and emotion recognition), generative AI models (e.g., GPT-4, DALL-E), speech synthesis libraries (e.g., Google TTS, Amazon Polly)
[0330] Processing Description
[0331] The server uses a generative AI model to generate an original story based on the theme and emotion data selected by the user. For example, if a user selects the theme "Space Adventure" and the emotion engine recognizes the user's excitement, the generative AI model (e.g., GPT-4) will generate a "Space Adventure" story that reflects the user's excitement.
[0332] Next, the server generates appropriate illustrations based on the generated story. Using a generative AI model (e.g., DALL-E), illustrations of scenes and characters within the story are created. For example, colorful illustrations of spaceships and aliens are generated.
[0333] Furthermore, the server generates a voice using a speech synthesis library based on the voice settings selected by the user (e.g., a parent's voice). This voice is adjusted in tone and emotional expression based on the emotion data. For example, a parent's voice generates a voice with a happy tone.
[0334] Finally, the server integrates the generated story, illustrations, and audio, and sends the animated picture book to the smart glasses, which then display the personalized animated picture book based on the emotion data on the user's front display.
[0335] Specific examples
[0336] The user selects the theme "Space Adventure," and the emotion engine recognizes the exciting emotion. Based on that data, a generative AI model (GPT-4) generates a unique story, and a generative AI model (DALL-E) generates illustrations. The user selects a parent's voice, and the speech synthesis library synthesizes the voice based on that voice. All elements are integrated to create a complete animated storybook, which is then sent to the smart glasses.
[0337] Prompt Sentence Examples
[0338] Theme: Space Adventure
[0339] User Emotion: Excitement
[0340] Generative Stories: Kids, Adventure, Educational
[0341] Story length: 5 minutes
[0342] Generated illustrations: Colorful, spaceship, alien
[0343] In this way, a highly personalized animated picture book is generated using a generative AI model, providing users with an interactive experience.
[0344] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0345] Step 1:
[0346] The user selects the theme of the animated picture book.
[0347] The input is the action of the user selecting a theme through the interface.
[0348] The terminal transmits the theme selected by the user to the server and provides the theme data to the server.
[0349] Step 2:
[0350] The server receives the themes sent by the user and inputs the theme data into the generative AI model.
[0351] The input is theme data, and prompt sentences for story generation are generated based on this.
[0352] The server generates a story using a generative AI model, which generates the necessary text data.
[0353] Step 3:
[0354] The terminal or server captures the user's facial expressions and voice and inputs them into an emotion recognition engine.
[0355] The input is the user's facial expression and voice data, based on which emotion data is generated.
[0356] The emotion recognition engine analyzes the data, recognizes the user's emotional state, and provides the data to the server.
[0357] Step 4:
[0358] The server uses the generated emotion data and theme data to input it back into the generative AI model and generate a prompt for generating an illustration.
[0359] The input is emotion data and thematic data, which are then used to generate prompts containing detailed instructions for generating illustrations.
[0360] The server uses a generative AI model to generate illustrations suited to the theme, generating image data.
[0361] Step 5:
[0362] Users select the audio settings they want to use.
[0363] The input is the audio settings selected by the user through the interface.
[0364] The terminal transmits the selected audio settings to the server and provides the server with the audio setting data.
[0365] Step 6:
[0366] The server generates the voice using a voice synthesis library based on the selected voice settings.
[0367] The inputs are the voice setting data and the generated story data, and voice data is generated based on this.
[0368] The server takes the synthesized voice data and adjusts its tone and emotional expression.
[0369] Step 7:
[0370] The server integrates the generated story, illustrations, and audio to complete the animated picture book.
[0371] The inputs are story data, illustration data, and audio data, and based on this, integrated video data is generated.
[0372] The server integrates these data and generates the completed animated picture book.
[0373] Step 8:
[0374] The terminal receives the completed animated picture book from the server and outputs it on a display device such as smart glasses.
[0375] The input is the synthesized video data, from which a video file in a format suitable for the display device is generated.
[0376] The terminal outputs the animated picture book to a display device, providing the user with a visually interactive experience.
[0377] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0378] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0379] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0380] [Second embodiment]
[0381] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0382] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0383] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0384] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0385] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0386] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0387] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0388] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0389] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0390] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0391] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0392] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0393] This invention is a system in which a generative AI model creates a story and illustrations based on a theme selected by the user, and then integrates them to generate a video picture book. This system allows the user to set the desired audio and integrate the synthesized audio into the video picture book, allowing the system to display an original video picture book for young children on a device.
[0394] 1. Select a theme for your video book
[0395] The user selects a theme for the animated picture book. For example, they can choose a theme that includes an educational message such as "dragon," "adventure," or "eat all your food." The device receives the selected theme and sends the theme information to the server.
[0396] 2. Creating original stories and illustrations
[0397] Once the server receives the theme information, it uses a generative AI model to create an original story. The story can include episodes and characters that fit the theme. The server also uses the generative AI model to generate colorful, child-friendly illustrations to accompany the story. These stories and illustrations form the basis of the animated picture book.
[0398] 3. Voice settings and synthesis
[0399] The user selects the voice their child prefers, such as a deepfake voice sampled from a parent's voice or the voice of an animated character. The device then sends the selected voice settings to the server, which uses speech synthesis technology to generate the voice that matches the video storybook in the selected audio format.
[0400] 4. Integration and display of animated picture books
[0401] The server integrates the generated story, illustrations, and synthesized voice to create a complete animated picture book. The terminal receives the completed animated picture book and displays it on a digital device such as an iPad. This allows parents to easily read original animated picture books to their children using their digital devices.
[0402] Specific examples
[0403] For example, if the user selects the "Dragon" theme and requests the "Parent Voice" audio setting, the process will proceed as follows:
[0404] The user selects the "Dragon" theme and the device sends the theme to the server.
[0405] The server uses a generative AI model to generate original stories and illustrations based on the theme of "dragons."
[0406] The user selects "Parent's Voice" as the audio setting, and the device sends the setting to the server.
[0407] The server synthesizes voice using the parent's voice and integrates the voice into the animated picture book.
[0408] The completed video picture book is sent to the terminal and displayed on a digital device such as an iPad.
[0409] In this way, the system of the present invention allows parents to easily create animated picture books containing original educational messages and read them to their young children, thereby reducing the burden on parents and supporting their children's growth and learning.
[0410] The processing flow will be explained below.
[0411] Step 1:
[0412] The user selects the "theme of the animated picture book." For example, they can choose themes such as "Dragons" or "Eat all your food."
[0413] Step 2:
[0414] The terminal receives the user's selected theme information and transmits the information to the server.
[0415] Step 3:
[0416] The server analyzes the received theme information.
[0417] Step 4:
[0418] Based on the selected theme, the server uses a generative AI model to generate an original story, which includes episodes and characters that fit the theme.
[0419] Step 5:
[0420] Similarly, the server uses generative AI models to generate illustrations that fit the story, with colorful, child-friendly designs.
[0421] Step 6:
[0422] The server integrates the generated story and illustrations to create an animated picture book.
[0423] Step 7:
[0424] Users can choose the voice their child prefers, such as a deepfake voice sampled from a parent's voice or the voice of an animated character.
[0425] Step 8:
[0426] The device sends the user's selected audio settings to the server.
[0427] Step 9:
[0428] The server synthesizes the voice based on the selected voice format, either using the sampled parent voice to generate the deepfake voice or using the specified character voice.
[0429] Step 10:
[0430] The server integrates the synthesized voice into the generated animated picture book.
[0431] Step 11:
[0432] The server transmits the completed animated picture book data to the terminal.
[0433] Step 12:
[0434] The terminal displays the received video picture book on a digital device (such as an iPad), allowing parents to read original video picture books to their children using the digital device.
[0435] Through the above steps, users can easily create original animated picture books containing educational messages and easily read them to their children.
[0436] Example 1
[0437] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0438] In conventional video picture book generation systems, the process of generating original stories and illustrations based on themes and audio selected by the user and integrating them into videos is complicated. Furthermore, the use of voice synthesis technology is limited, making it difficult to accommodate individual audio settings desired by users. Therefore, there is a need for a simple and efficient way to generate original video picture books containing educational messages.
[0439] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0440] In this invention, the server includes means for a user to select a theme for the animated picture book, means for transmitting the selected theme information to the server, means for generating an original story based on the selected theme using a generative AI model, means for generating illustrations appropriate for the theme using the generative AI model, means for transmitting the generated story and illustrations to a terminal, means for a user to select audio settings, means for transmitting audio setting information to the server, means for generating audio in a selected audio format using speech synthesis technology, means for integrating the generated audio with the story and illustrations to create the animated picture book, and means for outputting the animated picture book to a display device. This allows a user to simply select a theme and set the audio, automatically generating an original animated picture book and efficiently providing high-quality content including educational messages.
[0441] "User" refers to an individual or organization that uses the system to create animated picture books.
[0442] "Theme" refers to the basic concept or subject matter of the content of the animated picture book.
[0443] "Terminal" refers to a computer or digital device operated by a user.
[0444] "Server" refers to the central processing unit that receives thematic information and audio setting information and generates stories and illustrations using generative AI models.
[0445] A "generative AI model" refers to an algorithm that uses artificial intelligence technology to generate text and images.
[0446] A "prompt sentence" is text data that is input into a generative AI model to guide the generated content based on a specific theme or setting.
[0447] "Original story" refers to a unique narrative generated by a generative AI model based on a selected theme.
[0448] "Illustrations" refer to visual images generated by a generative AI model based on a story.
[0449] "Audio Settings" refers to the particular setting options a user selects regarding the audio for an animated picture book.
[0450] "Speech synthesis technology" refers to technology for converting text data into speech.
[0451] An "animated picture book" is a digital picture book that integrates an original story, illustrations, and synthesized audio.
[0452] "Display device" refers to a digital device for displaying the generated animated picture book.
[0453] This invention is a system in which a generative AI model creates a story and illustrations based on a theme selected by the user, and then integrates them to generate a video picture book. This system allows the user to set the desired audio and integrate the synthesized audio into the video picture book, allowing the system to display an original video picture book for young children on a device.
[0454] Hardware and software used
[0455] Device: A digital device operated by a user (e.g., tablet, smartphone)
[0456] Server: Central processing unit (e.g., cloud server)
[0457] Generative AI models: Natural language processing and image generation techniques (e.g., GPT-4, DALL-E 2)
[0458] Speech synthesis technology: Technology that converts text into speech (e.g., Google Text-to-Speech API)
[0459] How it works
[0460] 1. Choose a theme
[0461] The user operates the device to select a theme for the animated picture book from a list of themes provided, such as "Dragons" or "Eat all your food."
[0462] The terminal transmits the selected theme information to the server.
[0463] 2. Story and illustration generation
[0464] The server generates an original story based on the received theme information using a generative AI model (e.g., GPT-4), and also generates colorful, child-friendly illustrations that fit the story using a generative AI model (e.g., DALL-E 2).
[0465] 3. Voice settings and synthesis
[0466] The user selects the voice their child prefers on the device. For example, they can choose the voice of a parent or an anime character. The device then sends the selected voice setting information to the server. The server then uses speech synthesis technology (e.g., Google Text-to-Speech API) to generate speech in the selected voice format.
[0467] 4. Video picture book integration
[0468] The server completes the animated picture book by integrating the generated story, illustrations, and synthesized voice.
[0469] 5. Outputting a video picture book
[0470] The completed animated picture book data is sent from the server to the terminal and displayed on the user's device.
[0471] Specific examples
[0472] If the user selects the "Dragon" theme and requests the "Parent Voice" audio setting, the process will proceed as follows:
[0473] The user selects the "Dragon" theme and the device sends the theme to the server.
[0474] The server generates an original story based on the theme of "dragon" using a generative AI model (e.g., GPT-4) and generates illustrations using a generative AI model (e.g., DALL-E 2).
[0475] The user selects "Parent's Voice" as the audio setting, and the device sends the setting information to the server.
[0476] The server uses the parent's voice to synthesize the text and integrate it into the story and illustrations.
[0477] The completed animated picture book is sent to the terminal and displayed on the device.
[0478] Examples of prompt statements
[0479] "Create a dragon adventure story. The main character is a brave child who befriends a dragon. The illustrations should be colorful and child-friendly."
[0480] By inputting this prompt into a generative AI model, the system generates a story and illustrations based on the specified theme and setting.
[0481] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0482] Step 1:
[0483] The user selects the theme of the animated picture book.
[0484] The user operates the device to select a theme for the animated picture book (e.g., "Dragon") from a list of themes provided. The input data is the theme selected by the user, and the output data is this theme information. Specifically, the user taps the theme name using the device's touchscreen.
[0485] Step 2:
[0486] The terminal transmits the theme information to the server.
[0487] The terminal sends the selected theme information to the server as an HTTP POST request. The input data is the theme information, and the output data is the request sent to the server. In concrete terms, the terminal sends the data to the server via the network.
[0488] Step 3:
[0489] The server generates the original story.
[0490] The server generates an original story using a generative AI model (e.g., GPT-4) based on the received theme information. The input data is the theme information, and the output data is the generated story. Specifically, based on the "dragon theme," the server sends a prompt to the generative AI model and receives the generated story.
[0491] Step 4:
[0492] The server generates the illustration.
[0493] The server uses a generative AI model (e.g., DALL-E 2) to generate illustrations appropriate for the generated story. The input data is the generated story, and the output data is the generated illustration. Specifically, the server uses the generated story as a prompt, sends it to the generative AI model, and receives the illustration.
[0494] Step 5:
[0495] The server sends the generated story and illustrations to the device.
[0496] The server sends the generated story and illustration data to the terminal as an HTTP response. The input data is the generated story and illustration, and the output data is the request sent to the terminal. In concrete terms, the server sends the data to the terminal via the network.
[0497] Step 6:
[0498] The user selects an audio setting.
[0499] The user selects the desired voice (e.g., "Parent's Voice") from the voice setting options provided on the device. The input data is the user's voice setting selection, and the output data is the selected voice setting. Specifically, the user taps the voice option using the device's touchscreen.
[0500] Step 7:
[0501] The terminal transmits audio setting information to the server.
[0502] The terminal sends the audio setting information selected by the user to the server as an HTTP POST request. The input data is the audio setting information, and the output data is the request sent to the server. Specifically, the terminal sends the data to the server via the network.
[0503] Step 8:
[0504] The server generates the audio and integrates it into the story and illustrations.
[0505] The server uses speech synthesis technology (e.g., Google Text-to-Speech API) to generate audio in the selected audio format and integrates the generated audio into the story and illustrations. The input data is audio setting information and the generated story and illustrations, and the output data is a complete animated picture book with the integrated audio. Specifically, the server synthesizes audio using the "parent's voice" and integrates it into the pre-generated story and illustrations.
[0506] Step 9:
[0507] The server sends the completed animated picture book to the terminal.
[0508] The server sends the completed animated picture book data to the terminal as an HTTP response. The input data is the completed animated picture book, and the output data is the request sent to the terminal. In concrete terms, the server sends the data to the terminal via the network.
[0509] Step 10:
[0510] The device displays the animated picture book.
[0511] The terminal displays the received animated picture book on a display device (e.g., a tablet or smartphone). The input data is the completed animated picture book, and the output data is the displayed animated picture book. The specific operation is to launch the terminal's video playback application and play the animated picture book.
[0512] The above are the specific processing steps of the system.
[0513] (Application example 1)
[0514] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0515] While existing video picture book generation systems can easily generate original stories and illustrations based on a user-selected theme, integrating the generated content into a single video picture book and playing it on the user's device requires significant effort. Especially when applied to content distribution services, an effective approach is needed to ensure easy user access and use. Furthermore, it is difficult to integrate diverse functions, such as voice synthesis using deepfake technology and the selection of themes that include educational messages.
[0516] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0517] In this invention, the server includes: a means for a user to select a theme for the animated picture book; a means for generating an original story based on the selected theme using a generative AI model; a means for generating illustrations appropriate to the theme using the generative AI model; a means for integrating the generated story and illustrations to create an animated picture book; a means for selecting a user's preferred audio; a means for synthesizing the selected audio; a means for integrating the synthesized audio into the animated picture book; a means for outputting the animated picture book to a display device; a means for providing the content distribution service as an application for a smartphone or tablet; a means for inputting a prompt to the generative AI model; a means for integrating the generated story and illustrations into a single animated picture book; and a means for playing the animated picture book on a display device. This allows users to easily create original animated picture books and play them on digital devices. Furthermore, a multifunctional animated picture book generation system that can be used in a user-friendly environment can also be provided for content distribution services.
[0518] An animated picture book is a digital picture book that combines illustrations and a story created based on a theme selected by the user, and is played back along with audio.
[0519] A "generative AI model" is a type of artificial intelligence that automatically generates original stories and illustrations based on a given theme and prompt.
[0520] A "theme" is a user-selected subject or theme that forms the basis of the story and illustrations in the animated picture book.
[0521] "Speech synthesis" is a technology that creates narration and character voices corresponding to a generated story based on user-selected voice settings.
[0522] "Content distribution service" refers to a service that provides digital content to users via the Internet.
[0523] A "prompt" refers to a sentence or keyword that is input to a generative AI model to instruct it to produce a specific output.
[0524] A "display device" is a device (e.g., smartphone, tablet, smart TV) for visually playing the generated animated picture book.
[0525] "Deepfake voice" is a synthetic voice that is generated by sampling the voice of a specific individual and using artificial intelligence technology.
[0526] "User preferred voice" refers to the voice settings of narration and characters selected by the user for a specific purpose.
[0527] A system for implementing this invention is configured as follows: First, a user launches an application on a display device such as a smartphone, tablet, smart TV, etc. The user selects a theme for the animated picture book, and the theme information is sent from the terminal to a server.
[0528] The server uses a generative AI model (e.g., GPT-4, DALL-E) to generate an original story and illustrations based on the received theme information. This generative AI model generates appropriate output by inputting a prompt. For example, if the theme is "adventure," the following prompt is used:
[0529] Example story generation prompt: "The main character is a brave boy named Tarrou. He and his dragon friend explore a magical land. They..."
[0530] Example prompt for generating an illustration: "Draw a dragon flying through the sky, looking at the stars. Make sure the colors are colorful and child-friendly."
[0531] The generated story and illustrations are integrated on the server side to create a single animated picture book. The user selects their preferred audio, and the audio settings are sent from the device to the server. Using speech synthesis technology (e.g., Neural TTS), audio is generated in the selected audio format, and this audio is integrated into the animated picture book.
[0532] Finally, the completed animated picture book is sent to the user's device and played through the application, allowing users to easily create original animated picture books and view them on the spot.
[0533] The system uses display devices such as smartphones, tablets, and smart TVs, as well as cloud services that utilize a serverless architecture (e.g., AWS Lambda and Google Cloud Functions). Data is sent and received using API requests and responses in JSON format. Specific generative AI models used include GPT-4, DALL-E, and Stable Diffusion.
[0534] For example, if the user selects the "Dragon" theme and requests speech synthesis using the parent's voice, the following process occurs:
[0535] 1. The user selects the "Dragon" theme and the device sends the theme to the server.
[0536] 2. The server generates original stories and illustrations using a generative AI model based on the theme of "dragons."
[0537] 3. The user selects "Parent's Voice" as the audio setting, and the device sends the setting to the server.
[0538] 4. The server uses the parent's voice to synthesize the audio and integrate it into the animated picture book.
[0539] 5. The completed animated picture book is sent to the terminal and played on the display device through the application.
[0540] In this way, users can easily create original animated picture books and provide them to children as educational content.
[0541] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0542] Step 1:
[0543] The user selects a theme for the animated picture book. The user launches the application and selects one from the list of themes provided. The theme name selected by the user is input here, and the device sends this theme information to the server. The request data including the theme information is output.
[0544] Step 2:
[0545] The server receives the theme information. The server receives the theme information (e.g., "Adventure" or "Dragon") sent from the device and generates a prompt sentence to be input to the generative AI model based on this. The input is the theme information, and the output is the prompt sentence to be passed to the generative AI model.
[0546] Step 3:
[0547] The server generates a story using a generative AI model. Based on the generated prompt, it asks the generative AI model (e.g., GPT-4) to generate a story. The input is the prompt, and the output is the original story. For example, a story like this might be generated: "The main character is a brave boy named Taro. Taro and his friend, the dragon, go on an adventure in a magical land. They..."
[0548] Step 4:
[0549] The server generates illustrations using a generative AI model. Based on the generated story, a prompt is input into the illustration generation AI model (e.g., DALL-E, Stable Diffusion) to request the generation of an illustration. The input is a prompt related to the story text, and the output is an original illustration image. For example, a prompt might be, "Draw a dragon flying through the sky, looking at the shining stars in the night sky. Please make it colorful and suitable for children."
[0550] Step 5:
[0551] The server integrates the generated story and illustrations. The server interactively integrates the story and illustrations and compiles them into a single animated picture book. The input is the generated story text and illustration images, and the output is the constituent data of the animated picture book.
[0552] Step 6:
[0553] The user selects the voice they prefer. The user selects one of the voice options provided within the application (e.g., parent's voice, anime character's voice), and the setting information is sent from the device to the server. The input is the voice setting information selected by the user, and the output is the request data to the server.
[0554] Step 7:
[0555] The server generates the voice using speech synthesis technology. Based on the selected voice settings, it generates the story narration using speech synthesis technology (e.g., Neural TTS). The input is the voice setting information and the story text, and the output is the synthesized voice data.
[0556] Step 8:
[0557] The server integrates the synthesized audio into the animated picture book. The generated audio is added to the existing animated picture book's constituent data, and is finally integrated into an animated picture book with audio. The input is the synthesized audio data and the animated picture book's constituent data, and the output is the completed animated picture book data.
[0558] Step 9:
[0559] The server sends the completed animated picture book to the terminal. The completed animated picture book data is sent to the terminal so that it can be played in the application. The completed animated picture book data is input, and the result of data transfer to the terminal is output. The terminal plays the received animated picture book on a display device.
[0560] Through the above processing steps, users can easily create original animated picture books and view them on the spot.
[0561] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0562] This invention is a system that generates animated picture books by using a generative AI model to create and integrate stories and illustrations based on a theme selected by the user. This system allows users to set their desired audio and integrate the synthesized audio into the animated picture book. In addition, by combining it with an emotion engine that recognizes the user's emotions, it is possible to provide more personalized animated picture books.
[0563] 1. Select a theme for your video book
[0564] The user selects the "animated picture book theme." For example, they can choose themes such as "Dragons" or "Eat all your food." The device receives the selected theme and sends the theme information to the server.
[0565] 2. Emotion Recognition by Emotion Engine
[0566] An emotion engine installed on the device or server recognizes the user's emotions from their facial expressions and voice. This emotion data is used for processing in later steps.
[0567] 3. Creating original stories and illustrations
[0568] Once the server receives the theme information and emotion data, it uses a generative AI model to create an original story. Based on the emotion data, the character's actions and dialogue within the story are adjusted. Similarly, the generative AI model is used to generate illustrations that fit the theme and emotion, resulting in facial expressions and scenes that correspond to the emotion.
[0569] 4. Voice settings and synthesis
[0570] The user selects the voice their child prefers. For example, they can choose a deepfake voice sampled from a parent's voice or the voice of an animated character. The device then sends the selected voice settings to the server. The server then uses voice synthesis technology to generate a voice that matches the video book in the selected audio format. The tone and emotional expression of the voice are then adjusted based on the emotional data.
[0571] 5. Integration and display of animated picture books
[0572] The server integrates the generated story, illustrations, and synthesized voice to create a complete animated picture book. The terminal receives the completed animated picture book and displays it on a digital device such as an iPad. This allows users to easily read original animated picture books to their children using their digital devices.
[0573] Specific examples
[0574] For example, if the user selects the "Dragon" theme and requests the "Parent Voice" audio setting, the process will proceed as follows:
[0575] The user selects the "Dragon" theme and the device sends the theme to the server.
[0576] The emotion engine on the terminal or server recognizes the user's emotion and transmits the data to the server.
[0577] Based on the theme of "dragon," the server uses a generative AI model to generate original stories and illustrations, taking emotional data into consideration.
[0578] The user selects "Parent's Voice" as the audio setting, and the device sends the setting to the server.
[0579] The server synthesizes voice using the parent's voice, adjusts the tone and emotional expression of the voice based on the emotional data, and integrates the voice into the animated picture book.
[0580] The completed video picture book is sent to the terminal and displayed on a digital device such as an iPad.
[0581] In this way, the system of the present invention allows parents to easily create animated picture books containing original educational messages tailored to their children's emotions and interests, enabling them to read aloud to their young children more effectively, thereby reducing the burden on parents and providing individual support for their children's growth and learning.
[0582] The processing flow will be explained below.
[0583] Step 1:
[0584] The user selects the "theme of the animated picture book." For example, they can choose themes such as "Dragons" or "Eat all your food."
[0585] Step 2:
[0586] The terminal receives the user's selected theme information and transmits the information to the server.
[0587] Step 3:
[0588] The emotion engine uses the device's camera and microphone to recognize the user's emotions, collecting and analyzing the user's facial expressions and voice data to extract emotional data.
[0589] Step 4:
[0590] The device transmits the collected emotion data to a server.
[0591] Step 5:
[0592] The server analyzes the received theme information and emotion data.
[0593] Step 6:
[0594] The server generates an original story based on the theme information using a generative AI model, adjusting the characters' actions and dialogue based on the emotional data.
[0595] Step 7:
[0596] The server then uses a generative AI model to generate illustrations that fit the story, adjusting the character's facial expressions and the mood of the scene based on the emotional data.
[0597] Step 8:
[0598] The server integrates the generated story and illustrations to create an animated picture book.
[0599] Step 9:
[0600] Users can choose the voice their child prefers, such as a deepfake voice sampled from a parent's voice or the voice of an animated character.
[0601] Step 10:
[0602] The device sends the user's selected audio settings to the server.
[0603] Step 11:
[0604] The server synthesizes speech based on the selected voice format, adjusting the tone and emotional expression of the speech based on the emotional data.
[0605] Step 12:
[0606] The server integrates the synthesized audio into the generated animated picture book.
[0607] Step 13:
[0608] The server transmits the completed animated picture book data to the terminal.
[0609] Step 14:
[0610] The video picture book received by the terminal is displayed on a digital device (such as an iPad), allowing users to easily read original video picture books to their children using their digital device.
[0611] As a specific example, when the user selects "dragon" as the theme and "parent's voice" as the voice, the processing flow is as follows.
[0612] The user selects the "Dragon" theme and the terminal transmits the theme information to the server.
[0613] The device's emotion engine analyzes the user's facial expressions and voice and sends the emotion data to the server.
[0614] The server uses a generative AI model to generate original stories and illustrations based on thematic information and emotional data.
[0615] The user selects "Parent's Voice" as the audio setting, and the device sends the setting information to the server.
[0616] The server generates synthetic speech using the parent's voice and integrates the speech, which reflects emotional data, into the animated picture book.
[0617] The completed animated picture book is sent to the terminal and displayed on the user's digital device.
[0618] This allows parents to create original animated picture books that are personalized according to their emotions, enabling them to read to their children effectively.
[0619] Example 2
[0620] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0621] Conventional video picture book creation systems have had problems with personalization based on the user's emotions and interests, and lack of natural emotional expression in their voice synthesis. Furthermore, generating original stories and illustrations based on a theme and integrating them is not easy, placing a heavy burden on the user. Therefore, there is a need for a system that can recognize the user's emotions and easily create more personalized original video picture books.
[0622] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0623] In this invention, the server includes means for analyzing the user's facial expressions and voice to acquire emotional data, means for generating an original story based on the theme information and emotional data using a generative AI model, and means for adjusting the tone of the voice and emotional expression based on the emotional data using voice synthesis technology. This makes it possible to generate a story and illustrations that reflect the user's emotions and to perform voice synthesis that incorporates natural emotional expressions.
[0624] A "user" is an entity that uses the system to select the theme and audio settings for an animated picture book.
[0625] A "terminal" is a digital device that is operated by a user and has the role of transmitting theme information and audio setting information to a server.
[0626] The "server" is a central device that receives thematic information and emotional data selected by the user, and generates original stories and illustrations and synthesizes voices based on that information.
[0627] A "generative AI model" is an artificial intelligence algorithm that generates original stories and illustrations based on thematic information and emotional data.
[0628] "Theme information" is data indicating the content selected by the user as the subject of the animated picture book.
[0629] An "emotion engine" is a combination of software and hardware that analyzes a user's facial expressions and voice to recognize emotional data.
[0630] "Emotion data" is data that indicates the emotional state of the user as recognized by the emotion engine.
[0631] A "story" is a narrative plot that a generative AI model generates based on thematic information and emotional data.
[0632] "Illustrations" are visual images associated with a story that are generated by a generative AI model.
[0633] "Speech synthesis" is the process of using speech synthesis technology to generate narration or character voices in a user-defined voice format.
[0634] "Voice setting information" is data indicating the voice synthesis settings selected by the user.
[0635] An animated picture book is a digital picture book created by integrating generated stories, illustrations, and audio.
[0636] This invention is a system that generates animated picture books by using a generative AI model to create and integrate stories and illustrations based on a theme selected by the user. This system allows users to set their desired audio and integrate the synthesized audio into the animated picture book. In addition, by combining it with an emotion engine that recognizes the user's emotions, it is possible to provide more personalized animated picture books.
[0637] Specifically, the user first uses the device to select a theme for the animated picture book, such as "Dragons" or "Eat all your food." The device then sends the selected theme information to the server. The server is equipped with an emotion engine that analyzes the user's facial expressions and voice to obtain emotion data. This emotion data is then used for subsequent processing.
[0638] The server inputs the received theme information and emotional data into a generative AI model. The generative AI model generates an original story based on this information. Furthermore, the behavior and dialogue of the characters in the story are adjusted based on the emotional data. Similarly, illustrations appropriate for the theme and emotion are generated. This allows expressions and scenes to be depicted according to the emotion.
[0639] Next, the user uses the device to select the desired voice, such as "parent's voice" or "animated character's voice." The device then sends the selected voice setting to the server. The server then uses voice synthesis technology to generate voice that matches the selected voice format for the animated picture book. At this time, the tone and emotional expression of the voice are adjusted based on the emotional data.
[0640] Finally, the server integrates the generated story, illustrations, and synthesized voice to create a complete animated picture book. The completed animated picture book is sent to the terminal and displayed on a digital device such as an iPad. This allows users to easily read original animated picture books to their children using their digital devices.
[0641] For example, if the user selects the "Dragon" theme and uses "Parent Voice", the following happens:
[0642] 1. The user selects the "Dragon" theme, and the device sends the theme information to the server.
[0643] 2. The server's emotion engine analyzes the user's facial expressions and voice to obtain emotional data such as "it looks fun."
[0644] 3. The server inputs the theme information "dragon" and the emotion data "looks fun" into the AI model to generate an original story and illustrations.
[0645] 4. The user selects "Parent's voice" and the device sends the information to the server.
[0646] 5. The server uses speech synthesis technology to generate audio in the parent's voice and adjusts the tone and emotional expression based on the emotional data.
[0647] 6. The server combines the story, illustrations, and audio to complete the animated picture book and sends it to the device.
[0648] 7. The device will display the completed video picture book on a digital device such as an iPad.
[0649] An example prompt is, "The theme is dragons. Use your parent's voice to make the book character happy based on the emotion data."
[0650] In this way, the system of the present invention makes it possible to easily create original animated picture books that are in line with the user's emotions and interests, and can also include educational messages, thereby reducing the burden on the user and providing individual support for children's growth and learning.
[0651] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0652] Step 1: Select a theme
[0653] The user operates the device to select a theme for the animated picture book. The input is the theme selection information. Using this information as input, the device sends the theme information to the server. Specifically, the user selects a theme such as "dragon," and the device records this selection and sends it to the server. The output is the selected theme information.
[0654] Step 2: Emotion Recognition
[0655] The server's emotion engine analyzes the user's facial expressions and voice to obtain emotion data. The input is the user's facial expressions and voice information. The server analyzes this and outputs emotion data such as "happy" or "sad." Specifically, the emotion engine analyzes data captured by a camera or microphone.
[0656] Step 3: Generate an original story
[0657] The server uses a generative AI model to generate an original story based on thematic information and emotional data. The input is thematic information and emotional data. Based on this, the generative AI model generates a story, which is then output. Specifically, it generates a "fun story about a dragon's adventure."
[0658] Step 4: Creating an illustration
[0659] The server uses the same generative AI model to generate suitable illustrations based on thematic information and emotional data. The input is thematic information and emotional data. The output is the generated illustration. For example, an illustration of a "smiling dragon" is generated. Here, too, the generative AI model creates the illustration based on visual elements.
[0660] Step 5: Select your audio settings
[0661] The user selects the desired voice using the terminal. The input is the voice setting information selected by the user. The terminal sends this information to the server, and the voice setting information is output. Specifically, the user selects "parent's voice."
[0662] Step 6: Text-to-Speech
[0663] The server uses voice synthesis technology based on the voice setting information to generate voice in the voice format set by the user. The input is the voice setting information and emotional data. The server generates voice based on this, and outputs voice with the tone and emotional expression adjusted based on the emotional data. For example, a voice with a "parent's voice in a happy tone" may be generated.
[0664] Step 7: Integrating the video book
[0665] The server integrates the generated story, illustrations, and audio to create a moving picture book. The inputs are the story, illustrations, and audio. These are integrated and output as a single moving picture book. Specifically, the generated content is appropriately edited and compiled into a moving picture format.
[0666] Step 8: Output and display the animated picture book
[0667] The server sends the completed animated picture book to the terminal. The input is the integrated animated picture book file. The terminal receives this data and displays it on a digital device such as an iPad. The output is the animated picture book displayed on the digital device. Specifically, the terminal plays the data and allows the user to view the animated picture book.
[0668] (Application example 2)
[0669] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0670] Conventional animated picture book systems have difficulty personalizing content based on user emotions and preferences, resulting in generic content. Furthermore, they lack an interactive experience that utilizes the user's wearable device, leaving room for further improvement in user engagement. This invention aims to solve these problems by providing a personalization function that includes user emotion recognition and an experience that utilizes a wearable device.
[0671] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0672] In this invention, the server includes: means for a user to select a theme for the animated picture book; means for generating an original story based on the selected theme using a generative AI model; means for generating illustrations appropriate for the theme using the generative AI model; means for integrating the generated story and illustrations to create an animated picture book; means for selecting a voice preferred by the user; means for performing voice synthesis using the selected voice; means for integrating the synthesized voice into the animated picture book; means for outputting the animated picture book to a display device; means for recognizing emotions from the user's facial expressions and voice and generating a story and illustrations and adjusting the tone of the voice based on the emotions; and means for recognizing emotions and displaying the animated picture book using a wearable device such as smart glasses. This allows a personalized animated picture book to be generated based on the user's emotions and to be experienced interactively via the wearable device.
[0673] An "animated picture book" is a digital book that combines an original story and illustrations generated by a generative AI model based on a theme selected by the user, and displays them with synthesized audio.
[0674] A "means of selection" is a feature of an interface or application that allows a user to select a particular theme or voice.
[0675] A "generative AI model" is an artificial intelligence algorithm that automatically generates stories and illustrations based on given themes and conditions.
[0676] "Means for recognizing emotions" refers to technologies and systems for reading emotions from a user's facial expressions, voice, etc., and processing that emotional data.
[0677] "Voice synthesis" is a technology for generating specific voices and tones, and involves synthesizing sampled voices using deepfake technology, etc.
[0678] A "wearable device" is a digital device (e.g., smart glasses) that can be worn by the user and has functions such as displaying animated picture books and recognizing emotions.
[0679] A "display device" is a digital screen or projection device for displaying the generated animated picture book.
[0680] This invention is a system that generates animated picture books using a generative AI model by having the user select a theme, and displays them through a wearable device. Specific embodiments of each step are described below.
[0681] System Program
[0682] The server includes means for a user to select a theme for the animated picture book, means for generating an original story based on the selected theme using a generative AI model, means for generating illustrations suitable for the theme using a generative AI model, means for integrating the generated story and illustrations to create an animated picture book, means for selecting an audio preferred by the user, means for performing voice synthesis using the selected audio, means for integrating the synthesized audio into the animated picture book, means for outputting the animated picture book to a display device, means for recognizing emotions from the user's facial expressions and voice, and generating a story and illustrations and adjusting the tone of the voice based thereon, and means for recognizing emotions and displaying the animated picture book using a wearable device such as smart glasses.
[0683] Hardware and software used
[0684] Hardware: smart glasses, camera, microphone
[0685] Software: OpenCV (for facial and emotion recognition), generative AI models (e.g., GPT-4, DALL-E), speech synthesis libraries (e.g., Google TTS, Amazon Polly)
[0686] Processing Description
[0687] The server uses a generative AI model to generate an original story based on the theme and emotion data selected by the user. For example, if a user selects the theme "Space Adventure" and the emotion engine recognizes the user's excitement, the generative AI model (e.g., GPT-4) will generate a "Space Adventure" story that reflects the user's excitement.
[0688] Next, the server generates appropriate illustrations based on the generated story. Using a generative AI model (e.g., DALL-E), illustrations of scenes and characters within the story are created. For example, colorful illustrations of spaceships and aliens are generated.
[0689] Furthermore, the server generates a voice using a speech synthesis library based on the voice settings selected by the user (e.g., a parent's voice). This voice is adjusted in tone and emotional expression based on the emotion data. For example, a parent's voice generates a voice with a happy tone.
[0690] Finally, the server integrates the generated story, illustrations, and audio, and sends the animated picture book to the smart glasses, which then display the personalized animated picture book based on the emotion data on the user's front display.
[0691] Specific examples
[0692] The user selects the theme "Space Adventure," and the emotion engine recognizes the exciting emotion. Based on that data, a generative AI model (GPT-4) generates a unique story, and a generative AI model (DALL-E) generates illustrations. The user selects a parent's voice, and the speech synthesis library synthesizes the voice based on that voice. All elements are integrated to create a complete animated storybook, which is then sent to the smart glasses.
[0693] Prompt Sentence Examples
[0694] Theme: Space Adventure
[0695] User Emotion: Excitement
[0696] Generative Stories: Kids, Adventure, Educational
[0697] Story length: 5 minutes
[0698] Generated illustrations: Colorful, spaceship, alien
[0699] In this way, a highly personalized animated picture book is generated using a generative AI model, providing users with an interactive experience.
[0700] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0701] Step 1:
[0702] The user selects the theme of the animated picture book.
[0703] The input is the action of the user selecting a theme through the interface.
[0704] The terminal transmits the theme selected by the user to the server and provides the theme data to the server.
[0705] Step 2:
[0706] The server receives the themes sent by the user and inputs the theme data into the generative AI model.
[0707] The input is theme data, and prompt sentences for story generation are generated based on this.
[0708] The server generates a story using a generative AI model, which generates the necessary text data.
[0709] Step 3:
[0710] The terminal or server captures the user's facial expressions and voice and inputs them into an emotion recognition engine.
[0711] The input is the user's facial expression and voice data, based on which emotion data is generated.
[0712] The emotion recognition engine analyzes the data, recognizes the user's emotional state, and provides the data to the server.
[0713] Step 4:
[0714] The server uses the generated emotion data and theme data to input it back into the generative AI model and generate a prompt for generating an illustration.
[0715] The input is emotion data and thematic data, which are then used to generate prompts containing detailed instructions for generating illustrations.
[0716] The server uses a generative AI model to generate illustrations suited to the theme, generating image data.
[0717] Step 5:
[0718] Users select the audio settings they want to use.
[0719] The input is the audio settings selected by the user through the interface.
[0720] The terminal transmits the selected audio settings to the server and provides the server with the audio setting data.
[0721] Step 6:
[0722] The server generates the voice using a voice synthesis library based on the selected voice settings.
[0723] The inputs are the voice setting data and the generated story data, and voice data is generated based on this.
[0724] The server takes the synthesized voice data and adjusts its tone and emotional expression.
[0725] Step 7:
[0726] The server integrates the generated story, illustrations, and audio to complete the animated picture book.
[0727] The inputs are story data, illustration data, and audio data, and based on this, integrated video data is generated.
[0728] The server integrates these data and generates the completed animated picture book.
[0729] Step 8:
[0730] The terminal receives the completed animated picture book from the server and outputs it on a display device such as smart glasses.
[0731] The input is the synthesized video data, from which a video file in a format suitable for the display device is generated.
[0732] The terminal outputs the animated picture book to a display device, providing the user with a visually interactive experience.
[0733] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0734] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0735] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0736] [Third embodiment]
[0737] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0738] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0739] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0740] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0741] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0742] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0743] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0744] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0745] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0746] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0747] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0748] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0749] This invention is a system in which a generative AI model creates a story and illustrations based on a theme selected by the user, and then integrates them to generate a video picture book. This system allows the user to set the desired audio and integrate the synthesized audio into the video picture book, allowing the system to display an original video picture book for young children on a device.
[0750] 1. Select a theme for your video book
[0751] The user selects a theme for the animated picture book. For example, they can choose a theme that includes an educational message such as "dragon," "adventure," or "eat all your food." The device receives the selected theme and sends the theme information to the server.
[0752] 2. Creating original stories and illustrations
[0753] Once the server receives the theme information, it uses a generative AI model to create an original story. The story can include episodes and characters that fit the theme. The server also uses the generative AI model to generate colorful, child-friendly illustrations to accompany the story. These stories and illustrations form the basis of the animated picture book.
[0754] 3. Voice settings and synthesis
[0755] The user selects the voice their child prefers, such as a deepfake voice sampled from a parent's voice or the voice of an animated character. The device then sends the selected voice settings to the server, which uses speech synthesis technology to generate the voice that matches the video storybook in the selected audio format.
[0756] 4. Integration and display of animated picture books
[0757] The server integrates the generated story, illustrations, and synthesized voice to create a complete animated picture book. The terminal receives the completed animated picture book and displays it on a digital device such as an iPad. This allows parents to easily read original animated picture books to their children using their digital devices.
[0758] Specific examples
[0759] For example, if the user selects the "Dragon" theme and requests the "Parent Voice" audio setting, the process will proceed as follows:
[0760] The user selects the "Dragon" theme and the device sends the theme to the server.
[0761] The server uses a generative AI model to generate original stories and illustrations based on the theme of "dragons."
[0762] The user selects "Parent's Voice" as the audio setting, and the device sends the setting to the server.
[0763] The server synthesizes voice using the parent's voice and integrates the voice into the animated picture book.
[0764] The completed video picture book is sent to the terminal and displayed on a digital device such as an iPad.
[0765] In this way, the system of the present invention allows parents to easily create animated picture books containing original educational messages and read them to their young children, thereby reducing the burden on parents and supporting their children's growth and learning.
[0766] The processing flow will be explained below.
[0767] Step 1:
[0768] The user selects the "theme of the animated picture book." For example, they can choose themes such as "Dragons" or "Eat all your food."
[0769] Step 2:
[0770] The terminal receives the user's selected theme information and transmits the information to the server.
[0771] Step 3:
[0772] The server analyzes the received theme information.
[0773] Step 4:
[0774] Based on the selected theme, the server uses a generative AI model to generate an original story, which includes episodes and characters that fit the theme.
[0775] Step 5:
[0776] Similarly, the server uses generative AI models to generate illustrations that fit the story, with colorful, child-friendly designs.
[0777] Step 6:
[0778] The server integrates the generated story and illustrations to create an animated picture book.
[0779] Step 7:
[0780] Users can choose the voice their child prefers, such as a deepfake voice sampled from a parent's voice or the voice of an animated character.
[0781] Step 8:
[0782] The device sends the user's selected audio settings to the server.
[0783] Step 9:
[0784] The server synthesizes the voice based on the selected voice format, either using the sampled parent voice to generate the deepfake voice or using the specified character voice.
[0785] Step 10:
[0786] The server integrates the synthesized voice into the generated animated picture book.
[0787] Step 11:
[0788] The server transmits the completed animated picture book data to the terminal.
[0789] Step 12:
[0790] The terminal displays the received video picture book on a digital device (such as an iPad), allowing parents to read original video picture books to their children using the digital device.
[0791] Through the above steps, users can easily create original animated picture books containing educational messages and easily read them to their children.
[0792] Example 1
[0793] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0794] In conventional video picture book generation systems, the process of generating original stories and illustrations based on themes and audio selected by the user and integrating them into videos is complicated. Furthermore, the use of voice synthesis technology is limited, making it difficult to accommodate individual audio settings desired by users. Therefore, there is a need for a simple and efficient way to generate original video picture books containing educational messages.
[0795] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0796] In this invention, the server includes means for a user to select a theme for the animated picture book, means for transmitting the selected theme information to the server, means for generating an original story based on the selected theme using a generative AI model, means for generating illustrations appropriate for the theme using the generative AI model, means for transmitting the generated story and illustrations to a terminal, means for a user to select audio settings, means for transmitting audio setting information to the server, means for generating audio in a selected audio format using speech synthesis technology, means for integrating the generated audio with the story and illustrations to create the animated picture book, and means for outputting the animated picture book to a display device. This allows a user to simply select a theme and set the audio, automatically generating an original animated picture book and efficiently providing high-quality content including educational messages.
[0797] "User" refers to an individual or organization that uses the system to create animated picture books.
[0798] "Theme" refers to the basic concept or subject matter of the content of the animated picture book.
[0799] "Terminal" refers to a computer or digital device operated by a user.
[0800] "Server" refers to the central processing unit that receives thematic information and audio setting information and generates stories and illustrations using generative AI models.
[0801] A "generative AI model" refers to an algorithm that uses artificial intelligence technology to generate text and images.
[0802] A "prompt sentence" is text data that is input into a generative AI model to guide the generated content based on a specific theme or setting.
[0803] "Original story" refers to a unique narrative generated by a generative AI model based on a selected theme.
[0804] "Illustrations" refer to visual images generated by a generative AI model based on a story.
[0805] "Audio Settings" refers to the particular setting options a user selects regarding the audio for an animated picture book.
[0806] "Speech synthesis technology" refers to technology for converting text data into speech.
[0807] An "animated picture book" is a digital picture book that integrates an original story, illustrations, and synthesized audio.
[0808] "Display device" refers to a digital device for displaying the generated animated picture book.
[0809] This invention is a system in which a generative AI model creates a story and illustrations based on a theme selected by the user, and then integrates them to generate a video picture book. This system allows the user to set the desired audio and integrate the synthesized audio into the video picture book, allowing the system to display an original video picture book for young children on a device.
[0810] Hardware and software used
[0811] Device: A digital device operated by a user (e.g., tablet, smartphone)
[0812] Server: Central processing unit (e.g., cloud server)
[0813] Generative AI models: Natural language processing and image generation techniques (e.g., GPT-4, DALL-E 2)
[0814] Speech synthesis technology: Technology that converts text into speech (e.g., Google Text-to-Speech API)
[0815] How it works
[0816] 1. Choose a theme
[0817] The user operates the device to select a theme for the animated picture book from a list of themes provided, such as "Dragons" or "Eat all your food."
[0818] The terminal transmits the selected theme information to the server.
[0819] 2. Story and illustration generation
[0820] The server generates an original story based on the received theme information using a generative AI model (e.g., GPT-4), and also generates colorful, child-friendly illustrations that fit the story using a generative AI model (e.g., DALL-E 2).
[0821] 3. Voice settings and synthesis
[0822] The user selects the voice their child prefers on the device. For example, they can choose the voice of a parent or an anime character. The device then sends the selected voice setting information to the server. The server then uses speech synthesis technology (e.g., Google Text-to-Speech API) to generate speech in the selected voice format.
[0823] 4. Video picture book integration
[0824] The server completes the animated picture book by integrating the generated story, illustrations, and synthesized voice.
[0825] 5. Outputting a video picture book
[0826] The completed animated picture book data is sent from the server to the terminal and displayed on the user's device.
[0827] Specific examples
[0828] If the user selects the "Dragon" theme and requests the "Parent Voice" audio setting, the process will proceed as follows:
[0829] The user selects the "Dragon" theme and the device sends the theme to the server.
[0830] The server generates an original story based on the theme of "dragon" using a generative AI model (e.g., GPT-4) and generates illustrations using a generative AI model (e.g., DALL-E 2).
[0831] The user selects "Parent's Voice" as the audio setting, and the device sends the setting information to the server.
[0832] The server uses the parent's voice to synthesize the text and integrate it into the story and illustrations.
[0833] The completed animated picture book is sent to the terminal and displayed on the device.
[0834] Examples of prompt statements
[0835] "Create a dragon adventure story. The main character is a brave child who befriends a dragon. The illustrations should be colorful and child-friendly."
[0836] By inputting this prompt into a generative AI model, the system generates a story and illustrations based on the specified theme and setting.
[0837] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0838] Step 1:
[0839] The user selects the theme of the animated picture book.
[0840] The user operates the device to select a theme for the animated picture book (e.g., "Dragon") from a list of themes provided. The input data is the theme selected by the user, and the output data is this theme information. Specifically, the user taps the theme name using the device's touchscreen.
[0841] Step 2:
[0842] The terminal transmits the theme information to the server.
[0843] The terminal sends the selected theme information to the server as an HTTP POST request. The input data is the theme information, and the output data is the request sent to the server. In concrete terms, the terminal sends the data to the server via the network.
[0844] Step 3:
[0845] The server generates the original story.
[0846] The server generates an original story using a generative AI model (e.g., GPT-4) based on the received theme information. The input data is the theme information, and the output data is the generated story. Specifically, based on the "dragon theme," the server sends a prompt to the generative AI model and receives the generated story.
[0847] Step 4:
[0848] The server generates the illustration.
[0849] The server uses a generative AI model (e.g., DALL-E 2) to generate illustrations appropriate for the generated story. The input data is the generated story, and the output data is the generated illustration. Specifically, the server uses the generated story as a prompt, sends it to the generative AI model, and receives the illustration.
[0850] Step 5:
[0851] The server sends the generated story and illustrations to the device.
[0852] The server sends the generated story and illustration data to the terminal as an HTTP response. The input data is the generated story and illustration, and the output data is the request sent to the terminal. In concrete terms, the server sends the data to the terminal via the network.
[0853] Step 6:
[0854] The user selects an audio setting.
[0855] The user selects the desired voice (e.g., "Parent's Voice") from the voice setting options provided on the device. The input data is the user's voice setting selection, and the output data is the selected voice setting. Specifically, the user taps the voice option using the device's touchscreen.
[0856] Step 7:
[0857] The terminal transmits audio setting information to the server.
[0858] The terminal sends the audio setting information selected by the user to the server as an HTTP POST request. The input data is the audio setting information, and the output data is the request sent to the server. Specifically, the terminal sends the data to the server via the network.
[0859] Step 8:
[0860] The server generates the audio and integrates it into the story and illustrations.
[0861] The server uses speech synthesis technology (e.g., Google Text-to-Speech API) to generate audio in the selected audio format and integrates the generated audio into the story and illustrations. The input data is audio setting information and the generated story and illustrations, and the output data is a complete animated picture book with the integrated audio. Specifically, the server synthesizes audio using the "parent's voice" and integrates it into the pre-generated story and illustrations.
[0862] Step 9:
[0863] The server sends the completed animated picture book to the terminal.
[0864] The server sends the completed animated picture book data to the terminal as an HTTP response. The input data is the completed animated picture book, and the output data is the request sent to the terminal. In concrete terms, the server sends the data to the terminal via the network.
[0865] Step 10:
[0866] The device displays the animated picture book.
[0867] The terminal displays the received animated picture book on a display device (e.g., a tablet or smartphone). The input data is the completed animated picture book, and the output data is the displayed animated picture book. The specific operation is to launch the terminal's video playback application and play the animated picture book.
[0868] The above are the specific processing steps of the system.
[0869] (Application example 1)
[0870] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0871] While existing video picture book generation systems can easily generate original stories and illustrations based on a user-selected theme, integrating the generated content into a single video picture book and playing it on the user's device requires significant effort. Especially when applied to content distribution services, an effective approach is needed to ensure easy user access and use. Furthermore, it is difficult to integrate diverse functions, such as voice synthesis using deepfake technology and the selection of themes that include educational messages.
[0872] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0873] In this invention, the server includes: a means for a user to select a theme for the animated picture book; a means for generating an original story based on the selected theme using a generative AI model; a means for generating illustrations appropriate to the theme using the generative AI model; a means for integrating the generated story and illustrations to create an animated picture book; a means for selecting a user's preferred audio; a means for synthesizing the selected audio; a means for integrating the synthesized audio into the animated picture book; a means for outputting the animated picture book to a display device; a means for providing the content distribution service as an application for a smartphone or tablet; a means for inputting a prompt to the generative AI model; a means for integrating the generated story and illustrations into a single animated picture book; and a means for playing the animated picture book on a display device. This allows users to easily create original animated picture books and play them on digital devices. Furthermore, a multifunctional animated picture book generation system that can be used in a user-friendly environment can also be provided for content distribution services.
[0874] An animated picture book is a digital picture book that combines illustrations and a story created based on a theme selected by the user, and is played back along with audio.
[0875] A "generative AI model" is a type of artificial intelligence that automatically generates original stories and illustrations based on a given theme and prompt.
[0876] A "theme" is a user-selected subject or theme that forms the basis of the story and illustrations in the animated picture book.
[0877] "Speech synthesis" is a technology that creates narration and character voices corresponding to a generated story based on user-selected voice settings.
[0878] "Content distribution service" refers to a service that provides digital content to users via the Internet.
[0879] A "prompt" refers to a sentence or keyword that is input to a generative AI model to instruct it to produce a specific output.
[0880] A "display device" is a device (e.g., smartphone, tablet, smart TV) for visually playing the generated animated picture book.
[0881] "Deepfake voice" is a synthetic voice that is generated by sampling the voice of a specific individual and using artificial intelligence technology.
[0882] "User preferred voice" refers to the voice settings of narration and characters selected by the user for a specific purpose.
[0883] A system for implementing this invention is configured as follows: First, a user launches an application on a display device such as a smartphone, tablet, smart TV, etc. The user selects a theme for the animated picture book, and the theme information is sent from the terminal to a server.
[0884] The server uses a generative AI model (e.g., GPT-4, DALL-E) to generate an original story and illustrations based on the received theme information. This generative AI model generates appropriate output by inputting a prompt. For example, if the theme is "adventure," the following prompt is used:
[0885] Example story generation prompt: "The main character is a brave boy named Tarrou. He and his dragon friend explore a magical land. They..."
[0886] Example prompt for generating an illustration: "Draw a dragon flying through the sky, looking at the stars. Make sure the colors are colorful and child-friendly."
[0887] The generated story and illustrations are integrated on the server side to create a single animated picture book. The user selects their preferred audio, and the audio settings are sent from the device to the server. Using speech synthesis technology (e.g., Neural TTS), audio is generated in the selected audio format, and this audio is integrated into the animated picture book.
[0888] Finally, the completed animated picture book is sent to the user's device and played through the application, allowing users to easily create original animated picture books and view them on the spot.
[0889] The system uses display devices such as smartphones, tablets, and smart TVs, as well as cloud services that utilize a serverless architecture (e.g., AWS Lambda and Google Cloud Functions). Data is sent and received using API requests and responses in JSON format. Specific generative AI models used include GPT-4, DALL-E, and Stable Diffusion.
[0890] For example, if the user selects the "Dragon" theme and requests speech synthesis using the parent's voice, the following process occurs:
[0891] 1. The user selects the "Dragon" theme and the device sends the theme to the server.
[0892] 2. The server generates original stories and illustrations using a generative AI model based on the theme of "dragons."
[0893] 3. The user selects "Parent's Voice" as the audio setting, and the device sends the setting to the server.
[0894] 4. The server uses the parent's voice to synthesize the audio and integrate it into the animated picture book.
[0895] 5. The completed animated picture book is sent to the terminal and played on the display device through the application.
[0896] In this way, users can easily create original animated picture books and provide them to children as educational content.
[0897] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0898] Step 1:
[0899] The user selects a theme for the animated picture book. The user launches the application and selects one from the list of themes provided. The theme name selected by the user is input here, and the device sends this theme information to the server. The request data including the theme information is output.
[0900] Step 2:
[0901] The server receives the theme information. The server receives the theme information (e.g., "Adventure" or "Dragon") sent from the device and generates a prompt sentence to be input to the generative AI model based on this. The input is the theme information, and the output is the prompt sentence to be passed to the generative AI model.
[0902] Step 3:
[0903] The server generates a story using a generative AI model. Based on the generated prompt, it asks the generative AI model (e.g., GPT-4) to generate a story. The input is the prompt, and the output is the original story. For example, a story like this might be generated: "The main character is a brave boy named Taro. Taro and his friend, the dragon, go on an adventure in a magical land. They..."
[0904] Step 4:
[0905] The server generates illustrations using a generative AI model. Based on the generated story, a prompt is input into the illustration generation AI model (e.g., DALL-E, Stable Diffusion) to request the generation of an illustration. The input is a prompt related to the story text, and the output is an original illustration image. For example, a prompt might be, "Draw a dragon flying through the sky, looking at the shining stars in the night sky. Please make it colorful and suitable for children."
[0906] Step 5:
[0907] The server integrates the generated story and illustrations. The server interactively integrates the story and illustrations and compiles them into a single animated picture book. The input is the generated story text and illustration images, and the output is the constituent data of the animated picture book.
[0908] Step 6:
[0909] The user selects the voice they prefer. The user selects one of the voice options provided within the application (e.g., parent's voice, anime character's voice), and the setting information is sent from the device to the server. The input is the voice setting information selected by the user, and the output is the request data to the server.
[0910] Step 7:
[0911] The server generates the voice using speech synthesis technology. Based on the selected voice settings, it generates the story narration using speech synthesis technology (e.g., Neural TTS). The input is the voice setting information and the story text, and the output is the synthesized voice data.
[0912] Step 8:
[0913] The server integrates the synthesized audio into the animated picture book. The generated audio is added to the existing animated picture book's constituent data, and is finally integrated into an animated picture book with audio. The input is the synthesized audio data and the animated picture book's constituent data, and the output is the completed animated picture book data.
[0914] Step 9:
[0915] The server sends the completed animated picture book to the terminal. The completed animated picture book data is sent to the terminal so that it can be played in the application. The completed animated picture book data is input, and the result of data transfer to the terminal is output. The terminal plays the received animated picture book on a display device.
[0916] Through the above processing steps, users can easily create original animated picture books and view them on the spot.
[0917] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0918] This invention is a system that generates animated picture books by using a generative AI model to create and integrate stories and illustrations based on a theme selected by the user. This system allows users to set their desired audio and integrate the synthesized audio into the animated picture book. In addition, by combining it with an emotion engine that recognizes the user's emotions, it is possible to provide more personalized animated picture books.
[0919] 1. Select a theme for your video book
[0920] The user selects the "animated picture book theme." For example, they can choose themes such as "Dragons" or "Eat all your food." The device receives the selected theme and sends the theme information to the server.
[0921] 2. Emotion Recognition by Emotion Engine
[0922] An emotion engine installed on the device or server recognizes the user's emotions from their facial expressions and voice. This emotion data is used for processing in later steps.
[0923] 3. Creating original stories and illustrations
[0924] Once the server receives the theme information and emotion data, it uses a generative AI model to create an original story. Based on the emotion data, the character's actions and dialogue within the story are adjusted. Similarly, the generative AI model is used to generate illustrations that fit the theme and emotion, resulting in facial expressions and scenes that correspond to the emotion.
[0925] 4. Voice settings and synthesis
[0926] The user selects the voice their child prefers. For example, they can choose a deepfake voice sampled from a parent's voice or the voice of an animated character. The device then sends the selected voice settings to the server. The server then uses voice synthesis technology to generate a voice that matches the video book in the selected audio format. The tone and emotional expression of the voice are then adjusted based on the emotional data.
[0927] 5. Integration and display of animated picture books
[0928] The server integrates the generated story, illustrations, and synthesized voice to create a complete animated picture book. The terminal receives the completed animated picture book and displays it on a digital device such as an iPad. This allows users to easily read original animated picture books to their children using their digital devices.
[0929] Specific examples
[0930] For example, if the user selects the "Dragon" theme and requests the "Parent Voice" audio setting, the process will proceed as follows:
[0931] The user selects the "Dragon" theme and the device sends the theme to the server.
[0932] The emotion engine on the terminal or server recognizes the user's emotion and transmits the data to the server.
[0933] Based on the theme of "dragon," the server uses a generative AI model to generate original stories and illustrations, taking emotional data into consideration.
[0934] The user selects "Parent's Voice" as the audio setting, and the device sends the setting to the server.
[0935] The server synthesizes voice using the parent's voice, adjusts the tone and emotional expression of the voice based on the emotional data, and integrates the voice into the animated picture book.
[0936] The completed video picture book is sent to the terminal and displayed on a digital device such as an iPad.
[0937] In this way, the system of the present invention allows parents to easily create animated picture books containing original educational messages tailored to their children's emotions and interests, enabling them to read aloud to their young children more effectively, thereby reducing the burden on parents and providing individual support for their children's growth and learning.
[0938] The processing flow will be explained below.
[0939] Step 1:
[0940] The user selects the "theme of the animated picture book." For example, they can choose themes such as "Dragons" or "Eat all your food."
[0941] Step 2:
[0942] The terminal receives the user's selected theme information and transmits the information to the server.
[0943] Step 3:
[0944] The emotion engine uses the device's camera and microphone to recognize the user's emotions, collecting and analyzing the user's facial expressions and voice data to extract emotional data.
[0945] Step 4:
[0946] The device transmits the collected emotion data to a server.
[0947] Step 5:
[0948] The server analyzes the received theme information and emotion data.
[0949] Step 6:
[0950] The server generates an original story based on the theme information using a generative AI model, adjusting the characters' actions and dialogue based on the emotional data.
[0951] Step 7:
[0952] The server then uses a generative AI model to generate illustrations that fit the story, adjusting the character's facial expressions and the mood of the scene based on the emotional data.
[0953] Step 8:
[0954] The server integrates the generated story and illustrations to create an animated picture book.
[0955] Step 9:
[0956] Users can choose the voice their child prefers, such as a deepfake voice sampled from a parent's voice or the voice of an animated character.
[0957] Step 10:
[0958] The device sends the user's selected audio settings to the server.
[0959] Step 11:
[0960] The server synthesizes speech based on the selected voice format, adjusting the tone and emotional expression of the speech based on the emotional data.
[0961] Step 12:
[0962] The server integrates the synthesized audio into the generated animated picture book.
[0963] Step 13:
[0964] The server transmits the completed animated picture book data to the terminal.
[0965] Step 14:
[0966] The video picture book received by the terminal is displayed on a digital device (such as an iPad), allowing users to easily read original video picture books to their children using their digital device.
[0967] As a specific example, when the user selects "dragon" as the theme and "parent's voice" as the voice, the processing flow is as follows.
[0968] The user selects the "Dragon" theme and the terminal transmits the theme information to the server.
[0969] The device's emotion engine analyzes the user's facial expressions and voice and sends the emotion data to the server.
[0970] The server uses a generative AI model to generate original stories and illustrations based on thematic information and emotional data.
[0971] The user selects "Parent's Voice" as the audio setting, and the device sends the setting information to the server.
[0972] The server generates synthetic speech using the parent's voice and integrates the speech, which reflects emotional data, into the animated picture book.
[0973] The completed animated picture book is sent to the terminal and displayed on the user's digital device.
[0974] This allows parents to create original animated picture books that are personalized according to their emotions, enabling them to read to their children effectively.
[0975] Example 2
[0976] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0977] Conventional video picture book creation systems have had problems with personalization based on the user's emotions and interests, and lack of natural emotional expression in their voice synthesis. Furthermore, generating original stories and illustrations based on a theme and integrating them is not easy, placing a heavy burden on the user. Therefore, there is a need for a system that can recognize the user's emotions and easily create more personalized original video picture books.
[0978] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0979] In this invention, the server includes means for analyzing the user's facial expressions and voice to acquire emotional data, means for generating an original story based on the theme information and emotional data using a generative AI model, and means for adjusting the tone of the voice and emotional expression based on the emotional data using voice synthesis technology. This makes it possible to generate a story and illustrations that reflect the user's emotions and to perform voice synthesis that incorporates natural emotional expressions.
[0980] A "user" is an entity that uses the system to select the theme and audio settings for an animated picture book.
[0981] A "terminal" is a digital device that is operated by a user and has the role of transmitting theme information and audio setting information to a server.
[0982] The "server" is a central device that receives thematic information and emotional data selected by the user, and generates original stories and illustrations and synthesizes voices based on that information.
[0983] A "generative AI model" is an artificial intelligence algorithm that generates original stories and illustrations based on thematic information and emotional data.
[0984] "Theme information" is data indicating the content selected by the user as the subject of the animated picture book.
[0985] An "emotion engine" is a combination of software and hardware that analyzes a user's facial expressions and voice to recognize emotional data.
[0986] "Emotion data" is data that indicates the emotional state of the user as recognized by the emotion engine.
[0987] A "story" is a narrative plot that a generative AI model generates based on thematic information and emotional data.
[0988] "Illustrations" are visual images associated with a story that are generated by a generative AI model.
[0989] "Speech synthesis" is the process of using speech synthesis technology to generate narration or character voices in a user-defined voice format.
[0990] "Voice setting information" is data indicating the voice synthesis settings selected by the user.
[0991] An animated picture book is a digital picture book created by integrating generated stories, illustrations, and audio.
[0992] This invention is a system that generates animated picture books by using a generative AI model to create and integrate stories and illustrations based on a theme selected by the user. This system allows users to set their desired audio and integrate the synthesized audio into the animated picture book. In addition, by combining it with an emotion engine that recognizes the user's emotions, it is possible to provide more personalized animated picture books.
[0993] Specifically, the user first uses the device to select a theme for the animated picture book, such as "Dragons" or "Eat all your food." The device then sends the selected theme information to the server. The server is equipped with an emotion engine that analyzes the user's facial expressions and voice to obtain emotion data. This emotion data is then used for subsequent processing.
[0994] The server inputs the received theme information and emotional data into a generative AI model. The generative AI model generates an original story based on this information. Furthermore, the behavior and dialogue of the characters in the story are adjusted based on the emotional data. Similarly, illustrations appropriate for the theme and emotion are generated. This allows expressions and scenes to be depicted according to the emotion.
[0995] Next, the user uses the device to select the desired voice, such as "parent's voice" or "animated character's voice." The device then sends the selected voice setting to the server. The server then uses voice synthesis technology to generate voice that matches the selected voice format for the animated picture book. At this time, the tone and emotional expression of the voice are adjusted based on the emotional data.
[0996] Finally, the server integrates the generated story, illustrations, and synthesized voice to create a complete animated picture book. The completed animated picture book is sent to the terminal and displayed on a digital device such as an iPad. This allows users to easily read original animated picture books to their children using their digital devices.
[0997] For example, if the user selects the "Dragon" theme and uses "Parent Voice", the following happens:
[0998] 1. The user selects the "Dragon" theme, and the device sends the theme information to the server.
[0999] 2. The server's emotion engine analyzes the user's facial expressions and voice to obtain emotional data such as "it looks fun."
[1000] 3. The server inputs the theme information "dragon" and the emotion data "looks fun" into the AI model to generate an original story and illustrations.
[1001] 4. The user selects "Parent's voice" and the device sends the information to the server.
[1002] 5. The server uses speech synthesis technology to generate audio in the parent's voice and adjusts the tone and emotional expression based on the emotional data.
[1003] 6. The server combines the story, illustrations, and audio to complete the animated picture book and sends it to the device.
[1004] 7. The device will display the completed video picture book on a digital device such as an iPad.
[1005] An example prompt is, "The theme is dragons. Use your parent's voice to make the book character happy based on the emotion data."
[1006] In this way, the system of the present invention makes it possible to easily create original animated picture books that are in line with the user's emotions and interests, and can also include educational messages, thereby reducing the burden on the user and providing individual support for children's growth and learning.
[1007] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1008] Step 1: Select a theme
[1009] The user operates the device to select a theme for the animated picture book. The input is the theme selection information. Using this information as input, the device sends the theme information to the server. Specifically, the user selects a theme such as "dragon," and the device records this selection and sends it to the server. The output is the selected theme information.
[1010] Step 2: Emotion Recognition
[1011] The server's emotion engine analyzes the user's facial expressions and voice to obtain emotion data. The input is the user's facial expressions and voice information. The server analyzes this and outputs emotion data such as "happy" or "sad." Specifically, the emotion engine analyzes data captured by a camera or microphone.
[1012] Step 3: Generate an original story
[1013] The server uses a generative AI model to generate an original story based on thematic information and emotional data. The input is thematic information and emotional data. Based on this, the generative AI model generates a story, which is then output. Specifically, it generates a "fun story about a dragon's adventure."
[1014] Step 4: Creating an illustration
[1015] The server uses the same generative AI model to generate suitable illustrations based on thematic information and emotional data. The input is thematic information and emotional data. The output is the generated illustration. For example, an illustration of a "smiling dragon" is generated. Here, too, the generative AI model creates the illustration based on visual elements.
[1016] Step 5: Select your audio settings
[1017] The user selects the desired voice using the terminal. The input is the voice setting information selected by the user. The terminal sends this information to the server, and the voice setting information is output. Specifically, the user selects "parent's voice."
[1018] Step 6: Text-to-Speech
[1019] The server uses voice synthesis technology based on the voice setting information to generate voice in the voice format set by the user. The input is the voice setting information and emotional data. The server generates voice based on this, and outputs voice with the tone and emotional expression adjusted based on the emotional data. For example, a voice with a "parent's voice in a happy tone" may be generated.
[1020] Step 7: Integrating the video book
[1021] The server integrates the generated story, illustrations, and audio to create a moving picture book. The inputs are the story, illustrations, and audio. These are integrated and output as a single moving picture book. Specifically, the generated content is appropriately edited and compiled into a moving picture format.
[1022] Step 8: Output and display the animated picture book
[1023] The server sends the completed animated picture book to the terminal. The input is the integrated animated picture book file. The terminal receives this data and displays it on a digital device such as an iPad. The output is the animated picture book displayed on the digital device. Specifically, the terminal plays the data and allows the user to view the animated picture book.
[1024] (Application example 2)
[1025] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1026] Conventional animated picture book systems have difficulty personalizing content based on user emotions and preferences, resulting in generic content. Furthermore, they lack an interactive experience that utilizes the user's wearable device, leaving room for further improvement in user engagement. This invention aims to solve these problems by providing a personalization function that includes user emotion recognition and an experience that utilizes a wearable device.
[1027] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1028] In this invention, the server includes: means for a user to select a theme for the animated picture book; means for generating an original story based on the selected theme using a generative AI model; means for generating illustrations appropriate for the theme using the generative AI model; means for integrating the generated story and illustrations to create an animated picture book; means for selecting a voice preferred by the user; means for performing voice synthesis using the selected voice; means for integrating the synthesized voice into the animated picture book; means for outputting the animated picture book to a display device; means for recognizing emotions from the user's facial expressions and voice and generating a story and illustrations and adjusting the tone of the voice based on the emotions; and means for recognizing emotions and displaying the animated picture book using a wearable device such as smart glasses. This allows a personalized animated picture book to be generated based on the user's emotions and to be experienced interactively via the wearable device.
[1029] An "animated picture book" is a digital book that combines an original story and illustrations generated by a generative AI model based on a theme selected by the user, and displays them with synthesized audio.
[1030] A "means of selection" is a feature of an interface or application that allows a user to select a particular theme or voice.
[1031] A "generative AI model" is an artificial intelligence algorithm that automatically generates stories and illustrations based on given themes and conditions.
[1032] "Means for recognizing emotions" refers to technologies and systems for reading emotions from a user's facial expressions, voice, etc., and processing that emotional data.
[1033] "Voice synthesis" is a technology for generating specific voices and tones, and involves synthesizing sampled voices using deepfake technology, etc.
[1034] A "wearable device" is a digital device (e.g., smart glasses) that can be worn by the user and has functions such as displaying animated picture books and recognizing emotions.
[1035] A "display device" is a digital screen or projection device for displaying the generated animated picture book.
[1036] This invention is a system that generates animated picture books using a generative AI model by having the user select a theme, and displays them through a wearable device. Specific embodiments of each step are described below.
[1037] System Program
[1038] The server includes means for a user to select a theme for the animated picture book, means for generating an original story based on the selected theme using a generative AI model, means for generating illustrations suitable for the theme using a generative AI model, means for integrating the generated story and illustrations to create an animated picture book, means for selecting an audio preferred by the user, means for performing voice synthesis using the selected audio, means for integrating the synthesized audio into the animated picture book, means for outputting the animated picture book to a display device, means for recognizing emotions from the user's facial expressions and voice, and generating a story and illustrations and adjusting the tone of the voice based thereon, and means for recognizing emotions and displaying the animated picture book using a wearable device such as smart glasses.
[1039] Hardware and software used
[1040] Hardware: smart glasses, camera, microphone
[1041] Software: OpenCV (for facial and emotion recognition), generative AI models (e.g., GPT-4, DALL-E), speech synthesis libraries (e.g., Google TTS, Amazon Polly)
[1042] Processing Description
[1043] The server uses a generative AI model to generate an original story based on the theme and emotion data selected by the user. For example, if a user selects the theme "Space Adventure" and the emotion engine recognizes the user's excitement, the generative AI model (e.g., GPT-4) will generate a "Space Adventure" story that reflects the user's excitement.
[1044] Next, the server generates appropriate illustrations based on the generated story. Using a generative AI model (e.g., DALL-E), illustrations of scenes and characters within the story are created. For example, colorful illustrations of spaceships and aliens are generated.
[1045] Furthermore, the server generates a voice using a speech synthesis library based on the voice settings selected by the user (e.g., a parent's voice). This voice is adjusted in tone and emotional expression based on the emotion data. For example, a parent's voice generates a voice with a happy tone.
[1046] Finally, the server integrates the generated story, illustrations, and audio, and sends the animated picture book to the smart glasses, which then display the personalized animated picture book based on the emotion data on the user's front display.
[1047] Specific examples
[1048] The user selects the theme "Space Adventure," and the emotion engine recognizes the exciting emotion. Based on that data, a generative AI model (GPT-4) generates a unique story, and a generative AI model (DALL-E) generates illustrations. The user selects a parent's voice, and the speech synthesis library synthesizes the voice based on that voice. All elements are integrated to create a complete animated storybook, which is then sent to the smart glasses.
[1049] Prompt Sentence Examples
[1050] Theme: Space Adventure
[1051] User Emotion: Excitement
[1052] Generative Stories: Kids, Adventure, Educational
[1053] Story length: 5 minutes
[1054] Generated illustrations: Colorful, spaceship, alien
[1055] In this way, a highly personalized animated picture book is generated using a generative AI model, providing users with an interactive experience.
[1056] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1057] Step 1:
[1058] The user selects the theme of the animated picture book.
[1059] The input is the action of the user selecting a theme through the interface.
[1060] The terminal transmits the theme selected by the user to the server and provides the theme data to the server.
[1061] Step 2:
[1062] The server receives the themes sent by the user and inputs the theme data into the generative AI model.
[1063] The input is theme data, and prompt sentences for story generation are generated based on this.
[1064] The server generates a story using a generative AI model, which generates the necessary text data.
[1065] Step 3:
[1066] The terminal or server captures the user's facial expressions and voice and inputs them into an emotion recognition engine.
[1067] The input is the user's facial expression and voice data, based on which emotion data is generated.
[1068] The emotion recognition engine analyzes the data, recognizes the user's emotional state, and provides the data to the server.
[1069] Step 4:
[1070] The server uses the generated emotion data and theme data to input it back into the generative AI model and generate a prompt for generating an illustration.
[1071] The input is emotion data and thematic data, which are then used to generate prompts containing detailed instructions for generating illustrations.
[1072] The server uses a generative AI model to generate illustrations suited to the theme, generating image data.
[1073] Step 5:
[1074] Users select the audio settings they want to use.
[1075] The input is the audio settings selected by the user through the interface.
[1076] The terminal transmits the selected audio settings to the server and provides the server with the audio setting data.
[1077] Step 6:
[1078] The server generates the voice using a voice synthesis library based on the selected voice settings.
[1079] The inputs are the voice setting data and the generated story data, and voice data is generated based on this.
[1080] The server takes the synthesized voice data and adjusts its tone and emotional expression.
[1081] Step 7:
[1082] The server integrates the generated story, illustrations, and audio to complete the animated picture book.
[1083] The inputs are story data, illustration data, and audio data, and based on this, integrated video data is generated.
[1084] The server integrates these data and generates the completed animated picture book.
[1085] Step 8:
[1086] The terminal receives the completed animated picture book from the server and outputs it on a display device such as smart glasses.
[1087] The input is the synthesized video data, from which a video file in a format suitable for the display device is generated.
[1088] The terminal outputs the animated picture book to a display device, providing the user with a visually interactive experience.
[1089] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1090] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1091] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1092] [Fourth embodiment]
[1093] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1094] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1095] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1096] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1097] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1098] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1099] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1100] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1101] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1102] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1103] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1104] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1105] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1106] This invention is a system in which a generative AI model creates a story and illustrations based on a theme selected by the user, and then integrates them to generate a video picture book. This system allows the user to set the desired audio and integrate the synthesized audio into the video picture book, allowing the system to display an original video picture book for young children on a device.
[1107] 1. Select a theme for your video book
[1108] The user selects a theme for the animated picture book. For example, they can choose a theme that includes an educational message such as "dragon," "adventure," or "eat all your food." The device receives the selected theme and sends the theme information to the server.
[1109] 2. Creating original stories and illustrations
[1110] Once the server receives the theme information, it uses a generative AI model to create an original story. The story can include episodes and characters that fit the theme. The server also uses the generative AI model to generate colorful, child-friendly illustrations to accompany the story. These stories and illustrations form the basis of the animated picture book.
[1111] 3. Voice settings and synthesis
[1112] The user selects the voice their child prefers, such as a deepfake voice sampled from a parent's voice or the voice of an animated character. The device then sends the selected voice settings to the server, which uses speech synthesis technology to generate the voice that matches the video storybook in the selected audio format.
[1113] 4. Integration and display of animated picture books
[1114] The server integrates the generated story, illustrations, and synthesized voice to create a complete animated picture book. The terminal receives the completed animated picture book and displays it on a digital device such as an iPad. This allows parents to easily read original animated picture books to their children using their digital devices.
[1115] Specific examples
[1116] For example, if the user selects the "Dragon" theme and requests the "Parent Voice" audio setting, the process will proceed as follows:
[1117] The user selects the "Dragon" theme and the device sends the theme to the server.
[1118] The server uses a generative AI model to generate original stories and illustrations based on the theme of "dragons."
[1119] The user selects "Parent's Voice" as the audio setting, and the device sends the setting to the server.
[1120] The server synthesizes voice using the parent's voice and integrates the voice into the animated picture book.
[1121] The completed video picture book is sent to the terminal and displayed on a digital device such as an iPad.
[1122] In this way, the system of the present invention allows parents to easily create animated picture books containing original educational messages and read them to their young children, thereby reducing the burden on parents and supporting their children's growth and learning.
[1123] The processing flow will be explained below.
[1124] Step 1:
[1125] The user selects the "theme of the animated picture book." For example, they can choose themes such as "Dragons" or "Eat all your food."
[1126] Step 2:
[1127] The terminal receives the user's selected theme information and transmits the information to the server.
[1128] Step 3:
[1129] The server analyzes the received theme information.
[1130] Step 4:
[1131] Based on the selected theme, the server uses a generative AI model to generate an original story, which includes episodes and characters that fit the theme.
[1132] Step 5:
[1133] Similarly, the server uses generative AI models to generate illustrations that fit the story, with colorful, child-friendly designs.
[1134] Step 6:
[1135] The server integrates the generated story and illustrations to create an animated picture book.
[1136] Step 7:
[1137] Users can choose the voice their child prefers, such as a deepfake voice sampled from a parent's voice or the voice of an animated character.
[1138] Step 8:
[1139] The device sends the user's selected audio settings to the server.
[1140] Step 9:
[1141] The server synthesizes the voice based on the selected voice format, either using the sampled parent voice to generate the deepfake voice or using the specified character voice.
[1142] Step 10:
[1143] The server integrates the synthesized voice into the generated animated picture book.
[1144] Step 11:
[1145] The server transmits the completed animated picture book data to the terminal.
[1146] Step 12:
[1147] The terminal displays the received video picture book on a digital device (such as an iPad), allowing parents to read original video picture books to their children using the digital device.
[1148] Through the above steps, users can easily create original animated picture books containing educational messages and easily read them to their children.
[1149] Example 1
[1150] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1151] In conventional video picture book generation systems, the process of generating original stories and illustrations based on themes and audio selected by the user and integrating them into videos is complicated. Furthermore, the use of voice synthesis technology is limited, making it difficult to accommodate individual audio settings desired by users. Therefore, there is a need for a simple and efficient way to generate original video picture books containing educational messages.
[1152] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1153] In this invention, the server includes means for a user to select a theme for the animated picture book, means for transmitting the selected theme information to the server, means for generating an original story based on the selected theme using a generative AI model, means for generating illustrations appropriate for the theme using the generative AI model, means for transmitting the generated story and illustrations to a terminal, means for a user to select audio settings, means for transmitting audio setting information to the server, means for generating audio in a selected audio format using speech synthesis technology, means for integrating the generated audio with the story and illustrations to create the animated picture book, and means for outputting the animated picture book to a display device. This allows a user to simply select a theme and set the audio, automatically generating an original animated picture book and efficiently providing high-quality content including educational messages.
[1154] "User" refers to an individual or organization that uses the system to create animated picture books.
[1155] "Theme" refers to the basic concept or subject matter of the content of the animated picture book.
[1156] "Terminal" refers to a computer or digital device operated by a user.
[1157] "Server" refers to the central processing unit that receives thematic information and audio setting information and generates stories and illustrations using generative AI models.
[1158] A "generative AI model" refers to an algorithm that uses artificial intelligence technology to generate text and images.
[1159] A "prompt sentence" is text data that is input into a generative AI model to guide the generated content based on a specific theme or setting.
[1160] "Original story" refers to a unique narrative generated by a generative AI model based on a selected theme.
[1161] "Illustrations" refer to visual images generated by a generative AI model based on a story.
[1162] "Audio Settings" refers to the particular setting options a user selects regarding the audio for an animated picture book.
[1163] "Speech synthesis technology" refers to technology for converting text data into speech.
[1164] An "animated picture book" is a digital picture book that integrates an original story, illustrations, and synthesized audio.
[1165] "Display device" refers to a digital device for displaying the generated animated picture book.
[1166] This invention is a system in which a generative AI model creates a story and illustrations based on a theme selected by the user, and then integrates them to generate a video picture book. This system allows the user to set the desired audio and integrate the synthesized audio into the video picture book, allowing the system to display an original video picture book for young children on a device.
[1167] Hardware and software used
[1168] Device: A digital device operated by a user (e.g., tablet, smartphone)
[1169] Server: Central processing unit (e.g., cloud server)
[1170] Generative AI models: Natural language processing and image generation techniques (e.g., GPT-4, DALL-E 2)
[1171] Speech synthesis technology: Technology that converts text into speech (e.g., Google Text-to-Speech API)
[1172] How it works
[1173] 1. Choose a theme
[1174] The user operates the device to select a theme for the animated picture book from a list of themes provided, such as "Dragons" or "Eat all your food."
[1175] The terminal transmits the selected theme information to the server.
[1176] 2. Story and illustration generation
[1177] The server generates an original story based on the received theme information using a generative AI model (e.g., GPT-4), and also generates colorful, child-friendly illustrations that fit the story using a generative AI model (e.g., DALL-E 2).
[1178] 3. Voice settings and synthesis
[1179] The user selects the voice their child prefers on the device. For example, they can choose the voice of a parent or an anime character. The device then sends the selected voice setting information to the server. The server then uses speech synthesis technology (e.g., Google Text-to-Speech API) to generate speech in the selected voice format.
[1180] 4. Video picture book integration
[1181] The server completes the animated picture book by integrating the generated story, illustrations, and synthesized voice.
[1182] 5. Outputting a video picture book
[1183] The completed animated picture book data is sent from the server to the terminal and displayed on the user's device.
[1184] Specific examples
[1185] If the user selects the "Dragon" theme and requests the "Parent Voice" audio setting, the process will proceed as follows:
[1186] The user selects the "Dragon" theme and the device sends the theme to the server.
[1187] The server generates an original story based on the theme of "dragon" using a generative AI model (e.g., GPT-4) and generates illustrations using a generative AI model (e.g., DALL-E 2).
[1188] The user selects "Parent's Voice" as the audio setting, and the device sends the setting information to the server.
[1189] The server uses the parent's voice to synthesize the text and integrate it into the story and illustrations.
[1190] The completed animated picture book is sent to the terminal and displayed on the device.
[1191] Examples of prompt statements
[1192] "Create a dragon adventure story. The main character is a brave child who befriends a dragon. The illustrations should be colorful and child-friendly."
[1193] By inputting this prompt into a generative AI model, the system generates a story and illustrations based on the specified theme and setting.
[1194] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1195] Step 1:
[1196] The user selects the theme of the animated picture book.
[1197] The user operates the device to select a theme for the animated picture book (e.g., "Dragon") from a list of themes provided. The input data is the theme selected by the user, and the output data is this theme information. Specifically, the user taps the theme name using the device's touchscreen.
[1198] Step 2:
[1199] The terminal transmits the theme information to the server.
[1200] The terminal sends the selected theme information to the server as an HTTP POST request. The input data is the theme information, and the output data is the request sent to the server. In concrete terms, the terminal sends the data to the server via the network.
[1201] Step 3:
[1202] The server generates the original story.
[1203] The server generates an original story using a generative AI model (e.g., GPT-4) based on the received theme information. The input data is the theme information, and the output data is the generated story. Specifically, based on the "dragon theme," the server sends a prompt to the generative AI model and receives the generated story.
[1204] Step 4:
[1205] The server generates the illustration.
[1206] The server uses a generative AI model (e.g., DALL-E 2) to generate illustrations appropriate for the generated story. The input data is the generated story, and the output data is the generated illustration. Specifically, the server uses the generated story as a prompt, sends it to the generative AI model, and receives the illustration.
[1207] Step 5:
[1208] The server sends the generated story and illustrations to the device.
[1209] The server sends the generated story and illustration data to the terminal as an HTTP response. The input data is the generated story and illustration, and the output data is the request sent to the terminal. In concrete terms, the server sends the data to the terminal via the network.
[1210] Step 6:
[1211] The user selects an audio setting.
[1212] The user selects the desired voice (e.g., "Parent's Voice") from the voice setting options provided on the device. The input data is the user's voice setting selection, and the output data is the selected voice setting. Specifically, the user taps the voice option using the device's touchscreen.
[1213] Step 7:
[1214] The terminal transmits audio setting information to the server.
[1215] The terminal sends the audio setting information selected by the user to the server as an HTTP POST request. The input data is the audio setting information, and the output data is the request sent to the server. Specifically, the terminal sends the data to the server via the network.
[1216] Step 8:
[1217] The server generates the audio and integrates it into the story and illustrations.
[1218] The server uses speech synthesis technology (e.g., Google Text-to-Speech API) to generate audio in the selected audio format and integrates the generated audio into the story and illustrations. The input data is audio setting information and the generated story and illustrations, and the output data is a complete animated picture book with the integrated audio. Specifically, the server synthesizes audio using the "parent's voice" and integrates it into the pre-generated story and illustrations.
[1219] Step 9:
[1220] The server sends the completed animated picture book to the terminal.
[1221] The server sends the completed animated picture book data to the terminal as an HTTP response. The input data is the completed animated picture book, and the output data is the request sent to the terminal. In concrete terms, the server sends the data to the terminal via the network.
[1222] Step 10:
[1223] The device displays the animated picture book.
[1224] The terminal displays the received animated picture book on a display device (e.g., a tablet or smartphone). The input data is the completed animated picture book, and the output data is the displayed animated picture book. The specific operation is to launch the terminal's video playback application and play the animated picture book.
[1225] The above are the specific processing steps of the system.
[1226] (Application example 1)
[1227] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1228] While existing video picture book generation systems can easily generate original stories and illustrations based on a user-selected theme, integrating the generated content into a single video picture book and playing it on the user's device requires significant effort. Especially when applied to content distribution services, an effective approach is needed to ensure easy user access and use. Furthermore, it is difficult to integrate diverse functions, such as voice synthesis using deepfake technology and the selection of themes that include educational messages.
[1229] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1230] In this invention, the server includes: a means for a user to select a theme for the animated picture book; a means for generating an original story based on the selected theme using a generative AI model; a means for generating illustrations appropriate to the theme using the generative AI model; a means for integrating the generated story and illustrations to create an animated picture book; a means for selecting a user's preferred audio; a means for synthesizing the selected audio; a means for integrating the synthesized audio into the animated picture book; a means for outputting the animated picture book to a display device; a means for providing the content distribution service as an application for a smartphone or tablet; a means for inputting a prompt to the generative AI model; a means for integrating the generated story and illustrations into a single animated picture book; and a means for playing the animated picture book on a display device. This allows users to easily create original animated picture books and play them on digital devices. Furthermore, a multifunctional animated picture book generation system that can be used in a user-friendly environment can also be provided for content distribution services.
[1231] An animated picture book is a digital picture book that combines illustrations and a story created based on a theme selected by the user, and is played back along with audio.
[1232] A "generative AI model" is a type of artificial intelligence that automatically generates original stories and illustrations based on a given theme and prompt.
[1233] A "theme" is a user-selected subject or theme that forms the basis of the story and illustrations in the animated picture book.
[1234] "Speech synthesis" is a technology that creates narration and character voices corresponding to a generated story based on user-selected voice settings.
[1235] "Content distribution service" refers to a service that provides digital content to users via the Internet.
[1236] A "prompt" refers to a sentence or keyword that is input to a generative AI model to instruct it to produce a specific output.
[1237] A "display device" is a device (e.g., smartphone, tablet, smart TV) for visually playing the generated animated picture book.
[1238] "Deepfake voice" is a synthetic voice that is generated by sampling the voice of a specific individual and using artificial intelligence technology.
[1239] "User preferred voice" refers to the voice settings of narration and characters selected by the user for a specific purpose.
[1240] A system for implementing this invention is configured as follows: First, a user launches an application on a display device such as a smartphone, tablet, smart TV, etc. The user selects a theme for the animated picture book, and the theme information is sent from the terminal to a server.
[1241] The server uses a generative AI model (e.g., GPT-4, DALL-E) to generate an original story and illustrations based on the received theme information. This generative AI model generates appropriate output by inputting a prompt. For example, if the theme is "adventure," the following prompt is used:
[1242] Example story generation prompt: "The main character is a brave boy named Tarrou. He and his dragon friend explore a magical land. They..."
[1243] Example prompt for generating an illustration: "Draw a dragon flying through the sky, looking at the stars. Make sure the colors are colorful and child-friendly."
[1244] The generated story and illustrations are integrated on the server side to create a single animated picture book. The user selects their preferred audio, and the audio settings are sent from the device to the server. Using speech synthesis technology (e.g., Neural TTS), audio is generated in the selected audio format, and this audio is integrated into the animated picture book.
[1245] Finally, the completed animated picture book is sent to the user's device and played through the application, allowing users to easily create original animated picture books and view them on the spot.
[1246] The system uses display devices such as smartphones, tablets, and smart TVs, as well as cloud services that utilize a serverless architecture (e.g., AWS Lambda and Google Cloud Functions). Data is sent and received using API requests and responses in JSON format. Specific generative AI models used include GPT-4, DALL-E, and Stable Diffusion.
[1247] For example, if the user selects the "Dragon" theme and requests speech synthesis using the parent's voice, the following process occurs:
[1248] 1. The user selects the "Dragon" theme and the device sends the theme to the server.
[1249] 2. The server generates original stories and illustrations using a generative AI model based on the theme of "dragons."
[1250] 3. The user selects "Parent's Voice" as the audio setting, and the device sends the setting to the server.
[1251] 4. The server uses the parent's voice to synthesize the audio and integrate it into the animated picture book.
[1252] 5. The completed animated picture book is sent to the terminal and played on the display device through the application.
[1253] In this way, users can easily create original animated picture books and provide them to children as educational content.
[1254] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1255] Step 1:
[1256] The user selects a theme for the animated picture book. The user launches the application and selects one from the list of themes provided. The theme name selected by the user is input here, and the device sends this theme information to the server. The request data including the theme information is output.
[1257] Step 2:
[1258] The server receives the theme information. The server receives the theme information (e.g., "Adventure" or "Dragon") sent from the device and generates a prompt sentence to be input to the generative AI model based on this. The input is the theme information, and the output is the prompt sentence to be passed to the generative AI model.
[1259] Step 3:
[1260] The server generates a story using a generative AI model. Based on the generated prompt, it asks the generative AI model (e.g., GPT-4) to generate a story. The input is the prompt, and the output is the original story. For example, a story like this might be generated: "The main character is a brave boy named Taro. Taro and his friend, the dragon, go on an adventure in a magical land. They..."
[1261] Step 4:
[1262] The server generates illustrations using a generative AI model. Based on the generated story, a prompt is input into the illustration generation AI model (e.g., DALL-E, Stable Diffusion) to request the generation of an illustration. The input is a prompt related to the story text, and the output is an original illustration image. For example, a prompt might be, "Draw a dragon flying through the sky, looking at the shining stars in the night sky. Please make it colorful and suitable for children."
[1263] Step 5:
[1264] The server integrates the generated story and illustrations. The server interactively integrates the story and illustrations and compiles them into a single animated picture book. The input is the generated story text and illustration images, and the output is the constituent data of the animated picture book.
[1265] Step 6:
[1266] The user selects the voice they prefer. The user selects one of the voice options provided within the application (e.g., parent's voice, anime character's voice), and the setting information is sent from the device to the server. The input is the voice setting information selected by the user, and the output is the request data to the server.
[1267] Step 7:
[1268] The server generates the voice using speech synthesis technology. Based on the selected voice settings, it generates the story narration using speech synthesis technology (e.g., Neural TTS). The input is the voice setting information and the story text, and the output is the synthesized voice data.
[1269] Step 8:
[1270] The server integrates the synthesized audio into the animated picture book. The generated audio is added to the existing animated picture book's constituent data, and is finally integrated into an animated picture book with audio. The input is the synthesized audio data and the animated picture book's constituent data, and the output is the completed animated picture book data.
[1271] Step 9:
[1272] The server sends the completed animated picture book to the terminal. The completed animated picture book data is sent to the terminal so that it can be played in the application. The completed animated picture book data is input, and the result of data transfer to the terminal is output. The terminal plays the received animated picture book on a display device.
[1273] Through the above processing steps, users can easily create original animated picture books and view them on the spot.
[1274] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1275] This invention is a system that generates animated picture books by using a generative AI model to create and integrate stories and illustrations based on a theme selected by the user. This system allows users to set their desired audio and integrate the synthesized audio into the animated picture book. In addition, by combining it with an emotion engine that recognizes the user's emotions, it is possible to provide more personalized animated picture books.
[1276] 1. Select a theme for your video book
[1277] The user selects the "animated picture book theme." For example, they can choose themes such as "Dragons" or "Eat all your food." The device receives the selected theme and sends the theme information to the server.
[1278] 2. Emotion Recognition by Emotion Engine
[1279] An emotion engine installed on the device or server recognizes the user's emotions from their facial expressions and voice. This emotion data is used for processing in later steps.
[1280] 3. Creating original stories and illustrations
[1281] Once the server receives the theme information and emotion data, it uses a generative AI model to create an original story. Based on the emotion data, the character's actions and dialogue within the story are adjusted. Similarly, the generative AI model is used to generate illustrations that fit the theme and emotion, resulting in facial expressions and scenes that correspond to the emotion.
[1282] 4. Voice settings and synthesis
[1283] The user selects the voice their child prefers. For example, they can choose a deepfake voice sampled from a parent's voice or the voice of an animated character. The device then sends the selected voice settings to the server. The server then uses voice synthesis technology to generate a voice that matches the video book in the selected audio format. The tone and emotional expression of the voice are then adjusted based on the emotional data.
[1284] 5. Integration and display of animated picture books
[1285] The server integrates the generated story, illustrations, and synthesized voice to create a complete animated picture book. The terminal receives the completed animated picture book and displays it on a digital device such as an iPad. This allows users to easily read original animated picture books to their children using their digital devices.
[1286] Specific examples
[1287] For example, if the user selects the "Dragon" theme and requests the "Parent Voice" audio setting, the process will proceed as follows:
[1288] The user selects the "Dragon" theme and the device sends the theme to the server.
[1289] The emotion engine on the terminal or server recognizes the user's emotion and transmits the data to the server.
[1290] Based on the theme of "dragon," the server uses a generative AI model to generate original stories and illustrations, taking emotional data into consideration.
[1291] The user selects "Parent's Voice" as the audio setting, and the device sends the setting to the server.
[1292] The server synthesizes voice using the parent's voice, adjusts the tone and emotional expression of the voice based on the emotional data, and integrates the voice into the animated picture book.
[1293] The completed video picture book is sent to the terminal and displayed on a digital device such as an iPad.
[1294] In this way, the system of the present invention allows parents to easily create animated picture books containing original educational messages tailored to their children's emotions and interests, enabling them to read aloud to their young children more effectively, thereby reducing the burden on parents and providing individual support for their children's growth and learning.
[1295] The processing flow will be explained below.
[1296] Step 1:
[1297] The user selects the "theme of the animated picture book." For example, they can choose themes such as "Dragons" or "Eat all your food."
[1298] Step 2:
[1299] The terminal receives the user's selected theme information and transmits the information to the server.
[1300] Step 3:
[1301] The emotion engine uses the device's camera and microphone to recognize the user's emotions, collecting and analyzing the user's facial expressions and voice data to extract emotional data.
[1302] Step 4:
[1303] The device transmits the collected emotion data to a server.
[1304] Step 5:
[1305] The server analyzes the received theme information and emotion data.
[1306] Step 6:
[1307] The server generates an original story based on the theme information using a generative AI model, adjusting the characters' actions and dialogue based on the emotional data.
[1308] Step 7:
[1309] The server then uses a generative AI model to generate illustrations that fit the story, adjusting the character's facial expressions and the mood of the scene based on the emotional data.
[1310] Step 8:
[1311] The server integrates the generated story and illustrations to create an animated picture book.
[1312] Step 9:
[1313] Users can choose the voice their child prefers, such as a deepfake voice sampled from a parent's voice or the voice of an animated character.
[1314] Step 10:
[1315] The device sends the user's selected audio settings to the server.
[1316] Step 11:
[1317] The server synthesizes speech based on the selected voice format, adjusting the tone and emotional expression of the speech based on the emotional data.
[1318] Step 12:
[1319] The server integrates the synthesized audio into the generated animated picture book.
[1320] Step 13:
[1321] The server transmits the completed animated picture book data to the terminal.
[1322] Step 14:
[1323] The video picture book received by the terminal is displayed on a digital device (such as an iPad), allowing users to easily read original video picture books to their children using their digital device.
[1324] As a specific example, when the user selects "dragon" as the theme and "parent's voice" as the voice, the processing flow is as follows.
[1325] The user selects the "Dragon" theme and the terminal transmits the theme information to the server.
[1326] The device's emotion engine analyzes the user's facial expressions and voice and sends the emotion data to the server.
[1327] The server uses a generative AI model to generate original stories and illustrations based on thematic information and emotional data.
[1328] The user selects "Parent's Voice" as the audio setting, and the device sends the setting information to the server.
[1329] The server generates synthetic speech using the parent's voice and integrates the speech, which reflects emotional data, into the animated picture book.
[1330] The completed animated picture book is sent to the terminal and displayed on the user's digital device.
[1331] This allows parents to create original animated picture books that are personalized according to their emotions, enabling them to read to their children effectively.
[1332] Example 2
[1333] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1334] Conventional video picture book creation systems have had problems with personalization based on the user's emotions and interests, and lack of natural emotional expression in their voice synthesis. Furthermore, generating original stories and illustrations based on a theme and integrating them is not easy, placing a heavy burden on the user. Therefore, there is a need for a system that can recognize the user's emotions and easily create more personalized original video picture books.
[1335] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1336] In this invention, the server includes means for analyzing the user's facial expressions and voice to acquire emotional data, means for generating an original story based on the theme information and emotional data using a generative AI model, and means for adjusting the tone of the voice and emotional expression based on the emotional data using voice synthesis technology. This makes it possible to generate a story and illustrations that reflect the user's emotions and to perform voice synthesis that incorporates natural emotional expressions.
[1337] A "user" is an entity that uses the system to select the theme and audio settings for an animated picture book.
[1338] A "terminal" is a digital device that is operated by a user and has the role of transmitting theme information and audio setting information to a server.
[1339] The "server" is a central device that receives thematic information and emotional data selected by the user, and generates original stories and illustrations and synthesizes voices based on that information.
[1340] A "generative AI model" is an artificial intelligence algorithm that generates original stories and illustrations based on thematic information and emotional data.
[1341] "Theme information" is data indicating the content selected by the user as the subject of the animated picture book.
[1342] An "emotion engine" is a combination of software and hardware that analyzes a user's facial expressions and voice to recognize emotional data.
[1343] "Emotion data" is data that indicates the emotional state of the user as recognized by the emotion engine.
[1344] A "story" is a narrative plot that a generative AI model generates based on thematic information and emotional data.
[1345] "Illustrations" are visual images associated with a story that are generated by a generative AI model.
[1346] "Speech synthesis" is the process of using speech synthesis technology to generate narration or character voices in a user-defined voice format.
[1347] "Voice setting information" is data indicating the voice synthesis settings selected by the user.
[1348] An animated picture book is a digital picture book created by integrating generated stories, illustrations, and audio.
[1349] This invention is a system that generates animated picture books by using a generative AI model to create and integrate stories and illustrations based on a theme selected by the user. This system allows users to set their desired audio and integrate the synthesized audio into the animated picture book. In addition, by combining it with an emotion engine that recognizes the user's emotions, it is possible to provide more personalized animated picture books.
[1350] Specifically, the user first uses the device to select a theme for the animated picture book, such as "Dragons" or "Eat all your food." The device then sends the selected theme information to the server. The server is equipped with an emotion engine that analyzes the user's facial expressions and voice to obtain emotion data. This emotion data is then used for subsequent processing.
[1351] The server inputs the received theme information and emotional data into a generative AI model. The generative AI model generates an original story based on this information. Furthermore, the behavior and dialogue of the characters in the story are adjusted based on the emotional data. Similarly, illustrations appropriate for the theme and emotion are generated. This allows expressions and scenes to be depicted according to the emotion.
[1352] Next, the user uses the device to select the desired voice, such as "parent's voice" or "animated character's voice." The device then sends the selected voice setting to the server. The server then uses voice synthesis technology to generate voice that matches the selected voice format for the animated picture book. At this time, the tone and emotional expression of the voice are adjusted based on the emotional data.
[1353] Finally, the server integrates the generated story, illustrations, and synthesized voice to create a complete animated picture book. The completed animated picture book is sent to the terminal and displayed on a digital device such as an iPad. This allows users to easily read original animated picture books to their children using their digital devices.
[1354] For example, if the user selects the "Dragon" theme and uses "Parent Voice", the following happens:
[1355] 1. The user selects the "Dragon" theme, and the device sends the theme information to the server.
[1356] 2. The server's emotion engine analyzes the user's facial expressions and voice to obtain emotional data such as "it looks fun."
[1357] 3. The server inputs the theme information "dragon" and the emotion data "looks fun" into the AI model to generate an original story and illustrations.
[1358] 4. The user selects "Parent's voice" and the device sends the information to the server.
[1359] 5. The server uses speech synthesis technology to generate audio in the parent's voice and adjusts the tone and emotional expression based on the emotional data.
[1360] 6. The server combines the story, illustrations, and audio to complete the animated picture book and sends it to the device.
[1361] 7. The device will display the completed video picture book on a digital device such as an iPad.
[1362] An example prompt is, "The theme is dragons. Use your parent's voice to make the book character happy based on the emotion data."
[1363] In this way, the system of the present invention makes it possible to easily create original animated picture books that are in line with the user's emotions and interests, and can also include educational messages, thereby reducing the burden on the user and providing individual support for children's growth and learning.
[1364] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1365] Step 1: Select a theme
[1366] The user operates the device to select a theme for the animated picture book. The input is the theme selection information. Using this information as input, the device sends the theme information to the server. Specifically, the user selects a theme such as "dragon," and the device records this selection and sends it to the server. The output is the selected theme information.
[1367] Step 2: Emotion Recognition
[1368] The server's emotion engine analyzes the user's facial expressions and voice to obtain emotion data. The input is the user's facial expressions and voice information. The server analyzes this and outputs emotion data such as "happy" or "sad." Specifically, the emotion engine analyzes data captured by a camera or microphone.
[1369] Step 3: Generate an original story
[1370] The server uses a generative AI model to generate an original story based on thematic information and emotional data. The input is thematic information and emotional data. Based on this, the generative AI model generates a story, which is then output. Specifically, it generates a "fun story about a dragon's adventure."
[1371] Step 4: Creating an illustration
[1372] The server uses the same generative AI model to generate suitable illustrations based on thematic information and emotional data. The input is thematic information and emotional data. The output is the generated illustration. For example, an illustration of a "smiling dragon" is generated. Here, too, the generative AI model creates the illustration based on visual elements.
[1373] Step 5: Select your audio settings
[1374] The user selects the desired voice using the terminal. The input is the voice setting information selected by the user. The terminal sends this information to the server, and the voice setting information is output. Specifically, the user selects "parent's voice."
[1375] Step 6: Text-to-Speech
[1376] The server uses voice synthesis technology based on the voice setting information to generate voice in the voice format set by the user. The input is the voice setting information and emotional data. The server generates voice based on this, and outputs voice with the tone and emotional expression adjusted based on the emotional data. For example, a voice with a "parent's voice in a happy tone" may be generated.
[1377] Step 7: Integrating the video book
[1378] The server integrates the generated story, illustrations, and audio to create a moving picture book. The inputs are the story, illustrations, and audio. These are integrated and output as a single moving picture book. Specifically, the generated content is appropriately edited and compiled into a moving picture format.
[1379] Step 8: Output and display the animated picture book
[1380] The server sends the completed animated picture book to the terminal. The input is the integrated animated picture book file. The terminal receives this data and displays it on a digital device such as an iPad. The output is the animated picture book displayed on the digital device. Specifically, the terminal plays the data and allows the user to view the animated picture book.
[1381] (Application example 2)
[1382] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1383] Conventional animated picture book systems have difficulty personalizing content based on user emotions and preferences, resulting in generic content. Furthermore, they lack an interactive experience that utilizes the user's wearable device, leaving room for further improvement in user engagement. This invention aims to solve these problems by providing a personalization function that includes user emotion recognition and an experience that utilizes a wearable device.
[1384] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1385] In this invention, the server includes: means for a user to select a theme for the animated picture book; means for generating an original story based on the selected theme using a generative AI model; means for generating illustrations appropriate for the theme using the generative AI model; means for integrating the generated story and illustrations to create an animated picture book; means for selecting a voice preferred by the user; means for performing voice synthesis using the selected voice; means for integrating the synthesized voice into the animated picture book; means for outputting the animated picture book to a display device; means for recognizing emotions from the user's facial expressions and voice and generating a story and illustrations and adjusting the tone of the voice based on the emotions; and means for recognizing emotions and displaying the animated picture book using a wearable device such as smart glasses. This allows a personalized animated picture book to be generated based on the user's emotions and to be experienced interactively via the wearable device.
[1386] An "animated picture book" is a digital book that combines an original story and illustrations generated by a generative AI model based on a theme selected by the user, and displays them with synthesized audio.
[1387] A "means of selection" is a feature of an interface or application that allows a user to select a particular theme or voice.
[1388] A "generative AI model" is an artificial intelligence algorithm that automatically generates stories and illustrations based on given themes and conditions.
[1389] "Means for recognizing emotions" refers to technologies and systems for reading emotions from a user's facial expressions, voice, etc., and processing that emotional data.
[1390] "Voice synthesis" is a technology for generating specific voices and tones, and involves synthesizing sampled voices using deepfake technology, etc.
[1391] A "wearable device" is a digital device (e.g., smart glasses) that can be worn by the user and has functions such as displaying animated picture books and recognizing emotions.
[1392] A "display device" is a digital screen or projection device for displaying the generated animated picture book.
[1393] This invention is a system that generates animated picture books using a generative AI model by having the user select a theme, and displays them through a wearable device. Specific embodiments of each step are described below.
[1394] System Program
[1395] The server includes means for a user to select a theme for the animated picture book, means for generating an original story based on the selected theme using a generative AI model, means for generating illustrations suitable for the theme using a generative AI model, means for integrating the generated story and illustrations to create an animated picture book, means for selecting an audio preferred by the user, means for performing voice synthesis using the selected audio, means for integrating the synthesized audio into the animated picture book, means for outputting the animated picture book to a display device, means for recognizing emotions from the user's facial expressions and voice, and generating a story and illustrations and adjusting the tone of the voice based thereon, and means for recognizing emotions and displaying the animated picture book using a wearable device such as smart glasses.
[1396] Hardware and software used
[1397] Hardware: smart glasses, camera, microphone
[1398] Software: OpenCV (for facial and emotion recognition), generative AI models (e.g., GPT-4, DALL-E), speech synthesis libraries (e.g., Google TTS, Amazon Polly)
[1399] Processing Description
[1400] The server uses a generative AI model to generate an original story based on the theme and emotion data selected by the user. For example, if a user selects the theme "Space Adventure" and the emotion engine recognizes the user's excitement, the generative AI model (e.g., GPT-4) will generate a "Space Adventure" story that reflects the user's excitement.
[1401] Next, the server generates appropriate illustrations based on the generated story. Using a generative AI model (e.g., DALL-E), illustrations of scenes and characters within the story are created. For example, colorful illustrations of spaceships and aliens are generated.
[1402] Furthermore, the server generates a voice using a speech synthesis library based on the voice settings selected by the user (e.g., a parent's voice). This voice is adjusted in tone and emotional expression based on the emotion data. For example, a parent's voice generates a voice with a happy tone.
[1403] Finally, the server integrates the generated story, illustrations, and audio, and sends the animated picture book to the smart glasses, which then display the personalized animated picture book based on the emotion data on the user's front display.
[1404] Specific examples
[1405] The user selects the theme "Space Adventure," and the emotion engine recognizes the exciting emotion. Based on that data, a generative AI model (GPT-4) generates a unique story, and a generative AI model (DALL-E) generates illustrations. The user selects a parent's voice, and the speech synthesis library synthesizes the voice based on that voice. All elements are integrated to create a complete animated storybook, which is then sent to the smart glasses.
[1406] Prompt Sentence Examples
[1407] Theme: Space Adventure
[1408] User Emotion: Excitement
[1409] Generative Stories: Kids, Adventure, Educational
[1410] Story length: 5 minutes
[1411] Generated illustrations: Colorful, spaceship, alien
[1412] In this way, a highly personalized animated picture book is generated using a generative AI model, providing users with an interactive experience.
[1413] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1414] Step 1:
[1415] The user selects the theme of the animated picture book.
[1416] The input is the action of the user selecting a theme through the interface.
[1417] The terminal transmits the theme selected by the user to the server and provides the theme data to the server.
[1418] Step 2:
[1419] The server receives the themes sent by the user and inputs the theme data into the generative AI model.
[1420] The input is theme data, and prompt sentences for story generation are generated based on this.
[1421] The server generates a story using a generative AI model, which generates the necessary text data.
[1422] Step 3:
[1423] The terminal or server captures the user's facial expressions and voice and inputs them into an emotion recognition engine.
[1424] The input is the user's facial expression and voice data, based on which emotion data is generated.
[1425] The emotion recognition engine analyzes the data, recognizes the user's emotional state, and provides the data to the server.
[1426] Step 4:
[1427] The server uses the generated emotion data and theme data to input it back into the generative AI model and generate a prompt for generating an illustration.
[1428] The input is emotion data and thematic data, which are then used to generate prompts containing detailed instructions for generating illustrations.
[1429] The server uses a generative AI model to generate illustrations suited to the theme, generating image data.
[1430] Step 5:
[1431] Users select the audio settings they want to use.
[1432] The input is the audio settings selected by the user through the interface.
[1433] The terminal transmits the selected audio settings to the server and provides the server with the audio setting data.
[1434] Step 6:
[1435] The server generates the voice using a voice synthesis library based on the selected voice settings.
[1436] The inputs are the voice setting data and the generated story data, and voice data is generated based on this.
[1437] The server takes the synthesized voice data and adjusts its tone and emotional expression.
[1438] Step 7:
[1439] The server integrates the generated story, illustrations, and audio to complete the animated picture book.
[1440] The inputs are story data, illustration data, and audio data, and based on this, integrated video data is generated.
[1441] The server integrates these data and generates the completed animated picture book.
[1442] Step 8:
[1443] The terminal receives the completed animated picture book from the server and outputs it on a display device such as smart glasses.
[1444] The input is the synthesized video data, from which a video file in a format suitable for the display device is generated.
[1445] The terminal outputs the animated picture book to a display device, providing the user with a visually interactive experience.
[1446] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1447] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1448] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1449] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1450] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1451] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1452] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1453] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1454] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1455] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1456] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1457] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1458] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1459] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1460] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1461] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1462] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1463] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1464] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1465] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1466] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1467] The following is further disclosed regarding the above embodiment.
[1468] (Claim 1)
[1469] A means for a user to select a theme for the animated picture book;
[1470] a means for generating original stories based on selected themes using a generative AI model; and
[1471] A means to generate illustrations appropriate for a theme using generative AI models;
[1472] A method to integrate the generated story and illustrations to create a video picture book,
[1473] means for selecting a user's preferred voice;
[1474] means for performing speech synthesis using the selected voice;
[1475] A means of integrating the synthesized audio into an animated picture book;
[1476] A means of outputting animated picture books to a display device
[1477] A system including:
[1478] (Claim 2)
[1479] 10. The system of claim 1, further comprising means for generating deepfake audio using sampled parental audio.
[1480] (Claim 3)
[1481] 10. The system of claim 1, further comprising means for selecting a theme for the animated storybook that includes an educational message.
[1482] "Example 1"
[1483] (Claim 1)
[1484] A means for a user to select a theme for the animated picture book;
[1485] means for transmitting selected theme information to a server;
[1486] a means for generating original stories based on selected themes using a generative AI model;
[1487] A means to generate illustrations appropriate for a theme using generative AI models;
[1488] A means for transmitting the generated story and illustrations to a terminal;
[1489] a means for a user to select audio settings;
[1490] means for transmitting audio setting information to a server;
[1491] means for generating speech in a selected speech format using speech synthesis technology;
[1492] A means to integrate the generated audio with stories and illustrations to create animated picture books;
[1493] A means of outputting animated picture books to a display device
[1494] A system including:
[1495] (Claim 2)
[1496] The system of claim 1, further comprising means for inputting a prompt sentence into the generative AI model to generate a story and illustrations.
[1497] (Claim 3)
[1498] 10. The system of claim 1, further comprising means for selecting a theme containing an educational message for the generated animated storybook.
[1499] "Application Example 1"
[1500] (Claim 1)
[1501] A means for a user to select a theme for the animated picture book;
[1502] a means for generating original stories based on selected themes using a generative AI model; and
[1503] A means to generate illustrations appropriate for a theme using generative AI models;
[1504] A method to integrate the generated story and illustrations to create a video picture book,
[1505] means for selecting a user's preferred voice;
[1506] means for performing speech synthesis using the selected voice;
[1507] A means of integrating the synthesized audio into an animated picture book;
[1508] A means for outputting the animated picture book to a display device;
[1509] A means for providing the content distribution service as an application for a smartphone or tablet;
[1510] a means for inputting a prompt to the generative AI model;
[1511] A way to integrate the generated story and illustrations into a single animated picture book,
[1512] A system including means for playing an animated picture book on a display device.
[1513] (Claim 2)
[1514] 10. The system of claim 1, further comprising means for generating deepfake audio using sampled parental audio.
[1515] (Claim 3)
[1516] 10. The system of claim 1, further comprising means for selecting a theme for the animated storybook that includes an educational message.
[1517] "Example 2: Combining Emotion Engines"
[1518] (Claim 1)
[1519] A means for a user to select a theme for the animated picture book;
[1520] means for transmitting selected theme information to a server;
[1521] A means for the server to analyze the user's facial expressions and voice to acquire emotional data;
[1522] A means of generating original stories based on thematic information and sentiment data using a generative AI model;
[1523] A means of generating suitable illustrations based on thematic information and emotional data using a generative AI model;
[1524] A method for creating animated picture books by integrating the generated stories and illustrations,
[1525] means for transmitting audio setting information to a server;
[1526] a means for the server to generate a voice using a voice synthesis technology and adjust the tone and emotional expression of the voice based on the emotional data;
[1527] A means of integrating the synthesized audio into an animated picture book;
[1528] A means of outputting animated picture books to a display device
[1529] A system including:
[1530] (Claim 2)
[1531] 10. The system of claim 1, wherein the server includes means for generating deepfake audio using sampled parental audio.
[1532] (Claim 3)
[1533] 10. The system of claim 1, further comprising means for selecting a theme for the animated storybook that includes an educational message.
[1534] "Application example 2 when combining emotion engines"
[1535] (Claim 1)
[1536] A means for a user to select a theme for the animated picture book;
[1537] a means for generating original stories based on selected themes using a generative AI model; and
[1538] A means to generate illustrations appropriate for a theme using generative AI models;
[1539] A method to integrate the generated story and illustrations to create a video picture book,
[1540] means for selecting a user's preferred voice;
[1541] means for performing speech synthesis using the selected voice;
[1542] A means of integrating the synthesized audio into an animated picture book;
[1543] A means for outputting the animated picture book to a display device;
[1544] A means for recognizing emotions from a user's facial expressions and voice, and generating stories and illustrations based on the emotions, and adjusting the tone of the voice;
[1545] A method for using wearable devices such as smart glasses to recognize emotions and display animated picture books
[1546] A system including:
[1547] (Claim 2)
[1548] 10. The system of claim 1, further comprising means for generating deepfake audio using sampled parental audio.
[1549] (Claim 3)
[1550] 10. The system of claim 1, further comprising means for selecting a theme for the animated storybook that includes an educational message. [Explanation of symbols]
[1551] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for a user to select a theme for the animated picture book; a means for generating original stories based on selected themes using a generative AI model; and A means to generate illustrations appropriate for a theme using generative AI models; A method to integrate the generated story and illustrations to create a video picture book, means for selecting a user's preferred voice; means for performing speech synthesis using the selected voice; A means of integrating the synthesized audio into an animated picture book; A means of outputting animated picture books to a display device A system including:
2. The system of claim 1 , further comprising means for generating deepfake audio using sampled parental audio.
3. 10. The system of claim 1, further comprising means for selecting a theme containing an educational message for the animated storybook.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A