System

The system addresses the limitations of existing educational resources by generating and sharing educational stories and images in multiple languages, promoting parent-child communication and reducing educational disparities.

JP2026019852APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024121600
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing educational resources lack the ability to easily create safe and educational stories based on user-defined themes and characters, fail to unleash children's creative potential, and do not promote parent-child communication, while also being inadequate in providing multilingual support.

Method used

A system that includes means for receiving user input, filtering inappropriate content, generating stories and images, translating into multiple languages, and sharing the content with others, utilizing generative AI models and translation engines to create coherent and educational digital picture books.

Benefits of technology

Enables the generation of safe, educational, and creative stories that promote parent-child communication and reduce educational disparities by allowing users to set themes and characters, automatically generating stories and images, and sharing them in multiple languages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019852000001_ABST
    Figure 2026019852000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for receiving input data; means for filtering the received input data; means for generating a story based on the filtered data; means for generating an image based on the generated story; means for translating the generated story into a plurality of languages; means for distributing the translated story to a terminal; and means for sharing the story distributed to the terminal with other users.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In recent years, the widening educational gap and the influence of extremist ideology on social media have become social issues. Against this backdrop, there is a need to provide educational resources that are suitable for children of diverse cultural backgrounds while maintaining the neutrality and integrity of education. However, existing educational resources have limited means for parents and educators to easily create safe and educational stories based on themes and characters they desire. There is also a lack of tools to unleash children's creative potential and promote parent-child communication. [Means for solving the problem]

[0005] In order to solve the above problem, the present invention provides a system including a means for receiving input data and filtering the received input data, a means for generating a story based on the filtered data, a means for generating an image based on the generated story, a means for translating the generated story into multiple languages, a means for delivering the translated story to a terminal, and a means for sharing the story delivered to the terminal with other users.

[0006] Specifically, we provide a system as described in claim 1, which includes means for accepting themes and characters set by a user and automatically generating a story based on the accepted themes and characters. We also provide a system as described in claim 1, which includes means for simultaneously generating a story and images and integrating them, and means for evaluating the generated stories and images and ensuring their quality. This makes it possible to provide children with safe, educational, and creative stories, which is expected to improve the quality of education and promote communication between parents and children.

[0007] "Input data" refers to data that a user provides to the system, including information such as themes and character settings.

[0008] "Filtering" is the process of inspecting input data to see if it contains any extreme content or inappropriate language, and then excluding it.

[0009] A "story" is a story generated based on a set theme and characters, and is expressed in text format.

[0010] "Images" refer to visual illustrations or pictures that correspond to the generated story and serve to visually support the content of the story.

[0011] "Translation" is the process of converting the generated story into a different language.

[0012] A "terminal" is a device that allows a user to operate the system, and includes information processing devices such as smartphones, tablets, and personal computers.

[0013] "Sharing" refers to sharing the created story or image with other users using a communication means.

[0014] A "theme" is the main theme or central concept of a story, and is set by the user.

[0015] "Characters" are characters, animals, and other subjects that appear in the story, and are specifically defined by the user.

[0016] A "generative AI model" is an artificial intelligence algorithm that automatically generates stories or images based on input data.

[0017] "Quality assurance" is the process of verifying that the content of the generated stories and images is appropriate and guaranteeing their quality. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram illustrating a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] This system allows users to create safe and educational stories by setting themes and characters, and then translates and shares them in multiple languages. The program and processing of this system are described in detail below. The system mainly works in conjunction with three elements: the server, the terminal, and the user.

[0040] Description of system programs and processes

[0041] User Theme and Character Settings

[0042] 1. A user accesses the system through a terminal and inputs a theme (e.g., "friendship") and a character (e.g., "Rio the rabbit") using the input interface.

[0043] 2. The device sends the entered theme and character settings to the server.

[0044] Filtering input data

[0045] 1. The server analyzes and filters the theme and character settings received from the device, specifically by screening for prohibited words and explicit content.

[0046] 2. If the server determines that the input data is appropriate based on the filtering results, it proceeds to the next step using that data. If the data contains inappropriate data, it sends a message to the terminal requesting the user to correct it.

[0047] Narrative Generation

[0048] 1. The server calls a generative AI model based on the verified theme and character settings to generate a story. For example, it generates a story in which "Rio the Rabbit" cooperates with his friend "Ken the Turtle" to cross a large bridge.

[0049] 2. The server converts the generated story into a specified format, such as text format.

[0050] Illustration generation

[0051] 1. The server uses image generation AI to create illustrations that correspond to the generated story. Multiple illustrations that match the scenes in the story are generated.

[0052] 2. The server integrates the generated illustrations and story into a coherent picture book format.

[0053] Multilingual Translation

[0054] 1. The server invokes a translation engine to translate the generated story into the language specified by the user (e.g., English, French).

[0055] 2. The server reviews the translated story and makes any necessary corrections.

[0056] Story Distribution

[0057] 1. The server generates digital picture book data that integrates the translated story and illustrations and sends it to the user's device.

[0058] 2. The terminal provides an interface that displays the received picture book data to the user, allowing the user to view the picture book.

[0059] Share your story

[0060] 1. The user shares the created picture book with other users using the sharing function of the device. Sharing methods include social networking sites, email, and cloud services.

[0061] 2. The device sends the picture book data to the sharing device so that other users can view it.

[0062] For example, a parent or guardian can create a story with the theme of "friendship" and featuring characters Rio the Rabbit and Ken the Turtle. The story is then generated and filtered by the server. Illustrations corresponding to the story are then generated, and a digital picture book translated into Japanese, English, and French is finally generated. The picture book is then delivered to the parent's smartphone, and the parent can share it with other parents via social networking sites.

[0063] As described above, the system of the present invention can effectively generate and provide safe and educational stories to children, which is expected to promote communication between parents and children and contribute to reducing educational disparities.

[0064] The processing flow will be explained below.

[0065] Step 1:

[0066] The user opens the terminal interface and inputs the theme (e.g., "friendship") and character (e.g., "Rio the Rabbit").

[0067] Step 2:

[0068] The terminal receives the entered theme and character settings and transmits this data to the server.

[0069] Step 3:

[0070] The server analyzes the theme and character settings received from the device and begins the filtering process, specifically running a screening process for banned words and explicit content.

[0071] Step 4:

[0072] The server checks the filtering results, and if any inappropriate content is found, it generates a message requesting the user to correct the content and sends it to the terminal. If no inappropriate content is found, it proceeds to the next step.

[0073] Step 5:

[0074] The server calls the generative AI model based on the verified theme and character settings to generate a story. For example, it generates a story in which Rio the rabbit cooperates with his friend Ken the turtle to cross a large bridge.

[0075] Step 6:

[0076] The server converts the generated story into text format and then uses image generation AI to begin generating illustrations that correspond to the story.

[0077] Step 7:

[0078] The server integrates the generated illustrations with the story to create a coherent digital picture book.

[0079] Step 8:

[0080] The server executes a translation process to translate the generated digital picture book into a user-specified language (e.g., English, French).

[0081] Step 9:

[0082] The server checks the quality of the translated story, makes corrections if necessary, and finally generates the finished digital picture book.

[0083] Step 10:

[0084] The server transmits the completed digital picture book data to the user's terminal.

[0085] Step 11:

[0086] The terminal provides an interface for displaying the received digital picture book, allowing the user to view it.

[0087] Step 12:

[0088] The user shares the created digital picture book with other users using the sharing function of the device. For example, the picture book data is sent via social networking sites, email, or cloud services.

[0089] Step 13:

[0090] The terminal transmits picture book data to a sharing terminal, so that other users can view the picture book data.

[0091] The above is the specific processing flow of the system of the present invention, which enables safe and educational stories to be created, distributed, and shared.

[0092] Example 1

[0093] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0094] Previous story generation systems were complicated in the process of generating safe and educational stories based on themes and characters set by the user, and lacked a means to translate and share the generated stories in multiple languages. Furthermore, there were no systems that automatically generated illustrations along with the story and integrated them. As a result, users had to manually create the story, illustrations, and translation, which required a great deal of effort and time.

[0095] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0096] In this invention, the server includes an input means for a user to set a theme and characters, a means for receiving input data of the set theme and characters and analyzing and filtering the input data, a means for generating a story using a generative AI model based on the filtered data, a means for converting the generated story into a text format, a means for generating illustrations using an image generation AI model based on the generated story, a means for integrating the generated story and illustrations into a consistent format, a means for translating the generated story into multiple languages, a means for delivering the translated story to a terminal, and a means for sharing the story delivered to the terminal with other users. This makes it possible to effectively generate safe and educational stories based on themes and characters set by users, translate them into multiple languages, and share them.

[0097] A "user" is a person or entity that accesses the system and configures themes and characters.

[0098] A "terminal" is an electronic device used by a user, such as a computer, smartphone, or tablet.

[0099] "Server" means the central control unit that manages and executes the functions of the entire system.

[0100] A "theme" is the main theme or central concept of a story.

[0101] A "character" is a person, animal, or fictional being that appears in a story.

[0102] "Input means" refers to an interface that allows a user to input themes and characters into the system.

[0103] A "generative AI model" is an artificial intelligence model that generates stories based on themes and characters set by the user.

[0104] "Filtering" is the process of analyzing input data and eliminating inappropriate content.

[0105] "Text format" is a data format that expresses the generated story in a sentence format.

[0106] An "image generation AI model" is an artificial intelligence model for generating illustrations that correspond to a story.

[0107] A "translation engine" is software or a service that translates generated stories into other languages.

[0108] A "consistent format" means that the story and illustrations are in harmony and have a sense of unity.

[0109] "Distribution means" refers to a mechanism or method for transmitting the generated content to the user's terminal.

[0110] "Sharing means" refers to a function or method for sharing the generated content with other users.

[0111] This system allows users to create safe and educational stories by setting themes and characters, and then translates and shares them in multiple languages. This system operates in cooperation with three main elements: a server, a terminal, and users.

[0112] Users access the system using a terminal and set a theme (e.g., "Friendship") and a character (e.g., "Rio the Rabbit"). The terminal then sends the theme and character settings entered by the user to the server. The server analyzes the received data and filters it using natural language processing technology (NLP). Specifically, it screens for prohibited words and extreme content, and if inappropriate data is included, it sends a message to the terminal requesting the user to correct it.

[0113] The server then uses a generative AI model (e.g., GPT-3) to generate a story based on the filtered data. The generated story is converted into text format and formatted as needed. The server then uses an image generation AI (e.g., DALL-E) to create illustrations corresponding to each scene in the generated story. The prompt text also includes a description of the specific scene. The generated illustrations are integrated into the story and compiled into a coherent picture book format.

[0114] The server then uses a translation engine (e.g., Google Translate API) to translate the generated story into the user's specified language (e.g., English, French). The translation result undergoes grammar checks and style guide application, correcting any necessary parts. The server then sends the translated story and illustrations integrated into the digital picture book data to the user's device. The device then provides an interface for displaying the received picture book data, allowing the user to view the picture book.

[0115] As a concrete example, a user sets the characters "Rio the Rabbit" and "Ken the Turtle" as characters with the theme of "friendship." The server filters this and uses a generative AI model to generate a story in which "Rio the Rabbit" and "Ken the Turtle" cooperate to cross a bridge. Next, DALL-E is used to generate illustrations of the scene crossing the bridge. Finally, the generated story is translated into Japanese, English, and French and delivered to the user's smartphone as a digital picture book.

[0116] Example prompt sentence:

[0117] Theme: "Friendship"

[0118] Characters: "Rio the Rabbit" and "Ken the Turtle"

[0119] Objective: To generate stories that have educational value for children.

[0120] The system of the present invention can effectively generate safe and educational stories based on themes and characters set by the user, and can translate and share them in multiple languages, which is expected to promote communication between parents and children and contribute to reducing educational disparities.

[0121] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0122] Step 1:

[0123] A user accesses the system using a terminal and inputs a theme (e.g., "Friendship") and a character (e.g., "Rio the Rabbit"). The input theme and character settings are sent to the server by the terminal. The input data is sent in the form of a theme and a character.

[0124] Step 2:

[0125] The server analyzes the theme and character settings received from the device. It uses natural language processing (NLP) technology to parse the input data and understand its meaning. Based on the results of this analysis, it checks for inappropriate data by screening it against a list of prohibited words and for extreme content. As an output, it generates data that is deemed appropriate or that requires correction.

[0126] Step 3:

[0127] If the server determines that the filtered data is appropriate, it calls a generative AI model (e.g., GPT-3) based on the data to generate a story. At this time, the prompt text includes the theme and character information entered by the user. The input data is the theme and character settings, and the output data is the generated story.

[0128] Step 4:

[0129] The server converts the generated story into a text format and formats it, for example dividing the story into paragraphs, applying a particular font size and style, and adjusting line breaks and punctuation as needed. The input data is the generated story, and the output data is the formatted story text.

[0130] Step 5:

[0131] The server calls an image generation AI (e.g., DALL-E) to create illustrations corresponding to the generated story. By including a specific description of a particular scene in the story in the prompt, an illustration appropriate for that scene is generated. The input data is a description of the story scene, and the output data is the generated illustration.

[0132] Step 6:

[0133] The server integrates the generated illustrations and story into a coherent picture book format. This is the process of adjusting the placement of illustrations and text and determining the page layout. The input data is the story in text format and the generated illustrations, and the output data is the integrated digital picture book.

[0134] Step 7:

[0135] The server uses a translation engine (e.g., Google Translate API) to translate the generated story into the language specified by the user (e.g., English, French). It checks the translated content for grammar and makes any necessary corrections. The input data is the story in text format, and the output data is the translated story.

[0136] Step 8:

[0137] The server generates digital picture book data that integrates the translated story and illustrations and sends it to the user's terminal. The terminal provides an interface for displaying the received picture book data, allowing the user to view the picture book. The input data is the translated story and the generated illustrations, and the output data is the user's viewing interface.

[0138] Step 9:

[0139] Users can share the created picture book with other users using the sharing function of their device. Sharing methods include social networking sites, email, and cloud services. The input data is digital picture book data, and the output data is in a format that can be viewed by the recipient users.

[0140] (Application example 1)

[0141] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0142] In conventional story generation systems, it is common for users to set a theme and characters, and then generate a story based on those. However, they lack the functionality to compile the generated story and illustrations into a digital booklet in a consistent format, distribute it in multiple languages, or easily share it with other users. For this reason, there is a need for a system that increases users' freedom of expression and provides convenience in international multilingual support.

[0143] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0144] In this invention, the server includes means for receiving input data, means for filtering the received input data, means for generating a story based on the filtered data, means for generating images based on the generated story, means for translating the generated story into multiple languages, means for distributing the translated story to a terminal, means for sharing the story distributed to the terminal with other users, means for receiving theme and character settings from a user, means for automatically generating a story based on the received theme and characters, and means for integrating the story and images into a digital booklet format. This makes it possible to translate, distribute, and share educational and safe stories and corresponding illustrations into multiple languages ​​in a consistent digital booklet format based on the themes and characters set by the user.

[0145] The "means for receiving input data" is a function for receiving the theme and character information provided by the user via the terminal.

[0146] "Means for filtering received input data" refers to a function for filtering out inappropriate content from input data and narrowing it down to safe and appropriate data.

[0147] The "means for generating a story based on filtered data" is a function for automatically generating a story based on verified themes and character settings.

[0148] The "means for generating images based on the generated story" is a function for automatically generating illustrations according to the content of the story.

[0149] The "means for translating the generated story into multiple languages" is a function for automatically converting the generated story into multiple specified languages.

[0150] The "means for delivering the translated story to the terminal" is a function for transferring the translated story and corresponding illustrations to the user's terminal.

[0151] The "means for sharing a story delivered to the terminal with other users" is a function that enables a user to share a received story with other users via a social networking service or email.

[0152] The "means for accepting theme and character settings from the user" is a function that provides an input interface for the user to set the theme and character.

[0153] "Means for automatically generating a story based on the accepted theme and characters" is a function for automatically generating a story using a generative AI model based on the theme and character information set by the user.

[0154] "Means for integrating stories and images into a digital booklet format" is a function for combining the generated stories and illustrations and editing them into a coherent digital picture book format.

[0155] A "social networking service" is a platform for sharing information and interacting with other users over the Internet.

[0156] "Email" is a means of communication for sending and receiving text messages and files over the Internet.

[0157] The following describes an embodiment of the present invention. The system mainly consists of three elements: a server, a terminal, and a user. The specific roles and processing procedures of each element are described below.

[0158] System programs and their processing

[0159] User Theme and Character Settings

[0160] Users access the system through a device such as a smartphone or tablet and use an interface to input the theme (e.g., "Friendship") and character (e.g., "Rio the Rabbit") to set up the system. This information is then sent from the device to the server.

[0161] Filtering input data

[0162] The server analyzes the received themes and character settings, and screens them for banned words and explicit content. Scripts written in Python and Flask perform the filtering. If inappropriate data is detected, a message is sent to the user's device, prompting them to correct it.

[0163] Narrative Generation

[0164] Based on the filtered themes and character settings, the server invokes a generative AI model (e.g., GPT-4) to generate a story, which is then converted into text format and temporarily stored on the server.

[0165] Illustration generation

[0166] The server then uses image generation AI (e.g., Stable Diffusion or DALL-E) to generate illustrations that correspond to the generated story. These illustrations are aligned with the scenes in the story.

[0167] Multilingual Translation

[0168] The generated story is translated into the specified language (e.g., English, French) using the Google Translate API or DeepL API. The translated story is stored on the server and corrected as needed.

[0169] Story Distribution

[0170] The server then aggregates the translated stories and illustrations into a digital booklet and delivers it to users' devices, where they can view it through an application on their smartphones or tablets.

[0171] Share your story

[0172] Users can use the sharing function of their terminals to share the generated digital booklet with other users via social networking services or email.

[0173] Overview of the technologies and programs used

[0174] Server side: Python, Flask (or Django), a generative AI model (e.g. GPT-4), a translation engine (Google Translate API or DeepL).

[0175] Frontend: React Native (for mobile apps) or React.js (for web).

[0176] AI model: OpenAI's GPT-4, Stable Diffusion or DALL-E (illustration generation).

[0177] Database: PostgreSQL (storage of user data, generated stories, and illustrations).

[0178] Specific examples

[0179] For example, a child might create a story with the theme of "friendship" and featuring characters Rio the rabbit and Ken the turtle. The story is then generated and filtered by the server. Illustrations corresponding to the story are then generated, and a digital booklet translated into Japanese, English, and French is finally generated. This picture book is then delivered to the user's smartphone, and the user can share it with other users via social networking sites.

[0180] Generative AI model prompt example

[0181] prompt:

[0182] Theme: "Friendship"

[0183] Characters: "Rio the Rabbit" and "Ken the Turtle"

[0184] Instructions: Create an educational and safe story for children. The story should be based on the adventures of Rio the rabbit and Ken the turtle who team up to cross a big bridge. Write in a friendly, easy-to-understand voice.

[0185] In this way, the system of the present invention can effectively generate and provide safe and educational stories to users, which is expected to promote communication between users and contribute to the spread of education.

[0186] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0187] Step 1:

[0188] Users access the system through a device such as a smartphone or tablet and use the interface to input the theme (e.g., "Friendship") and character (e.g., "Rio the Rabbit") to set the game. The set information is sent from the device to the server as input data. The input is the theme and character setting data from the user's interface, and the output is the input data sent to the server.

[0189] Step 2:

[0190] The server filters the input data for themes and character settings received from the device. Specifically, it uses scripts using Python and Flask to screen for banned words and explicit content. The input is the theme and character setting data, and filtering is performed as data processing, with the output being the appropriate filtered data.

[0191] Step 3:

[0192] The server generates a story using a generative AI model (e.g., GPT-4) based on the filtered theme and character settings. The input is the filtered data, and a story is generated using the generative AI model as data calculation. The output is the text data of the generated story. The generated story is temporarily stored on the server.

[0193] Step 4:

[0194] The server uses an image generation AI (e.g., Stable Diffusion or DALL-E) based on the text data of the generated story to generate illustrations corresponding to each scene in the story. The input is the text data of the generated story, and the image generation AI is used as data calculation to generate illustrations, and the output is the generated illustration.

[0195] Step 5:

[0196] The server translates the generated story into multiple languages. Specifically, it uses the Google Translate API or DeepL API to translate the input data, which is the generated story text. The translation API is called as data processing, and the output is the translated multilingual story text data.

[0197] Step 6:

[0198] The server integrates the translated story and the generated illustrations into a digital booklet. The input is the text data of the translated story and the generated illustration data, which are integrated as data processing, and the output is the digital booklet data.

[0199] Step 7:

[0200] The server distributes the digital booklet data to the user's terminal. The input is the digital booklet data, which is distributed to the terminal as data transfer, and the output is the digital booklet displayed on the user's terminal.

[0201] Step 8:

[0202] The user uses the sharing function of the terminal to share the generated digital booklet with other users via social networking services or email. The input is the distributed digital booklet data, and the sharing function is used as data transfer, and the output is the digital booklet shared with other users.

[0203] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0204] This invention is a system that combines an emotion engine that recognizes the user's emotions, generates safe and educational stories, and translates and shares them in multiple languages. The program and processing of this system are explained in detail below. The system mainly works in cooperation with three elements: the server, the terminal, and the user.

[0205] Description of system programs and processes

[0206] User theme and character settings, and emotion recognition

[0207] 1. The user opens the terminal interface and inputs the theme (e.g., "Friendship") and character (e.g., "Rio the Rabbit").

[0208] 2. The device acquires emotion data from the user's facial expressions and voice and sends that data to the emotion engine.

[0209] 3. The emotion engine analyzes the acquired emotion data and recognizes the user's current emotion. The analysis results are reflected in the theme and character settings.

[0210] Filtering input data

[0211] 1. The server analyzes all input data, including the theme and character settings received from the device and the emotion data obtained from the emotion engine, and starts the filtering process. Specifically, it runs a screening process for banned words and explicit content.

[0212] 2. The server checks the filtering results and if any inappropriate content is found, it generates a message asking the user to correct it and sends it to the terminal. If there is no inappropriate content, it proceeds to the next step.

[0213] Story and illustration generation

[0214] 1. The server calls the generative AI model based on the verified theme, character settings, and emotional data to generate a story. For example, if the user is happy, it generates a story with a happy ending, depending on the user's emotions.

[0215] 2. The server converts the generated story into text format, and then uses image generation AI to generate illustrations corresponding to the story. Based on the emotional data, the illustrations also reflect emotional elements.

[0216] Multilingual Translation

[0217] 1. The server invokes a translation engine that translates the generated story into the language specified by the user (e.g., English, French).

[0218] 2. The server reviews the translated story and makes corrections if necessary.

[0219] Story Distribution

[0220] 1. The server generates digital picture book data that integrates the translated story and illustrations and sends it to the user's device.

[0221] 2. The terminal provides an interface for displaying the received digital picture book, allowing the user to view it.

[0222] Share your story

[0223] 1. The user shares the created digital picture book with other users using the sharing function of their device. Sharing methods include social networking sites, email, and cloud services.

[0224] 2. The device sends the picture book data to the sharing device so that other users can view it.

[0225] For example, if a parent or guardian selects characters Rio the Rabbit and Ken the Turtle as characters with the theme of "friendship" and the emotion they are feeling is "fun," a story about the fun adventures of Rio the Rabbit and Ken the Turtle is generated based on this emotion. The story is filtered, appropriate illustrations are generated, and a digital picture book translated into Japanese, English, and French is delivered to the parent or guardian's smartphone. The parent or guardian can then share the picture book with friends via social media.

[0226] As described above, the system of the present invention can provide more personalized educational resources by generating stories and illustrations that reflect the user's emotions, which is expected to promote communication between parents and children and contribute to reducing educational disparities.

[0227] The processing flow will be explained below.

[0228] Step 1:

[0229] The user opens the terminal interface and inputs the theme (e.g., "friendship") and character (e.g., "Rio the Rabbit").

[0230] Step 2:

[0231] The device captures the user's facial expressions and voice through the user's camera and microphone, and obtains emotional data in real time.

[0232] Step 3:

[0233] The device sends the acquired emotion data to the emotion engine, which analyzes the user's current emotion and identifies emotion categories such as "joy," "sadness," and "surprise."

[0234] Step 4:

[0235] The emotion engine returns the analysis results to the device, which then transmits the information to the server, thereby transmitting the user's emotion data to the server.

[0236] Step 5:

[0237] The server performs filtering based on the theme, character settings, and emotion data received from the device, specifically checking for taboo words and explicit content.

[0238] Step 6:

[0239] The server checks the filtering results, and if any inappropriate content is found, it generates a message requesting the user to correct the content and sends it to the terminal. If no inappropriate content is found, it proceeds to the next step.

[0240] Step 7:

[0241] The server then invokes a generative AI model based on verified themes, character settings, and emotional data to generate a story. For example, if the user expresses the emotion of "enjoyment," a story containing positive and enjoyable episodes will be generated.

[0242] Step 8:

[0243] The server converts the generated story into a text format, which can be modified as needed.

[0244] Step 9:

[0245] The server uses image generation AI to generate a corresponding illustration based on the story. The illustration also reflects the user's emotional data. For example, if the emotion of "fun" is conveyed, an illustration of a bright and cheerful scene is generated.

[0246] Step 10:

[0247] The server integrates the generated illustrations with the story to create a coherent digital picture book.

[0248] Step 11:

[0249] The server invokes a translation engine to translate the generated digital picture book into a language specified by the user (e.g., English, French).

[0250] Step 12:

[0251] The server checks the quality of the translated stories and makes corrections if necessary.

[0252] Step 13:

[0253] The server completes the digital picture book integrating the translation and illustrations and transmits the data to the user's terminal.

[0254] Step 14:

[0255] The terminal provides an interface for displaying the received digital picture book, allowing the user to view it.

[0256] Step 15:

[0257] The user can share the created digital picture book with other users using the sharing function of the device. Sharing methods include social networking sites, email, and cloud services.

[0258] Step 16:

[0259] The terminal transmits picture book data to a sharing terminal, so that other users can view the picture book data.

[0260] The above is the specific processing flow of the system of the present invention, which combines an emotion engine. This flow makes it possible to generate, distribute, and share personalized stories that reflect the user's emotions.

[0261] Example 2

[0262] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0263] Current story generation systems struggle to generate content that reflects the user's emotions, resulting in a lack of personalized educational resources. Furthermore, they lack the ability to automatically translate the generated stories and images into multiple languages ​​to match the user's emotions and themes, making them unable to meet the needs of global users. Furthermore, they have yet to provide digital content in a format that can be easily shared among users.

[0264] The specification process by the specification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving input data, means for filtering the received input data, means for generating a story based on the filtered data, means for generating images based on the generated story, means for translating the generated story into multiple languages, means for delivering the translated story to the terminal, means for sharing the story delivered to the terminal with other users, means for acquiring and analyzing user emotion data, and means for generating a story and images based on the acquired emotion data. This makes it possible to generate a personalized story and images based on the user's emotions, translate the story into multiple languages, and provide digital content that can be easily shared among users.

[0265] The "means for receiving input data" is a method by which a user inputs theme and character settings and transmits them to the system.

[0266] A "means for filtering received input data" is a method for analyzing received data and removing inappropriate content.

[0267] A "means for generating a story" is a method for automatically creating a story based on the filtered data and the user's emotional data.

[0268] An "image generation means" is a method for creating related visual content based on the generated narrative.

[0269] A "means for translating stories" is a method for translating the generated stories into multiple languages.

[0270] A "means for delivering a story to a terminal" is a method for transferring a translated story to a user's device.

[0271] "Means for sharing stories with other users" refers to methods that provide an interface or protocol for sharing distributed stories.

[0272] "Means for acquiring and analyzing emotion data" refers to a method for collecting and analyzing emotions from the user's facial expressions, voice, etc.

[0273] The "means for generating a story and images based on the acquired emotion data" is a method for generating a story and images using the analyzed emotion data.

[0274] The present invention is a system that combines an emotion engine that recognizes a user's emotions, generates safe and educational stories, and translates and shares them in multiple languages. Specific embodiments of this system and their operation are described in detail below.

[0275] System configuration

[0276] The system includes the following elements:

[0277] 1. Device: A device on which users input theme and character settings and acquire facial and voice data. This includes smartphones, tablets, and PCs.

[0278] 2. Server: A central processing unit that filters input data, generates stories and images, translates, and distributes data.

[0279] 3. Emotion engine: Software for acquiring and analyzing emotional data from the user's facial expressions and voice. Uses an API for emotion analysis.

[0280] 4. Generative AI model: An artificial intelligence model for generating stories and images based on user input and emotional data. Specifically, it uses large-scale language models and image generation models such as GPT (Generative Pre-trained Transformer).

[0281] How it works

[0282] User theme and character settings, and emotion recognition

[0283] The user opens the device's interface and inputs a theme (e.g., "Friendship") and a character (e.g., "Rio the Rabbit"). The device captures the user's facial expressions and voice through a camera and microphone, and sends the data to the emotion engine. The emotion engine analyzes the acquired emotion data and recognizes the user's current emotion. The results of this analysis are sent to the server and reflected in the theme and character settings.

[0284] Filtering input data

[0285] The server analyzes all input data, including the theme and character settings received from the device and the emotional data obtained from the emotion engine, and then begins the filtering process, specifically running a screening process for banned words and explicit content.

[0286] Story and illustration generation

[0287] The server generates a story by calling a generative AI model based on verified themes, character settings, and emotional data. For example, if the user is happy, it generates a story with a happy ending.

[0288] Example prompt sentence:

[0289] "Generate a story about friendship, featuring Rio the rabbit and Ken the turtle going on a fun adventure. The user emotion is 'fun'."

[0290] The server then converts the generated story into text format and uses image generation AI to generate illustrations that correspond to the story. The generated illustrations also reflect emotional elements based on the emotional data.

[0291] Example prompt sentence:

[0292] "Generate an illustration that reflects the emotion of enjoyment based on the following story: 'Rio the rabbit and Ken the turtle go on a fun adventure.'"

[0293] Multilingual Translation

[0294] The server calls a translation engine to translate the generated story into the language specified by the user (e.g., English, French). It checks the translated text obtained from the translation engine and makes corrections if necessary.

[0295] Story Distribution

[0296] The server generates digital picture book data that integrates the translated story and illustrations and sends it to the user's device, which provides an interface for displaying the received digital picture book so that the user can view it.

[0297] Share your story

[0298] Users can share the digital picture book they have created with other users using the sharing function of their device. Sharing methods include social networking sites, email, and cloud services. The device sends the picture book data to the destination device, allowing other users to view it.

[0299] Specific examples

[0300] For example, if a parent or guardian selects characters Rio the Rabbit and Ken the Turtle as characters with the theme of "friendship" and the emotion at that time is "fun," a story about the fun adventures of Rio the Rabbit and Ken the Turtle is generated based on this emotion. This story is filtered, appropriate illustrations are generated, and a digital picture book translated into Japanese, English, and French is delivered to the parent or guardian's smartphone. The parent or guardian can then share this picture book with friends via social media.

[0301] This invention can provide more personalized educational resources by generating stories and illustrations that reflect the user's emotions, which is expected to promote communication between parents and children and contribute to reducing educational disparities.

[0302] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0303] Step 1:

[0304] The user opens the device interface and inputs the theme (e.g., "Friendship") and character (e.g., "Rio the Rabbit"), and when the user confirms the input data, the device sends the data to the server.

[0305] Input: Input data for the theme "Friendship" and the character "Rabbit Rio"

[0306] Output: The input data is sent to the server

[0307] Step 2:

[0308] The device acquires the user's facial expressions and voice and sends them to the emotion engine. The emotion engine analyzes the facial and voice data to recognize the user's current emotion. The result is then sent to the server.

[0309] Input: User's facial expression data and voice data

[0310] Output: Recognized emotion data

[0311] Step 3:

[0312] The server receives the input data (theme and character settings) and analyzed emotional data, and performs filtering. Specifically, it screens for prohibited words and extreme content. Once filtering is complete, it notifies the user of the results.

[0313] Input: Theme "Friendship", Character "Rabbit Rio", Emotion data

[0314] Output: Filtering results and a message requesting corrections if necessary

[0315] Step 4:

[0316] The server invokes the generative AI model based on the verified input data and emotion data to generate a story. For example, if the user's emotion is "fun," it generates a story with a happy ending.

[0317] Input: Filtered data and sentiment data

[0318] Output: Generated narrative text

[0319] Specific operation example:

[0320] The server sends a prompt to the AI ​​model, such as "Generate a story about friendship, with Rio the rabbit and Ken the turtle going on a fun adventure. The user's emotion is 'fun'."

[0321] Step 5:

[0322] The server converts the generated story text into a text format and calls an image generation AI to generate illustrations corresponding to the story, which reflect the user's emotions.

[0323] Input: Generated narrative text

[0324] Output: Generated illustration

[0325] Specific operation example:

[0326] The prompt text "Generate an illustration that reflects the emotion of enjoyment based on the following story: 'Rio the rabbit and Ken the turtle go on a fun adventure'" is sent to the image generation AI.

[0327] Step 6:

[0328] The server invokes a translation engine that translates the generated story text into the language specified by the user (e.g., English, French). The translated text is returned to the server, which reviews it and makes corrections if necessary.

[0329] Input: Japanese story text

[0330] Output: Translated story text

[0331] Step 7:

[0332] The server integrates the translated story and illustrations to create digital picture book data, which is then sent to the user's device.

[0333] Input: translated story text, generated illustrations

[0334] Output: Digital picture book data

[0335] Step 8:

[0336] An interface for displaying the digital picture book received by the terminal is provided, allowing the user to view it.

[0337] Input: Digital picture book data

[0338] Output: A displayed digital picture book

[0339] Step 9:

[0340] Users can share the digital picture book they have created with other users using the sharing function of their device. Sharing methods include social networking sites, email, and cloud services.

[0341] Input: Digital picture book data

[0342] Output: Digital picture books shared via social media, email, and the cloud

[0343] Step 10:

[0344] The terminal transmits picture book data to a sharing device so that other users can view it.

[0345] Input: Share request

[0346] Output: Digital picture book data transferred to the sharing device

[0347] (Application example 2)

[0348] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0349] Conventional content generation and sharing systems were unable to generate and provide individual stories that took the user's emotions into consideration. This made it difficult to provide personalized content that responded to the user's emotions, and the educational and entertainment benefits were not fully realized. Furthermore, the inability to translate and share content in multiple languages ​​limited global sharing. This created challenges that made it difficult to improve the user experience and correct educational disparities.

[0350] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0351] In this invention, the server includes means for receiving input data and emotional data, means for filtering the received input data and emotional data, means for generating a story based on the filtered data and emotional data, means for generating images based on the generated story, means for translating the generated story into multiple languages, means for distributing the translated story to a terminal, means for sharing the story distributed to the terminal with other users, means for accepting theme and character settings and emotional data from a user, means for automatically generating a story based on the accepted theme and characters, means for personalizing the story based on the accepted emotional data, and means for simultaneously generating a story and images, integrating them, evaluating them based on the user's emotional data, and ensuring quality. This enables the generation of personalized stories and illustrations that take user emotions into consideration, realizes multilingual translation and easy sharing, and enables an improved user experience and mitigation of educational disparities.

[0352] "Input Data" refers to the theme, characters, and other setting information that a user provides to the system.

[0353] "Emotion data" refers to information about the emotional state of a user obtained from facial expressions, voice, etc.

[0354] "Filtering" refers to the process of removing inappropriate content or prohibited words from received data.

[0355] "Story" refers to a document or text that is generated based on themes, characters, and user emotional data.

[0356] "Images" refers to illustrations and visual content generated based on the story content and user emotional data.

[0357] "Translation" refers to the process of converting a generated story into multiple languages.

[0358] "Terminal" refers to an electronic device used by a user, such as a smartphone, tablet, or computer.

[0359] "Sharing" refers to the process of providing or distributing the generated stories and images to other users and platforms.

[0360] "Theme" refers to a concept that indicates the main point or main point of a story.

[0361] "Characters" refers to characters that appear in stories or animated characters.

[0362] "Personalization" refers to the process of adapting content based on a user's individual emotional data and preferences.

[0363] "Synthesis" refers to the process of bringing together the generated stories and images into a single piece of content.

[0364] "Evaluation" refers to the process used to judge the quality and appropriateness of the stories and images produced.

[0365] "Quality assurance" refers to the process of verifying that the generated content is appropriate and guaranteeing its quality.

[0366] The present invention is a system that combines an emotion engine that recognizes a user's emotions, generates safe and educational stories, and translates and shares them in multiple languages. An embodiment of this system will be described in detail below.

[0367] Description of system programs and processes

[0368] User theme and character settings, and emotion recognition

[0369] The user opens the device's interface and inputs the theme and character. For example, the user may select the theme "friendship" and "Rio the rabbit" as the character. Emotional data is also acquired from the user's facial expressions and voice. The device sends this data to an emotion engine, which analyzes it to recognize the user's current emotion. The results of this analysis are reflected in the theme and character settings.

[0370] Filtering input data

[0371] The server analyzes and filters all input data, including the theme and character settings received from the device and the emotion data obtained from the emotion engine. Specifically, a screening process is carried out for a list of prohibited words and explicit content. If inappropriate content is included, the server generates a message requesting the user to correct it and sends it to the device. If there is no inappropriate content, the process proceeds to the next step.

[0372] Story and illustration generation

[0373] The server generates a story by calling a generative AI model based on verified themes, character settings, and emotional data. For example, if the user is happy, a story with a happy ending is generated. This story is converted into text format, and then an image generation AI is used to generate illustrations corresponding to the story. Based on the emotional data, the illustrations also reflect emotional elements.

[0374] Multilingual Translation

[0375] The server invokes a translation engine to translate the generated story into a language specified by the user, such as Japanese, English, French, etc. The translated story is then reviewed by the server and any necessary corrections are made.

[0376] Story Distribution

[0377] The server generates digital picture book data that integrates the translated story and illustrations and transmits it to the user's terminal, which provides an interface for displaying the received digital picture book so that the user can view it.

[0378] Share your story

[0379] Users can share the digital picture book they have created with other users using the sharing function of their device. Sharing methods include social networking sites, email, and cloud services. The device sends the picture book data to the destination device, allowing other users to view it.

[0380] Hardware and software used

[0381] The server runs software such as generative AI models, emotion recognition engines, image generation AI, translation engines, and data filtering systems. Specific software and hardware used include the Google Translate API, EmotionRecognizer, StoryGenerator, and ImageGenerator. These work together to enable the generation, translation, and sharing of stories and illustrations based on user requests.

[0382] Specific examples

[0383] For example, if a parent or guardian selects characters Rio the Rabbit and Ken the Turtle as characters with the theme of "friendship" and the emotion at that time is "fun," a story about the fun adventures of Rio the Rabbit and Ken the Turtle is generated based on this emotion. The story is filtered and appropriate illustrations are generated, and then a digital picture book translated into Japanese, English, and French is delivered to the parent or guardian's smartphone. The parent or guardian can share this picture book with friends via social media.

[0384] Prompt Sentence Examples

[0385] Theme: Friendship

[0386] Character: Rio the Rabbit

[0387] Emotion data: {'emotion': 'joy'}

[0388] prompt:

[0389] "Rio the rabbit was having a great time playing with his friends. They decided to go on a new adventure together. Along the way, Rio..."

[0390] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0391] Step 1:

[0392] The user opens the device interface and inputs the theme and character. For example, the theme is set to "friendship" and the character is set to "Rio the rabbit." The user's facial expressions and voice data are also input. This input data is sent to the device (input data and emotion data).

[0393] Step 2:

[0394] The device processes the received theme, character, and emotion data and sends it to the server. The emotion data includes the results of the user's facial expression analysis and voice analysis. The device calls the emotion recognition engine to identify the user's emotion and sends the result as data (input data: theme, character, emotion data / output data: theme, character, emotion analysis result).

[0395] Step 3:

[0396] The server analyzes the theme, character, and emotion data received from the device and performs a filtering process. The filtering process screens the data based on a list of prohibited words and extreme content to check for inappropriate content. If inappropriate content is found, a message is generated requesting the user to correct it (input data: theme, character, emotion analysis results / output data: filtered data or correction request message).

[0397] Step 4:

[0398] Based on the filtered data, the server calls a generative AI model to generate a story. The generative AI model adjusts the tone and ending of the story based on the emotional data. For example, if the emotional data indicates joy, a story with a happy ending will be generated (input data: filtered data / output data: generated story).

[0399] Step 5:

[0400] Next, the server calls an image generation AI based on the generated story, which generates illustrations that correspond to the content of the story. Emotional data is also taken into account, and emotional elements are reflected in the illustrations (input data: generated story, emotion data / output data: generated illustrations).

[0401] Step 6:

[0402] The server invokes a translation engine to translate the generated story into multiple languages, for example, from Japanese to English and French. The translation is done automatically and can be manually corrected if necessary (input data: generated story / output data: translated story).

[0403] Step 7:

[0404] The server generates digital picture book data that integrates the translated story and generated illustrations and sends it to the user's device. The device provides an interface for displaying the received digital picture book, allowing the user to easily browse it (input data: translated story, generated illustrations / output data: digital picture book data).

[0405] Step 8:

[0406] Users can share the digital picture book they have created with other users using the sharing function of their device. Social networking sites, email, cloud services, and other methods of sharing are available. The device sends the picture book data to the destination device, allowing other users to view it (input data: digital picture book data; output data: data sent to the destination device).

[0407] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0408] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0409] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0410] [Second embodiment]

[0411] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0412] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0413] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0414] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0415] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0416] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0417] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0418] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0419] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0420] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0421] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0422] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0423] This system allows users to create safe and educational stories by setting themes and characters, and then translates and shares them in multiple languages. The program and processing of this system are described in detail below. The system mainly works in conjunction with three elements: the server, the terminal, and the user.

[0424] Description of system programs and processes

[0425] User Theme and Character Settings

[0426] 1. A user accesses the system through a terminal and inputs a theme (e.g., "friendship") and a character (e.g., "Rio the rabbit") using the input interface.

[0427] 2. The device sends the entered theme and character settings to the server.

[0428] Filtering input data

[0429] 1. The server analyzes and filters the theme and character settings received from the device, specifically by screening for prohibited words and explicit content.

[0430] 2. If the server determines that the input data is appropriate based on the filtering results, it proceeds to the next step using that data. If the data contains inappropriate data, it sends a message to the terminal requesting the user to correct it.

[0431] Narrative Generation

[0432] 1. The server calls a generative AI model based on the verified theme and character settings to generate a story. For example, it generates a story in which "Rio the Rabbit" cooperates with his friend "Ken the Turtle" to cross a large bridge.

[0433] 2. The server converts the generated story into a specified format, such as text format.

[0434] Illustration generation

[0435] 1. The server uses image generation AI to create illustrations that correspond to the generated story. Multiple illustrations that match the scenes in the story are generated.

[0436] 2. The server integrates the generated illustrations and story into a coherent picture book format.

[0437] Multilingual Translation

[0438] 1. The server invokes a translation engine to translate the generated story into the language specified by the user (e.g., English, French).

[0439] 2. The server reviews the translated story and makes any necessary corrections.

[0440] Story Distribution

[0441] 1. The server generates digital picture book data that integrates the translated story and illustrations and sends it to the user's device.

[0442] 2. The terminal provides an interface that displays the received picture book data to the user, allowing the user to view the picture book.

[0443] Share your story

[0444] 1. The user shares the created picture book with other users using the sharing function of the device. Sharing methods include social networking sites, email, and cloud services.

[0445] 2. The device sends the picture book data to the sharing device so that other users can view it.

[0446] For example, a parent or guardian can create a story with the theme of "friendship" and featuring characters Rio the Rabbit and Ken the Turtle. The story is then generated and filtered by the server. Illustrations corresponding to the story are then generated, and a digital picture book translated into Japanese, English, and French is finally generated. The picture book is then delivered to the parent's smartphone, and the parent can share it with other parents via social networking sites.

[0447] As described above, the system of the present invention can effectively generate and provide safe and educational stories to children, which is expected to promote communication between parents and children and contribute to reducing educational disparities.

[0448] The processing flow will be explained below.

[0449] Step 1:

[0450] The user opens the terminal interface and inputs the theme (e.g., "friendship") and character (e.g., "Rio the Rabbit").

[0451] Step 2:

[0452] The terminal receives the entered theme and character settings and transmits this data to the server.

[0453] Step 3:

[0454] The server analyzes the theme and character settings received from the device and begins the filtering process, specifically running a screening process for banned words and explicit content.

[0455] Step 4:

[0456] The server checks the filtering results, and if any inappropriate content is found, it generates a message requesting the user to correct the content and sends it to the terminal. If no inappropriate content is found, it proceeds to the next step.

[0457] Step 5:

[0458] The server calls the generative AI model based on the verified theme and character settings to generate a story. For example, it generates a story in which Rio the rabbit cooperates with his friend Ken the turtle to cross a large bridge.

[0459] Step 6:

[0460] The server converts the generated story into text format and then uses image generation AI to begin generating illustrations that correspond to the story.

[0461] Step 7:

[0462] The server integrates the generated illustrations with the story to create a coherent digital picture book.

[0463] Step 8:

[0464] The server executes a translation process to translate the generated digital picture book into a user-specified language (e.g., English, French).

[0465] Step 9:

[0466] The server checks the quality of the translated story, makes corrections if necessary, and finally generates the finished digital picture book.

[0467] Step 10:

[0468] The server transmits the completed digital picture book data to the user's terminal.

[0469] Step 11:

[0470] The terminal provides an interface for displaying the received digital picture book, allowing the user to view it.

[0471] Step 12:

[0472] The user shares the created digital picture book with other users using the sharing function of the device. For example, the picture book data is sent via social networking sites, email, or cloud services.

[0473] Step 13:

[0474] The terminal transmits picture book data to a sharing terminal, so that other users can view the picture book data.

[0475] The above is the specific processing flow of the system of the present invention, which enables safe and educational stories to be created, distributed, and shared.

[0476] Example 1

[0477] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0478] Previous story generation systems were complicated in the process of generating safe and educational stories based on themes and characters set by the user, and lacked a means to translate and share the generated stories in multiple languages. Furthermore, there were no systems that automatically generated illustrations along with the story and integrated them. As a result, users had to manually create the story, illustrations, and translation, which required a great deal of effort and time.

[0479] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0480] In this invention, the server includes an input means for a user to set a theme and characters, a means for receiving input data of the set theme and characters and analyzing and filtering the input data, a means for generating a story using a generative AI model based on the filtered data, a means for converting the generated story into a text format, a means for generating illustrations using an image generation AI model based on the generated story, a means for integrating the generated story and illustrations into a consistent format, a means for translating the generated story into multiple languages, a means for delivering the translated story to a terminal, and a means for sharing the story delivered to the terminal with other users. This makes it possible to effectively generate safe and educational stories based on themes and characters set by users, translate them into multiple languages, and share them.

[0481] A "user" is a person or entity that accesses the system and configures themes and characters.

[0482] A "terminal" is an electronic device used by a user, such as a computer, smartphone, or tablet.

[0483] "Server" means the central control unit that manages and executes the functions of the entire system.

[0484] A "theme" is the main theme or central concept of a story.

[0485] A "character" is a person, animal, or fictional being that appears in a story.

[0486] "Input means" refers to an interface that allows a user to input themes and characters into the system.

[0487] A "generative AI model" is an artificial intelligence model that generates stories based on themes and characters set by the user.

[0488] "Filtering" is the process of analyzing input data and eliminating inappropriate content.

[0489] "Text format" is a data format that expresses the generated story in a sentence format.

[0490] An "image generation AI model" is an artificial intelligence model for generating illustrations that correspond to a story.

[0491] A "translation engine" is software or a service that translates generated stories into other languages.

[0492] A "consistent format" means that the story and illustrations are in harmony and have a sense of unity.

[0493] "Distribution means" refers to a mechanism or method for transmitting the generated content to the user's terminal.

[0494] "Sharing means" refers to a function or method for sharing the generated content with other users.

[0495] This system allows users to create safe and educational stories by setting themes and characters, and then translates and shares them in multiple languages. This system operates in cooperation with three main elements: a server, a terminal, and users.

[0496] Users access the system using a terminal and set a theme (e.g., "Friendship") and a character (e.g., "Rio the Rabbit"). The terminal then sends the theme and character settings entered by the user to the server. The server analyzes the received data and filters it using natural language processing technology (NLP). Specifically, it screens for prohibited words and extreme content, and if inappropriate data is included, it sends a message to the terminal requesting the user to correct it.

[0497] The server then uses a generative AI model (e.g., GPT-3) to generate a story based on the filtered data. The generated story is converted into text format and formatted as needed. The server then uses an image generation AI (e.g., DALL-E) to create illustrations corresponding to each scene in the generated story. The prompt text also includes a description of the specific scene. The generated illustrations are integrated into the story and compiled into a coherent picture book format.

[0498] The server then uses a translation engine (e.g., Google Translate API) to translate the generated story into the user's specified language (e.g., English, French). The translation result undergoes grammar checks and style guide application, correcting any necessary parts. The server then sends the translated story and illustrations integrated into the digital picture book data to the user's device. The device then provides an interface for displaying the received picture book data, allowing the user to view the picture book.

[0499] As a concrete example, a user sets the characters "Rio the Rabbit" and "Ken the Turtle" as characters with the theme of "friendship." The server filters this and uses a generative AI model to generate a story in which "Rio the Rabbit" and "Ken the Turtle" cooperate to cross a bridge. Next, DALL-E is used to generate illustrations of the scene crossing the bridge. Finally, the generated story is translated into Japanese, English, and French and delivered to the user's smartphone as a digital picture book.

[0500] Example prompt sentence:

[0501] Theme: "Friendship"

[0502] Characters: "Rio the Rabbit" and "Ken the Turtle"

[0503] Objective: To generate stories that have educational value for children.

[0504] The system of the present invention can effectively generate safe and educational stories based on themes and characters set by the user, and can translate and share them in multiple languages, which is expected to promote communication between parents and children and contribute to reducing educational disparities.

[0505] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0506] Step 1:

[0507] A user accesses the system using a terminal and inputs a theme (e.g., "Friendship") and a character (e.g., "Rio the Rabbit"). The input theme and character settings are sent to the server by the terminal. The input data is sent in the form of a theme and a character.

[0508] Step 2:

[0509] The server analyzes the theme and character settings received from the device. It uses natural language processing (NLP) technology to parse the input data and understand its meaning. Based on the results of this analysis, it checks for inappropriate data by screening it against a list of prohibited words and for extreme content. As an output, it generates data that is deemed appropriate or that requires correction.

[0510] Step 3:

[0511] If the server determines that the filtered data is appropriate, it calls a generative AI model (e.g., GPT-3) based on the data to generate a story. At this time, the prompt text includes the theme and character information entered by the user. The input data is the theme and character settings, and the output data is the generated story.

[0512] Step 4:

[0513] The server converts the generated story into a text format and formats it, for example dividing the story into paragraphs, applying a particular font size and style, and adjusting line breaks and punctuation as needed. The input data is the generated story, and the output data is the formatted story text.

[0514] Step 5:

[0515] The server calls an image generation AI (e.g., DALL-E) to create illustrations corresponding to the generated story. By including a specific description of a particular scene in the story in the prompt, an illustration appropriate for that scene is generated. The input data is a description of the story scene, and the output data is the generated illustration.

[0516] Step 6:

[0517] The server integrates the generated illustrations and story into a coherent picture book format. This is the process of adjusting the placement of illustrations and text and determining the page layout. The input data is the story in text format and the generated illustrations, and the output data is the integrated digital picture book.

[0518] Step 7:

[0519] The server uses a translation engine (e.g., Google Translate API) to translate the generated story into the language specified by the user (e.g., English, French). It checks the translated content for grammar and makes any necessary corrections. The input data is the story in text format, and the output data is the translated story.

[0520] Step 8:

[0521] The server generates digital picture book data that integrates the translated story and illustrations and sends it to the user's terminal. The terminal provides an interface for displaying the received picture book data, allowing the user to view the picture book. The input data is the translated story and the generated illustrations, and the output data is the user's viewing interface.

[0522] Step 9:

[0523] Users can share the created picture book with other users using the sharing function of their device. Sharing methods include social networking sites, email, and cloud services. The input data is digital picture book data, and the output data is in a format that can be viewed by the recipient users.

[0524] (Application example 1)

[0525] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0526] In conventional story generation systems, it is common for users to set a theme and characters, and then generate a story based on those. However, they lack the functionality to compile the generated story and illustrations into a digital booklet in a consistent format, distribute it in multiple languages, or easily share it with other users. For this reason, there is a need for a system that increases users' freedom of expression and provides convenience in international multilingual support.

[0527] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0528] In this invention, the server includes means for receiving input data, means for filtering the received input data, means for generating a story based on the filtered data, means for generating images based on the generated story, means for translating the generated story into multiple languages, means for distributing the translated story to a terminal, means for sharing the story distributed to the terminal with other users, means for receiving theme and character settings from a user, means for automatically generating a story based on the received theme and characters, and means for integrating the story and images into a digital booklet format. This makes it possible to translate, distribute, and share educational and safe stories and corresponding illustrations into multiple languages ​​in a consistent digital booklet format based on the themes and characters set by the user.

[0529] The "means for receiving input data" is a function for receiving the theme and character information provided by the user via the terminal.

[0530] "Means for filtering received input data" refers to a function for filtering out inappropriate content from input data and narrowing it down to safe and appropriate data.

[0531] The "means for generating a story based on filtered data" is a function for automatically generating a story based on verified themes and character settings.

[0532] The "means for generating images based on the generated story" is a function for automatically generating illustrations according to the content of the story.

[0533] The "means for translating the generated story into multiple languages" is a function for automatically converting the generated story into multiple specified languages.

[0534] The "means for delivering the translated story to the terminal" is a function for transferring the translated story and corresponding illustrations to the user's terminal.

[0535] The "means for sharing a story delivered to the terminal with other users" is a function that enables a user to share a received story with other users via a social networking service or email.

[0536] The "means for accepting theme and character settings from the user" is a function that provides an input interface for the user to set the theme and character.

[0537] "Means for automatically generating a story based on the accepted theme and characters" is a function for automatically generating a story using a generative AI model based on the theme and character information set by the user.

[0538] "Means for integrating stories and images into a digital booklet format" is a function for combining the generated stories and illustrations and editing them into a coherent digital picture book format.

[0539] A "social networking service" is a platform for sharing information and interacting with other users over the Internet.

[0540] "Email" is a means of communication for sending and receiving text messages and files over the Internet.

[0541] The following describes an embodiment of the present invention. The system mainly consists of three elements: a server, a terminal, and a user. The specific roles and processing procedures of each element are described below.

[0542] System programs and their processing

[0543] User Theme and Character Settings

[0544] Users access the system through a device such as a smartphone or tablet and use an interface to input the theme (e.g., "Friendship") and character (e.g., "Rio the Rabbit") to set up the system. This information is then sent from the device to the server.

[0545] Filtering input data

[0546] The server analyzes the received themes and character settings, and screens them for banned words and explicit content. Scripts written in Python and Flask perform the filtering. If inappropriate data is detected, a message is sent to the user's device, prompting them to correct it.

[0547] Narrative Generation

[0548] Based on the filtered themes and character settings, the server invokes a generative AI model (e.g., GPT-4) to generate a story, which is then converted into text format and temporarily stored on the server.

[0549] Illustration generation

[0550] The server then uses image generation AI (e.g., Stable Diffusion or DALL-E) to generate illustrations that correspond to the generated story. These illustrations are aligned with the scenes in the story.

[0551] Multilingual Translation

[0552] The generated story is translated into the specified language (e.g., English, French) using the Google Translate API or DeepL API. The translated story is stored on the server and corrected as needed.

[0553] Story Distribution

[0554] The server then aggregates the translated stories and illustrations into a digital booklet and delivers it to users' devices, where they can view it through an application on their smartphones or tablets.

[0555] Share your story

[0556] Users can use the sharing function of their terminals to share the generated digital booklet with other users via social networking services or email.

[0557] Overview of the technologies and programs used

[0558] Server side: Python, Flask (or Django), a generative AI model (e.g. GPT-4), a translation engine (Google Translate API or DeepL).

[0559] Frontend: React Native (for mobile apps) or React.js (for web).

[0560] AI model: OpenAI's GPT-4, Stable Diffusion or DALL-E (illustration generation).

[0561] Database: PostgreSQL (storage of user data, generated stories, and illustrations).

[0562] Specific examples

[0563] For example, a child might create a story with the theme of "friendship" and featuring characters Rio the rabbit and Ken the turtle. The story is then generated and filtered by the server. Illustrations corresponding to the story are then generated, and a digital booklet translated into Japanese, English, and French is finally generated. This picture book is then delivered to the user's smartphone, and the user can share it with other users via social networking sites.

[0564] Generative AI model prompt example

[0565] prompt:

[0566] Theme: "Friendship"

[0567] Characters: "Rio the Rabbit" and "Ken the Turtle"

[0568] Instructions: Create an educational and safe story for children. The story should be based on the adventures of Rio the rabbit and Ken the turtle who team up to cross a big bridge. Write in a friendly, easy-to-understand voice.

[0569] In this way, the system of the present invention can effectively generate and provide safe and educational stories to users, which is expected to promote communication between users and contribute to the spread of education.

[0570] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0571] Step 1:

[0572] Users access the system through a device such as a smartphone or tablet and use the interface to input the theme (e.g., "Friendship") and character (e.g., "Rio the Rabbit") to set the game. The set information is sent from the device to the server as input data. The input is the theme and character setting data from the user's interface, and the output is the input data sent to the server.

[0573] Step 2:

[0574] The server filters the input data for themes and character settings received from the device. Specifically, it uses scripts using Python and Flask to screen for banned words and explicit content. The input is the theme and character setting data, and filtering is performed as data processing, with the output being the appropriate filtered data.

[0575] Step 3:

[0576] The server generates a story using a generative AI model (e.g., GPT-4) based on the filtered theme and character settings. The input is the filtered data, and a story is generated using the generative AI model as data calculation. The output is the text data of the generated story. The generated story is temporarily stored on the server.

[0577] Step 4:

[0578] The server uses an image generation AI (e.g., Stable Diffusion or DALL-E) based on the text data of the generated story to generate illustrations corresponding to each scene in the story. The input is the text data of the generated story, and the image generation AI is used as data calculation to generate illustrations, and the output is the generated illustration.

[0579] Step 5:

[0580] The server translates the generated story into multiple languages. Specifically, it uses the Google Translate API or DeepL API to translate the input data, which is the generated story text. The translation API is called as data processing, and the output is the translated multilingual story text data.

[0581] Step 6:

[0582] The server integrates the translated story and the generated illustrations into a digital booklet. The input is the text data of the translated story and the generated illustration data, which are integrated as data processing, and the output is the digital booklet data.

[0583] Step 7:

[0584] The server distributes the digital booklet data to the user's terminal. The input is the digital booklet data, which is distributed to the terminal as data transfer, and the output is the digital booklet displayed on the user's terminal.

[0585] Step 8:

[0586] The user uses the sharing function of the terminal to share the generated digital booklet with other users via social networking services or email. The input is the distributed digital booklet data, and the sharing function is used as data transfer, and the output is the digital booklet shared with other users.

[0587] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0588] This invention is a system that combines an emotion engine that recognizes the user's emotions, generates safe and educational stories, and translates and shares them in multiple languages. The program and processing of this system are explained in detail below. The system mainly works in cooperation with three elements: the server, the terminal, and the user.

[0589] Description of system programs and processes

[0590] User theme and character settings, and emotion recognition

[0591] 1. The user opens the terminal interface and inputs the theme (e.g., "Friendship") and character (e.g., "Rio the Rabbit").

[0592] 2. The device acquires emotion data from the user's facial expressions and voice and sends that data to the emotion engine.

[0593] 3. The emotion engine analyzes the acquired emotion data and recognizes the user's current emotion. The analysis results are reflected in the theme and character settings.

[0594] Filtering input data

[0595] 1. The server analyzes all input data, including the theme and character settings received from the device and the emotion data obtained from the emotion engine, and starts the filtering process. Specifically, it runs a screening process for banned words and explicit content.

[0596] 2. The server checks the filtering results and if any inappropriate content is found, it generates a message asking the user to correct it and sends it to the terminal. If there is no inappropriate content, it proceeds to the next step.

[0597] Story and illustration generation

[0598] 1. The server calls the generative AI model based on the verified theme, character settings, and emotional data to generate a story. For example, if the user is happy, it generates a story with a happy ending, depending on the user's emotions.

[0599] 2. The server converts the generated story into text format, and then uses image generation AI to generate illustrations corresponding to the story. Based on the emotional data, the illustrations also reflect emotional elements.

[0600] Multilingual Translation

[0601] 1. The server invokes a translation engine that translates the generated story into the language specified by the user (e.g., English, French).

[0602] 2. The server reviews the translated story and makes corrections if necessary.

[0603] Story Distribution

[0604] 1. The server generates digital picture book data that integrates the translated story and illustrations and sends it to the user's device.

[0605] 2. The terminal provides an interface for displaying the received digital picture book, allowing the user to view it.

[0606] Share your story

[0607] 1. The user shares the created digital picture book with other users using the sharing function of their device. Sharing methods include social networking sites, email, and cloud services.

[0608] 2. The device sends the picture book data to the sharing device so that other users can view it.

[0609] For example, if a parent or guardian selects characters Rio the Rabbit and Ken the Turtle as characters with the theme of "friendship" and the emotion they are feeling is "fun," a story about the fun adventures of Rio the Rabbit and Ken the Turtle is generated based on this emotion. The story is filtered, appropriate illustrations are generated, and a digital picture book translated into Japanese, English, and French is delivered to the parent or guardian's smartphone. The parent or guardian can then share the picture book with friends via social media.

[0610] As described above, the system of the present invention can provide more personalized educational resources by generating stories and illustrations that reflect the user's emotions, which is expected to promote communication between parents and children and contribute to reducing educational disparities.

[0611] The processing flow will be explained below.

[0612] Step 1:

[0613] The user opens the terminal interface and inputs the theme (e.g., "friendship") and character (e.g., "Rio the Rabbit").

[0614] Step 2:

[0615] The device captures the user's facial expressions and voice through the user's camera and microphone, and obtains emotional data in real time.

[0616] Step 3:

[0617] The device sends the acquired emotion data to the emotion engine, which analyzes the user's current emotion and identifies emotion categories such as "joy," "sadness," and "surprise."

[0618] Step 4:

[0619] The emotion engine returns the analysis results to the device, which then transmits the information to the server, thereby transmitting the user's emotion data to the server.

[0620] Step 5:

[0621] The server performs filtering based on the theme, character settings, and emotion data received from the device, specifically checking for taboo words and explicit content.

[0622] Step 6:

[0623] The server checks the filtering results, and if any inappropriate content is found, it generates a message requesting the user to correct the content and sends it to the terminal. If no inappropriate content is found, it proceeds to the next step.

[0624] Step 7:

[0625] The server then invokes a generative AI model based on verified themes, character settings, and emotional data to generate a story. For example, if the user expresses the emotion of "enjoyment," a story containing positive and enjoyable episodes will be generated.

[0626] Step 8:

[0627] The server converts the generated story into a text format, which can be modified as needed.

[0628] Step 9:

[0629] The server uses image generation AI to generate a corresponding illustration based on the story. The illustration also reflects the user's emotional data. For example, if the emotion of "fun" is conveyed, an illustration of a bright and cheerful scene is generated.

[0630] Step 10:

[0631] The server integrates the generated illustrations with the story to create a coherent digital picture book.

[0632] Step 11:

[0633] The server invokes a translation engine to translate the generated digital picture book into a language specified by the user (e.g., English, French).

[0634] Step 12:

[0635] The server checks the quality of the translated stories and makes corrections if necessary.

[0636] Step 13:

[0637] The server completes the digital picture book integrating the translation and illustrations and transmits the data to the user's terminal.

[0638] Step 14:

[0639] The terminal provides an interface for displaying the received digital picture book, allowing the user to view it.

[0640] Step 15:

[0641] The user can share the created digital picture book with other users using the sharing function of the device. Sharing methods include social networking sites, email, and cloud services.

[0642] Step 16:

[0643] The terminal transmits picture book data to a sharing terminal, so that other users can view the picture book data.

[0644] The above is the specific processing flow of the system of the present invention, which combines an emotion engine. This flow makes it possible to generate, distribute, and share personalized stories that reflect the user's emotions.

[0645] Example 2

[0646] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0647] Current story generation systems struggle to generate content that reflects the user's emotions, resulting in a lack of personalized educational resources. Furthermore, they lack the ability to automatically translate the generated stories and images into multiple languages ​​to match the user's emotions and themes, making them unable to meet the needs of global users. Furthermore, they have yet to provide digital content in a format that can be easily shared among users.

[0648] The specification process by the specification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving input data, means for filtering the received input data, means for generating a story based on the filtered data, means for generating images based on the generated story, means for translating the generated story into multiple languages, means for delivering the translated story to the terminal, means for sharing the story delivered to the terminal with other users, means for acquiring and analyzing user emotion data, and means for generating a story and images based on the acquired emotion data. This makes it possible to generate a personalized story and images based on the user's emotions, translate the story into multiple languages, and provide digital content that can be easily shared among users.

[0649] The "means for receiving input data" is a method by which a user inputs theme and character settings and transmits them to the system.

[0650] A "means for filtering received input data" is a method for analyzing received data and removing inappropriate content.

[0651] A "means for generating a story" is a method for automatically creating a story based on the filtered data and the user's emotional data.

[0652] An "image generation means" is a method for creating related visual content based on the generated narrative.

[0653] A "means for translating stories" is a method for translating the generated stories into multiple languages.

[0654] A "means for delivering a story to a terminal" is a method for transferring a translated story to a user's device.

[0655] "Means for sharing stories with other users" refers to methods that provide an interface or protocol for sharing distributed stories.

[0656] "Means for acquiring and analyzing emotion data" refers to a method for collecting and analyzing emotions from the user's facial expressions, voice, etc.

[0657] The "means for generating a story and images based on the acquired emotion data" is a method for generating a story and images using the analyzed emotion data.

[0658] The present invention is a system that combines an emotion engine that recognizes a user's emotions, generates safe and educational stories, and translates and shares them in multiple languages. Specific embodiments of this system and their operation are described in detail below.

[0659] System configuration

[0660] The system includes the following elements:

[0661] 1. Device: A device on which users input theme and character settings and acquire facial and voice data. This includes smartphones, tablets, and PCs.

[0662] 2. Server: A central processing unit that filters input data, generates stories and images, translates, and distributes data.

[0663] 3. Emotion engine: Software for acquiring and analyzing emotional data from the user's facial expressions and voice. Uses an API for emotion analysis.

[0664] 4. Generative AI model: An artificial intelligence model for generating stories and images based on user input and emotional data. Specifically, it uses large-scale language models and image generation models such as GPT (Generative Pre-trained Transformer).

[0665] How it works

[0666] User theme and character settings, and emotion recognition

[0667] The user opens the device's interface and inputs a theme (e.g., "Friendship") and a character (e.g., "Rio the Rabbit"). The device captures the user's facial expressions and voice through a camera and microphone, and sends the data to the emotion engine. The emotion engine analyzes the acquired emotion data and recognizes the user's current emotion. The results of this analysis are sent to the server and reflected in the theme and character settings.

[0668] Filtering input data

[0669] The server analyzes all input data, including the theme and character settings received from the device and the emotional data obtained from the emotion engine, and then begins the filtering process, specifically running a screening process for banned words and explicit content.

[0670] Story and illustration generation

[0671] The server generates a story by calling a generative AI model based on verified themes, character settings, and emotional data. For example, if the user is happy, it generates a story with a happy ending.

[0672] Example prompt sentence:

[0673] "Generate a story about friendship, featuring Rio the rabbit and Ken the turtle going on a fun adventure. The user emotion is 'fun'."

[0674] The server then converts the generated story into text format and uses image generation AI to generate illustrations that correspond to the story. The generated illustrations also reflect emotional elements based on the emotional data.

[0675] Example prompt sentence:

[0676] "Generate an illustration that reflects the emotion of enjoyment based on the following story: 'Rio the rabbit and Ken the turtle go on a fun adventure.'"

[0677] Multilingual Translation

[0678] The server calls a translation engine to translate the generated story into the language specified by the user (e.g., English, French). It checks the translated text obtained from the translation engine and makes corrections if necessary.

[0679] Story Distribution

[0680] The server generates digital picture book data that integrates the translated story and illustrations and sends it to the user's device, which provides an interface for displaying the received digital picture book so that the user can view it.

[0681] Share your story

[0682] Users can share the digital picture book they have created with other users using the sharing function of their device. Sharing methods include social networking sites, email, and cloud services. The device sends the picture book data to the destination device, allowing other users to view it.

[0683] Specific examples

[0684] For example, if a parent or guardian selects characters Rio the Rabbit and Ken the Turtle as characters with the theme of "friendship" and the emotion at that time is "fun," a story about the fun adventures of Rio the Rabbit and Ken the Turtle is generated based on this emotion. This story is filtered, appropriate illustrations are generated, and a digital picture book translated into Japanese, English, and French is delivered to the parent or guardian's smartphone. The parent or guardian can then share this picture book with friends via social media.

[0685] This invention can provide more personalized educational resources by generating stories and illustrations that reflect the user's emotions, which is expected to promote communication between parents and children and contribute to reducing educational disparities.

[0686] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0687] Step 1:

[0688] The user opens the device interface and inputs the theme (e.g., "Friendship") and character (e.g., "Rio the Rabbit"), and when the user confirms the input data, the device sends the data to the server.

[0689] Input: Input data for the theme "Friendship" and the character "Rabbit Rio"

[0690] Output: The input data is sent to the server

[0691] Step 2:

[0692] The device acquires the user's facial expressions and voice and sends them to the emotion engine. The emotion engine analyzes the facial and voice data to recognize the user's current emotion. The result is then sent to the server.

[0693] Input: User's facial expression data and voice data

[0694] Output: Recognized emotion data

[0695] Step 3:

[0696] The server receives the input data (theme and character settings) and analyzed emotional data, and performs filtering. Specifically, it screens for prohibited words and extreme content. Once filtering is complete, it notifies the user of the results.

[0697] Input: Theme "Friendship", Character "Rabbit Rio", Emotion data

[0698] Output: Filtering results and a message requesting corrections if necessary

[0699] Step 4:

[0700] The server invokes the generative AI model based on the verified input data and emotion data to generate a story. For example, if the user's emotion is "fun," it generates a story with a happy ending.

[0701] Input: Filtered data and sentiment data

[0702] Output: Generated narrative text

[0703] Specific operation example:

[0704] The server sends a prompt to the AI ​​model, such as "Generate a story about friendship, with Rio the rabbit and Ken the turtle going on a fun adventure. The user's emotion is 'fun'."

[0705] Step 5:

[0706] The server converts the generated story text into a text format and calls an image generation AI to generate illustrations corresponding to the story, which reflect the user's emotions.

[0707] Input: Generated narrative text

[0708] Output: Generated illustration

[0709] Specific operation example:

[0710] The prompt text "Generate an illustration that reflects the emotion of enjoyment based on the following story: 'Rio the rabbit and Ken the turtle go on a fun adventure'" is sent to the image generation AI.

[0711] Step 6:

[0712] The server invokes a translation engine that translates the generated story text into the language specified by the user (e.g., English, French). The translated text is returned to the server, which reviews it and makes corrections if necessary.

[0713] Input: Japanese story text

[0714] Output: Translated story text

[0715] Step 7:

[0716] The server integrates the translated story and illustrations to create digital picture book data, which is then sent to the user's device.

[0717] Input: translated story text, generated illustrations

[0718] Output: Digital picture book data

[0719] Step 8:

[0720] An interface for displaying the digital picture book received by the terminal is provided, allowing the user to view it.

[0721] Input: Digital picture book data

[0722] Output: A displayed digital picture book

[0723] Step 9:

[0724] Users can share the digital picture book they have created with other users using the sharing function of their device. Sharing methods include social networking sites, email, and cloud services.

[0725] Input: Digital picture book data

[0726] Output: Digital picture books shared via social media, email, and the cloud

[0727] Step 10:

[0728] The terminal transmits picture book data to a sharing device so that other users can view it.

[0729] Input: Share request

[0730] Output: Digital picture book data transferred to the sharing device

[0731] (Application example 2)

[0732] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0733] Conventional content generation and sharing systems were unable to generate and provide individual stories that took the user's emotions into consideration. This made it difficult to provide personalized content that responded to the user's emotions, and the educational and entertainment benefits were not fully realized. Furthermore, the inability to translate and share content in multiple languages ​​limited global sharing. This created challenges that made it difficult to improve the user experience and correct educational disparities.

[0734] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0735] In this invention, the server includes means for receiving input data and emotional data, means for filtering the received input data and emotional data, means for generating a story based on the filtered data and emotional data, means for generating images based on the generated story, means for translating the generated story into multiple languages, means for distributing the translated story to a terminal, means for sharing the story distributed to the terminal with other users, means for accepting theme and character settings and emotional data from a user, means for automatically generating a story based on the accepted theme and characters, means for personalizing the story based on the accepted emotional data, and means for simultaneously generating a story and images, integrating them, evaluating them based on the user's emotional data, and ensuring quality. This enables the generation of personalized stories and illustrations that take user emotions into consideration, realizes multilingual translation and easy sharing, and enables an improved user experience and mitigation of educational disparities.

[0736] "Input Data" refers to the theme, characters, and other setting information that a user provides to the system.

[0737] "Emotion data" refers to information about the emotional state of a user obtained from facial expressions, voice, etc.

[0738] "Filtering" refers to the process of removing inappropriate content or prohibited words from received data.

[0739] "Story" refers to a document or text that is generated based on themes, characters, and user emotional data.

[0740] "Images" refers to illustrations and visual content generated based on the story content and user emotional data.

[0741] "Translation" refers to the process of converting a generated story into multiple languages.

[0742] "Terminal" refers to an electronic device used by a user, such as a smartphone, tablet, or computer.

[0743] "Sharing" refers to the process of providing or distributing the generated stories and images to other users and platforms.

[0744] "Theme" refers to a concept that indicates the main point or main point of a story.

[0745] "Characters" refers to characters that appear in stories or animated characters.

[0746] "Personalization" refers to the process of adapting content based on a user's individual emotional data and preferences.

[0747] "Synthesis" refers to the process of bringing together the generated stories and images into a single piece of content.

[0748] "Evaluation" refers to the process used to judge the quality and appropriateness of the stories and images produced.

[0749] "Quality assurance" refers to the process of verifying that the generated content is appropriate and guaranteeing its quality.

[0750] The present invention is a system that combines an emotion engine that recognizes a user's emotions, generates safe and educational stories, and translates and shares them in multiple languages. An embodiment of this system will be described in detail below.

[0751] Description of system programs and processes

[0752] User theme and character settings, and emotion recognition

[0753] The user opens the device's interface and inputs the theme and character. For example, the user may select the theme "friendship" and "Rio the rabbit" as the character. Emotional data is also acquired from the user's facial expressions and voice. The device sends this data to an emotion engine, which analyzes it to recognize the user's current emotion. The results of this analysis are reflected in the theme and character settings.

[0754] Filtering input data

[0755] The server analyzes and filters all input data, including the theme and character settings received from the device and the emotion data obtained from the emotion engine. Specifically, a screening process is carried out for a list of prohibited words and explicit content. If inappropriate content is included, the server generates a message requesting the user to correct it and sends it to the device. If there is no inappropriate content, the process proceeds to the next step.

[0756] Story and illustration generation

[0757] The server generates a story by calling a generative AI model based on verified themes, character settings, and emotional data. For example, if the user is happy, a story with a happy ending is generated. This story is converted into text format, and then an image generation AI is used to generate illustrations corresponding to the story. Based on the emotional data, the illustrations also reflect emotional elements.

[0758] Multilingual Translation

[0759] The server invokes a translation engine to translate the generated story into a language specified by the user, such as Japanese, English, French, etc. The translated story is then reviewed by the server and any necessary corrections are made.

[0760] Story Distribution

[0761] The server generates digital picture book data that integrates the translated story and illustrations and transmits it to the user's terminal, which provides an interface for displaying the received digital picture book so that the user can view it.

[0762] Share your story

[0763] Users can share the digital picture book they have created with other users using the sharing function of their device. Sharing methods include social networking sites, email, and cloud services. The device sends the picture book data to the destination device, allowing other users to view it.

[0764] Hardware and software used

[0765] The server runs software such as generative AI models, emotion recognition engines, image generation AI, translation engines, and data filtering systems. Specific software and hardware used include the Google Translate API, EmotionRecognizer, StoryGenerator, and ImageGenerator. These work together to enable the generation, translation, and sharing of stories and illustrations based on user requests.

[0766] Specific examples

[0767] For example, if a parent or guardian selects characters Rio the Rabbit and Ken the Turtle as characters with the theme of "friendship" and the emotion at that time is "fun," a story about the fun adventures of Rio the Rabbit and Ken the Turtle is generated based on this emotion. The story is filtered and appropriate illustrations are generated, and then a digital picture book translated into Japanese, English, and French is delivered to the parent or guardian's smartphone. The parent or guardian can share this picture book with friends via social media.

[0768] Prompt Sentence Examples

[0769] Theme: Friendship

[0770] Character: Rio the Rabbit

[0771] Emotion data: {'emotion': 'joy'}

[0772] prompt:

[0773] "Rio the rabbit was having a great time playing with his friends. They decided to go on a new adventure together. Along the way, Rio..."

[0774] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0775] Step 1:

[0776] The user opens the device interface and inputs the theme and character. For example, the theme is set to "friendship" and the character is set to "Rio the rabbit." The user's facial expressions and voice data are also input. This input data is sent to the device (input data and emotion data).

[0777] Step 2:

[0778] The device processes the received theme, character, and emotion data and sends it to the server. The emotion data includes the results of the user's facial expression analysis and voice analysis. The device calls the emotion recognition engine to identify the user's emotion and sends the result as data (input data: theme, character, emotion data / output data: theme, character, emotion analysis result).

[0779] Step 3:

[0780] The server analyzes the theme, character, and emotion data received from the device and performs a filtering process. The filtering process screens the data based on a list of prohibited words and extreme content to check for inappropriate content. If inappropriate content is found, a message is generated requesting the user to correct it (input data: theme, character, emotion analysis results / output data: filtered data or correction request message).

[0781] Step 4:

[0782] Based on the filtered data, the server calls a generative AI model to generate a story. The generative AI model adjusts the tone and ending of the story based on the emotional data. For example, if the emotional data indicates joy, a story with a happy ending will be generated (input data: filtered data / output data: generated story).

[0783] Step 5:

[0784] Next, the server calls an image generation AI based on the generated story, which generates illustrations that correspond to the content of the story. Emotional data is also taken into account, and emotional elements are reflected in the illustrations (input data: generated story, emotion data / output data: generated illustrations).

[0785] Step 6:

[0786] The server invokes a translation engine to translate the generated story into multiple languages, for example, from Japanese to English and French. The translation is done automatically and can be manually corrected if necessary (input data: generated story / output data: translated story).

[0787] Step 7:

[0788] The server generates digital picture book data that integrates the translated story and generated illustrations and sends it to the user's device. The device provides an interface for displaying the received digital picture book, allowing the user to easily browse it (input data: translated story, generated illustrations / output data: digital picture book data).

[0789] Step 8:

[0790] Users can share the digital picture book they have created with other users using the sharing function of their device. Social networking sites, email, cloud services, and other methods of sharing are available. The device sends the picture book data to the destination device, allowing other users to view it (input data: digital picture book data; output data: data sent to the destination device).

[0791] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0792] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0793] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0794] [Third embodiment]

[0795] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0796] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0797] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0798] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0799] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0800] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0801] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0802] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0803] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0804] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0805] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0806] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0807] This system allows users to create safe and educational stories by setting themes and characters, and then translates and shares them in multiple languages. The program and processing of this system are described in detail below. The system mainly works in conjunction with three elements: the server, the terminal, and the user.

[0808] Description of system programs and processes

[0809] User Theme and Character Settings

[0810] 1. A user accesses the system through a terminal and inputs a theme (e.g., "friendship") and a character (e.g., "Rio the rabbit") using the input interface.

[0811] 2. The device sends the entered theme and character settings to the server.

[0812] Filtering input data

[0813] 1. The server analyzes and filters the theme and character settings received from the device, specifically by screening for prohibited words and explicit content.

[0814] 2. If the server determines that the input data is appropriate based on the filtering results, it proceeds to the next step using that data. If the data contains inappropriate data, it sends a message to the terminal requesting the user to correct it.

[0815] Narrative Generation

[0816] 1. The server calls a generative AI model based on the verified theme and character settings to generate a story. For example, it generates a story in which "Rio the Rabbit" cooperates with his friend "Ken the Turtle" to cross a large bridge.

[0817] 2. The server converts the generated story into a specified format, such as text format.

[0818] Illustration generation

[0819] 1. The server uses image generation AI to create illustrations that correspond to the generated story. Multiple illustrations that match the scenes in the story are generated.

[0820] 2. The server integrates the generated illustrations and story into a coherent picture book format.

[0821] Multilingual Translation

[0822] 1. The server invokes a translation engine to translate the generated story into the language specified by the user (e.g., English, French).

[0823] 2. The server reviews the translated story and makes any necessary corrections.

[0824] Story Distribution

[0825] 1. The server generates digital picture book data that integrates the translated story and illustrations and sends it to the user's device.

[0826] 2. The terminal provides an interface that displays the received picture book data to the user, allowing the user to view the picture book.

[0827] Share your story

[0828] 1. The user shares the created picture book with other users using the sharing function of the device. Sharing methods include social networking sites, email, and cloud services.

[0829] 2. The device sends the picture book data to the sharing device so that other users can view it.

[0830] For example, a parent or guardian can create a story with the theme of "friendship" and featuring characters Rio the Rabbit and Ken the Turtle. The story is then generated and filtered by the server. Illustrations corresponding to the story are then generated, and a digital picture book translated into Japanese, English, and French is finally generated. The picture book is then delivered to the parent's smartphone, and the parent can share it with other parents via social networking sites.

[0831] As described above, the system of the present invention can effectively generate and provide safe and educational stories to children, which is expected to promote communication between parents and children and contribute to reducing educational disparities.

[0832] The processing flow will be explained below.

[0833] Step 1:

[0834] The user opens the terminal interface and inputs the theme (e.g., "friendship") and character (e.g., "Rio the Rabbit").

[0835] Step 2:

[0836] The terminal receives the entered theme and character settings and transmits this data to the server.

[0837] Step 3:

[0838] The server analyzes the theme and character settings received from the device and begins the filtering process, specifically running a screening process for banned words and explicit content.

[0839] Step 4:

[0840] The server checks the filtering results, and if any inappropriate content is found, it generates a message requesting the user to correct the content and sends it to the terminal. If no inappropriate content is found, it proceeds to the next step.

[0841] Step 5:

[0842] The server calls the generative AI model based on the verified theme and character settings to generate a story. For example, it generates a story in which Rio the rabbit cooperates with his friend Ken the turtle to cross a large bridge.

[0843] Step 6:

[0844] The server converts the generated story into text format and then uses image generation AI to begin generating illustrations that correspond to the story.

[0845] Step 7:

[0846] The server integrates the generated illustrations with the story to create a coherent digital picture book.

[0847] Step 8:

[0848] The server executes a translation process to translate the generated digital picture book into a user-specified language (e.g., English, French).

[0849] Step 9:

[0850] The server checks the quality of the translated story, makes corrections if necessary, and finally generates the finished digital picture book.

[0851] Step 10:

[0852] The server transmits the completed digital picture book data to the user's terminal.

[0853] Step 11:

[0854] The terminal provides an interface for displaying the received digital picture book, allowing the user to view it.

[0855] Step 12:

[0856] The user shares the created digital picture book with other users using the sharing function of the device. For example, the picture book data is sent via social networking sites, email, or cloud services.

[0857] Step 13:

[0858] The terminal transmits picture book data to a sharing terminal, so that other users can view the picture book data.

[0859] The above is the specific processing flow of the system of the present invention, which enables safe and educational stories to be created, distributed, and shared.

[0860] Example 1

[0861] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0862] Previous story generation systems were complicated in the process of generating safe and educational stories based on themes and characters set by the user, and lacked a means to translate and share the generated stories in multiple languages. Furthermore, there were no systems that automatically generated illustrations along with the story and integrated them. As a result, users had to manually create the story, illustrations, and translation, which required a great deal of effort and time.

[0863] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0864] In this invention, the server includes an input means for a user to set a theme and characters, a means for receiving input data of the set theme and characters and analyzing and filtering the input data, a means for generating a story using a generative AI model based on the filtered data, a means for converting the generated story into a text format, a means for generating illustrations using an image generation AI model based on the generated story, a means for integrating the generated story and illustrations into a consistent format, a means for translating the generated story into multiple languages, a means for delivering the translated story to a terminal, and a means for sharing the story delivered to the terminal with other users. This makes it possible to effectively generate safe and educational stories based on themes and characters set by users, translate them into multiple languages, and share them.

[0865] A "user" is a person or entity that accesses the system and configures themes and characters.

[0866] A "terminal" is an electronic device used by a user, such as a computer, smartphone, or tablet.

[0867] "Server" means the central control unit that manages and executes the functions of the entire system.

[0868] A "theme" is the main theme or central concept of a story.

[0869] A "character" is a person, animal, or fictional being that appears in a story.

[0870] "Input means" refers to an interface that allows a user to input themes and characters into the system.

[0871] A "generative AI model" is an artificial intelligence model that generates stories based on themes and characters set by the user.

[0872] "Filtering" is the process of analyzing input data and eliminating inappropriate content.

[0873] "Text format" is a data format that expresses the generated story in a sentence format.

[0874] An "image generation AI model" is an artificial intelligence model for generating illustrations that correspond to a story.

[0875] A "translation engine" is software or a service that translates generated stories into other languages.

[0876] A "consistent format" means that the story and illustrations are in harmony and have a sense of unity.

[0877] "Distribution means" refers to a mechanism or method for transmitting the generated content to the user's terminal.

[0878] "Sharing means" refers to a function or method for sharing the generated content with other users.

[0879] This system allows users to create safe and educational stories by setting themes and characters, and then translates and shares them in multiple languages. This system operates in cooperation with three main elements: a server, a terminal, and users.

[0880] Users access the system using a terminal and set a theme (e.g., "Friendship") and a character (e.g., "Rio the Rabbit"). The terminal then sends the theme and character settings entered by the user to the server. The server analyzes the received data and filters it using natural language processing technology (NLP). Specifically, it screens for prohibited words and extreme content, and if inappropriate data is included, it sends a message to the terminal requesting the user to correct it.

[0881] The server then uses a generative AI model (e.g., GPT-3) to generate a story based on the filtered data. The generated story is converted into text format and formatted as needed. The server then uses an image generation AI (e.g., DALL-E) to create illustrations corresponding to each scene in the generated story. The prompt text also includes a description of the specific scene. The generated illustrations are integrated into the story and compiled into a coherent picture book format.

[0882] The server then uses a translation engine (e.g., Google Translate API) to translate the generated story into the user's specified language (e.g., English, French). The translation result undergoes grammar checks and style guide application, correcting any necessary parts. The server then sends the translated story and illustrations integrated into the digital picture book data to the user's device. The device then provides an interface for displaying the received picture book data, allowing the user to view the picture book.

[0883] As a concrete example, a user sets the characters "Rio the Rabbit" and "Ken the Turtle" as characters with the theme of "friendship." The server filters this and uses a generative AI model to generate a story in which "Rio the Rabbit" and "Ken the Turtle" cooperate to cross a bridge. Next, DALL-E is used to generate illustrations of the scene crossing the bridge. Finally, the generated story is translated into Japanese, English, and French and delivered to the user's smartphone as a digital picture book.

[0884] Example prompt sentence:

[0885] Theme: "Friendship"

[0886] Characters: "Rio the Rabbit" and "Ken the Turtle"

[0887] Objective: To generate stories that have educational value for children.

[0888] The system of the present invention can effectively generate safe and educational stories based on themes and characters set by the user, and can translate and share them in multiple languages, which is expected to promote communication between parents and children and contribute to reducing educational disparities.

[0889] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0890] Step 1:

[0891] A user accesses the system using a terminal and inputs a theme (e.g., "Friendship") and a character (e.g., "Rio the Rabbit"). The input theme and character settings are sent to the server by the terminal. The input data is sent in the form of a theme and a character.

[0892] Step 2:

[0893] The server analyzes the theme and character settings received from the device. It uses natural language processing (NLP) technology to parse the input data and understand its meaning. Based on the results of this analysis, it checks for inappropriate data by screening it against a list of prohibited words and for extreme content. As an output, it generates data that is deemed appropriate or that requires correction.

[0894] Step 3:

[0895] If the server determines that the filtered data is appropriate, it calls a generative AI model (e.g., GPT-3) based on the data to generate a story. At this time, the prompt text includes the theme and character information entered by the user. The input data is the theme and character settings, and the output data is the generated story.

[0896] Step 4:

[0897] The server converts the generated story into a text format and formats it, for example dividing the story into paragraphs, applying a particular font size and style, and adjusting line breaks and punctuation as needed. The input data is the generated story, and the output data is the formatted story text.

[0898] Step 5:

[0899] The server calls an image generation AI (e.g., DALL-E) to create illustrations corresponding to the generated story. By including a specific description of a particular scene in the story in the prompt, an illustration appropriate for that scene is generated. The input data is a description of the story scene, and the output data is the generated illustration.

[0900] Step 6:

[0901] The server integrates the generated illustrations and story into a coherent picture book format. This is the process of adjusting the placement of illustrations and text and determining the page layout. The input data is the story in text format and the generated illustrations, and the output data is the integrated digital picture book.

[0902] Step 7:

[0903] The server uses a translation engine (e.g., Google Translate API) to translate the generated story into the language specified by the user (e.g., English, French). It checks the translated content for grammar and makes any necessary corrections. The input data is the story in text format, and the output data is the translated story.

[0904] Step 8:

[0905] The server generates digital picture book data that integrates the translated story and illustrations and sends it to the user's terminal. The terminal provides an interface for displaying the received picture book data, allowing the user to view the picture book. The input data is the translated story and the generated illustrations, and the output data is the user's viewing interface.

[0906] Step 9:

[0907] Users can share the created picture book with other users using the sharing function of their device. Sharing methods include social networking sites, email, and cloud services. The input data is digital picture book data, and the output data is in a format that can be viewed by the recipient users.

[0908] (Application example 1)

[0909] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0910] In conventional story generation systems, it is common for users to set a theme and characters, and then generate a story based on those. However, they lack the functionality to compile the generated story and illustrations into a digital booklet in a consistent format, distribute it in multiple languages, or easily share it with other users. For this reason, there is a need for a system that increases users' freedom of expression and provides convenience in international multilingual support.

[0911] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0912] In this invention, the server includes means for receiving input data, means for filtering the received input data, means for generating a story based on the filtered data, means for generating images based on the generated story, means for translating the generated story into multiple languages, means for distributing the translated story to a terminal, means for sharing the story distributed to the terminal with other users, means for receiving theme and character settings from a user, means for automatically generating a story based on the received theme and characters, and means for integrating the story and images into a digital booklet format. This makes it possible to translate, distribute, and share educational and safe stories and corresponding illustrations into multiple languages ​​in a consistent digital booklet format based on the themes and characters set by the user.

[0913] The "means for receiving input data" is a function for receiving the theme and character information provided by the user via the terminal.

[0914] "Means for filtering received input data" refers to a function for filtering out inappropriate content from input data and narrowing it down to safe and appropriate data.

[0915] The "means for generating a story based on filtered data" is a function for automatically generating a story based on verified themes and character settings.

[0916] The "means for generating images based on the generated story" is a function for automatically generating illustrations according to the content of the story.

[0917] The "means for translating the generated story into multiple languages" is a function for automatically converting the generated story into multiple specified languages.

[0918] The "means for delivering the translated story to the terminal" is a function for transferring the translated story and corresponding illustrations to the user's terminal.

[0919] The "means for sharing a story delivered to the terminal with other users" is a function that enables a user to share a received story with other users via a social networking service or email.

[0920] The "means for accepting theme and character settings from the user" is a function that provides an input interface for the user to set the theme and character.

[0921] "Means for automatically generating a story based on the accepted theme and characters" is a function for automatically generating a story using a generative AI model based on the theme and character information set by the user.

[0922] "Means for integrating stories and images into a digital booklet format" is a function for combining the generated stories and illustrations and editing them into a coherent digital picture book format.

[0923] A "social networking service" is a platform for sharing information and interacting with other users over the Internet.

[0924] "Email" is a means of communication for sending and receiving text messages and files over the Internet.

[0925] The following describes an embodiment of the present invention. The system mainly consists of three elements: a server, a terminal, and a user. The specific roles and processing procedures of each element are described below.

[0926] System programs and their processing

[0927] User Theme and Character Settings

[0928] Users access the system through a device such as a smartphone or tablet and use an interface to input the theme (e.g., "Friendship") and character (e.g., "Rio the Rabbit") to set up the system. This information is then sent from the device to the server.

[0929] Filtering input data

[0930] The server analyzes the received themes and character settings, and screens them for banned words and explicit content. Scripts written in Python and Flask perform the filtering. If inappropriate data is detected, a message is sent to the user's device, prompting them to correct it.

[0931] Narrative Generation

[0932] Based on the filtered themes and character settings, the server invokes a generative AI model (e.g., GPT-4) to generate a story, which is then converted into text format and temporarily stored on the server.

[0933] Illustration generation

[0934] The server then uses image generation AI (e.g., Stable Diffusion or DALL-E) to generate illustrations that correspond to the generated story. These illustrations are aligned with the scenes in the story.

[0935] Multilingual Translation

[0936] The generated story is translated into the specified language (e.g., English, French) using the Google Translate API or DeepL API. The translated story is stored on the server and corrected as needed.

[0937] Story Distribution

[0938] The server then aggregates the translated stories and illustrations into a digital booklet and delivers it to users' devices, where they can view it through an application on their smartphones or tablets.

[0939] Share your story

[0940] Users can use the sharing function of their terminals to share the generated digital booklet with other users via social networking services or email.

[0941] Overview of the technologies and programs used

[0942] Server side: Python, Flask (or Django), a generative AI model (e.g. GPT-4), a translation engine (Google Translate API or DeepL).

[0943] Frontend: React Native (for mobile apps) or React.js (for web).

[0944] AI model: OpenAI's GPT-4, Stable Diffusion or DALL-E (illustration generation).

[0945] Database: PostgreSQL (storage of user data, generated stories, and illustrations).

[0946] Specific examples

[0947] For example, a child might create a story with the theme of "friendship" and featuring characters Rio the rabbit and Ken the turtle. The story is then generated and filtered by the server. Illustrations corresponding to the story are then generated, and a digital booklet translated into Japanese, English, and French is finally generated. This picture book is then delivered to the user's smartphone, and the user can share it with other users via social networking sites.

[0948] Generative AI model prompt example

[0949] prompt:

[0950] Theme: "Friendship"

[0951] Characters: "Rio the Rabbit" and "Ken the Turtle"

[0952] Instructions: Create an educational and safe story for children. The story should be based on the adventures of Rio the rabbit and Ken the turtle who team up to cross a big bridge. Write in a friendly, easy-to-understand voice.

[0953] In this way, the system of the present invention can effectively generate and provide safe and educational stories to users, which is expected to promote communication between users and contribute to the spread of education.

[0954] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0955] Step 1:

[0956] Users access the system through a device such as a smartphone or tablet and use the interface to input the theme (e.g., "Friendship") and character (e.g., "Rio the Rabbit") to set the game. The set information is sent from the device to the server as input data. The input is the theme and character setting data from the user's interface, and the output is the input data sent to the server.

[0957] Step 2:

[0958] The server filters the input data for themes and character settings received from the device. Specifically, it uses scripts using Python and Flask to screen for banned words and explicit content. The input is the theme and character setting data, and filtering is performed as data processing, with the output being the appropriate filtered data.

[0959] Step 3:

[0960] The server generates a story using a generative AI model (e.g., GPT-4) based on the filtered theme and character settings. The input is the filtered data, and a story is generated using the generative AI model as data calculation. The output is the text data of the generated story. The generated story is temporarily stored on the server.

[0961] Step 4:

[0962] The server uses an image generation AI (e.g., Stable Diffusion or DALL-E) based on the text data of the generated story to generate illustrations corresponding to each scene in the story. The input is the text data of the generated story, and the image generation AI is used as data calculation to generate illustrations, and the output is the generated illustration.

[0963] Step 5:

[0964] The server translates the generated story into multiple languages. Specifically, it uses the Google Translate API or DeepL API to translate the input data, which is the generated story text. The translation API is called as data processing, and the output is the translated multilingual story text data.

[0965] Step 6:

[0966] The server integrates the translated story and the generated illustrations into a digital booklet. The input is the text data of the translated story and the generated illustration data, which are integrated as data processing, and the output is the digital booklet data.

[0967] Step 7:

[0968] The server distributes the digital booklet data to the user's terminal. The input is the digital booklet data, which is distributed to the terminal as data transfer, and the output is the digital booklet displayed on the user's terminal.

[0969] Step 8:

[0970] The user uses the sharing function of the terminal to share the generated digital booklet with other users via social networking services or email. The input is the distributed digital booklet data, and the sharing function is used as data transfer, and the output is the digital booklet shared with other users.

[0971] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0972] This invention is a system that combines an emotion engine that recognizes the user's emotions, generates safe and educational stories, and translates and shares them in multiple languages. The program and processing of this system are explained in detail below. The system mainly works in cooperation with three elements: the server, the terminal, and the user.

[0973] Description of system programs and processes

[0974] User theme and character settings, and emotion recognition

[0975] 1. The user opens the terminal interface and inputs the theme (e.g., "Friendship") and character (e.g., "Rio the Rabbit").

[0976] 2. The device acquires emotion data from the user's facial expressions and voice and sends that data to the emotion engine.

[0977] 3. The emotion engine analyzes the acquired emotion data and recognizes the user's current emotion. The analysis results are reflected in the theme and character settings.

[0978] Filtering input data

[0979] 1. The server analyzes all input data, including the theme and character settings received from the device and the emotion data obtained from the emotion engine, and starts the filtering process. Specifically, it runs a screening process for banned words and explicit content.

[0980] 2. The server checks the filtering results and if any inappropriate content is found, it generates a message asking the user to correct it and sends it to the terminal. If there is no inappropriate content, it proceeds to the next step.

[0981] Story and illustration generation

[0982] 1. The server calls the generative AI model based on the verified theme, character settings, and emotional data to generate a story. For example, if the user is happy, it generates a story with a happy ending, depending on the user's emotions.

[0983] 2. The server converts the generated story into text format, and then uses image generation AI to generate illustrations corresponding to the story. Based on the emotional data, the illustrations also reflect emotional elements.

[0984] Multilingual Translation

[0985] 1. The server invokes a translation engine that translates the generated story into the language specified by the user (e.g., English, French).

[0986] 2. The server reviews the translated story and makes corrections if necessary.

[0987] Story Distribution

[0988] 1. The server generates digital picture book data that integrates the translated story and illustrations and sends it to the user's device.

[0989] 2. The terminal provides an interface for displaying the received digital picture book, allowing the user to view it.

[0990] Share your story

[0991] 1. The user shares the created digital picture book with other users using the sharing function of their device. Sharing methods include social networking sites, email, and cloud services.

[0992] 2. The device sends the picture book data to the sharing device so that other users can view it.

[0993] For example, if a parent or guardian selects characters Rio the Rabbit and Ken the Turtle as characters with the theme of "friendship" and the emotion they are feeling is "fun," a story about the fun adventures of Rio the Rabbit and Ken the Turtle is generated based on this emotion. The story is filtered, appropriate illustrations are generated, and a digital picture book translated into Japanese, English, and French is delivered to the parent or guardian's smartphone. The parent or guardian can then share the picture book with friends via social media.

[0994] As described above, the system of the present invention can provide more personalized educational resources by generating stories and illustrations that reflect the user's emotions, which is expected to promote communication between parents and children and contribute to reducing educational disparities.

[0995] The processing flow will be explained below.

[0996] Step 1:

[0997] The user opens the terminal interface and inputs the theme (e.g., "friendship") and character (e.g., "Rio the Rabbit").

[0998] Step 2:

[0999] The device captures the user's facial expressions and voice through the user's camera and microphone, and obtains emotional data in real time.

[1000] Step 3:

[1001] The device sends the acquired emotion data to the emotion engine, which analyzes the user's current emotion and identifies emotion categories such as "joy," "sadness," and "surprise."

[1002] Step 4:

[1003] The emotion engine returns the analysis results to the device, which then transmits the information to the server, thereby transmitting the user's emotion data to the server.

[1004] Step 5:

[1005] The server performs filtering based on the theme, character settings, and emotion data received from the device, specifically checking for taboo words and explicit content.

[1006] Step 6:

[1007] The server checks the filtering results, and if any inappropriate content is found, it generates a message requesting the user to correct the content and sends it to the terminal. If no inappropriate content is found, it proceeds to the next step.

[1008] Step 7:

[1009] The server then invokes a generative AI model based on verified themes, character settings, and emotional data to generate a story. For example, if the user expresses the emotion of "enjoyment," a story containing positive and enjoyable episodes will be generated.

[1010] Step 8:

[1011] The server converts the generated story into a text format, which can be modified as needed.

[1012] Step 9:

[1013] The server uses image generation AI to generate a corresponding illustration based on the story. The illustration also reflects the user's emotional data. For example, if the emotion of "fun" is conveyed, an illustration of a bright and cheerful scene is generated.

[1014] Step 10:

[1015] The server integrates the generated illustrations with the story to create a coherent digital picture book.

[1016] Step 11:

[1017] The server invokes a translation engine to translate the generated digital picture book into a language specified by the user (e.g., English, French).

[1018] Step 12:

[1019] The server checks the quality of the translated stories and makes corrections if necessary.

[1020] Step 13:

[1021] The server completes the digital picture book integrating the translation and illustrations and transmits the data to the user's terminal.

[1022] Step 14:

[1023] The terminal provides an interface for displaying the received digital picture book, allowing the user to view it.

[1024] Step 15:

[1025] The user can share the created digital picture book with other users using the sharing function of the device. Sharing methods include social networking sites, email, and cloud services.

[1026] Step 16:

[1027] The terminal transmits picture book data to a sharing terminal, so that other users can view the picture book data.

[1028] The above is the specific processing flow of the system of the present invention, which combines an emotion engine. This flow makes it possible to generate, distribute, and share personalized stories that reflect the user's emotions.

[1029] Example 2

[1030] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1031] Current story generation systems struggle to generate content that reflects the user's emotions, resulting in a lack of personalized educational resources. Furthermore, they lack the ability to automatically translate the generated stories and images into multiple languages ​​to match the user's emotions and themes, making them unable to meet the needs of global users. Furthermore, they have yet to provide digital content in a format that can be easily shared among users.

[1032] The specification process by the specification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving input data, means for filtering the received input data, means for generating a story based on the filtered data, means for generating images based on the generated story, means for translating the generated story into multiple languages, means for delivering the translated story to the terminal, means for sharing the story delivered to the terminal with other users, means for acquiring and analyzing user emotion data, and means for generating a story and images based on the acquired emotion data. This makes it possible to generate a personalized story and images based on the user's emotions, translate the story into multiple languages, and provide digital content that can be easily shared among users.

[1033] The "means for receiving input data" is a method by which a user inputs theme and character settings and transmits them to the system.

[1034] A "means for filtering received input data" is a method for analyzing received data and removing inappropriate content.

[1035] A "means for generating a story" is a method for automatically creating a story based on the filtered data and the user's emotional data.

[1036] An "image generation means" is a method for creating related visual content based on the generated narrative.

[1037] A "means for translating stories" is a method for translating the generated stories into multiple languages.

[1038] A "means for delivering a story to a terminal" is a method for transferring a translated story to a user's device.

[1039] "Means for sharing stories with other users" refers to methods that provide an interface or protocol for sharing distributed stories.

[1040] "Means for acquiring and analyzing emotion data" refers to a method for collecting and analyzing emotions from the user's facial expressions, voice, etc.

[1041] The "means for generating a story and images based on the acquired emotion data" is a method for generating a story and images using the analyzed emotion data.

[1042] The present invention is a system that combines an emotion engine that recognizes a user's emotions, generates safe and educational stories, and translates and shares them in multiple languages. Specific embodiments of this system and their operation are described in detail below.

[1043] System configuration

[1044] The system includes the following elements:

[1045] 1. Device: A device on which users input theme and character settings and acquire facial and voice data. This includes smartphones, tablets, and PCs.

[1046] 2. Server: A central processing unit that filters input data, generates stories and images, translates, and distributes data.

[1047] 3. Emotion engine: Software for acquiring and analyzing emotional data from the user's facial expressions and voice. Uses an API for emotion analysis.

[1048] 4. Generative AI model: An artificial intelligence model for generating stories and images based on user input and emotional data. Specifically, it uses large-scale language models and image generation models such as GPT (Generative Pre-trained Transformer).

[1049] How it works

[1050] User theme and character settings, and emotion recognition

[1051] The user opens the device's interface and inputs a theme (e.g., "Friendship") and a character (e.g., "Rio the Rabbit"). The device captures the user's facial expressions and voice through a camera and microphone, and sends the data to the emotion engine. The emotion engine analyzes the acquired emotion data and recognizes the user's current emotion. The results of this analysis are sent to the server and reflected in the theme and character settings.

[1052] Filtering input data

[1053] The server analyzes all input data, including the theme and character settings received from the device and the emotional data obtained from the emotion engine, and then begins the filtering process, specifically running a screening process for banned words and explicit content.

[1054] Story and illustration generation

[1055] The server generates a story by calling a generative AI model based on verified themes, character settings, and emotional data. For example, if the user is happy, it generates a story with a happy ending.

[1056] Example prompt sentence:

[1057] "Generate a story about friendship, featuring Rio the rabbit and Ken the turtle going on a fun adventure. The user emotion is 'fun'."

[1058] The server then converts the generated story into text format and uses image generation AI to generate illustrations that correspond to the story. The generated illustrations also reflect emotional elements based on the emotional data.

[1059] Example prompt sentence:

[1060] "Generate an illustration that reflects the emotion of enjoyment based on the following story: 'Rio the rabbit and Ken the turtle go on a fun adventure.'"

[1061] Multilingual Translation

[1062] The server calls a translation engine to translate the generated story into the language specified by the user (e.g., English, French). It checks the translated text obtained from the translation engine and makes corrections if necessary.

[1063] Story Distribution

[1064] The server generates digital picture book data that integrates the translated story and illustrations and sends it to the user's device, which provides an interface for displaying the received digital picture book so that the user can view it.

[1065] Share your story

[1066] Users can share the digital picture book they have created with other users using the sharing function of their device. Sharing methods include social networking sites, email, and cloud services. The device sends the picture book data to the destination device, allowing other users to view it.

[1067] Specific examples

[1068] For example, if a parent or guardian selects characters Rio the Rabbit and Ken the Turtle as characters with the theme of "friendship" and the emotion at that time is "fun," a story about the fun adventures of Rio the Rabbit and Ken the Turtle is generated based on this emotion. This story is filtered, appropriate illustrations are generated, and a digital picture book translated into Japanese, English, and French is delivered to the parent or guardian's smartphone. The parent or guardian can then share this picture book with friends via social media.

[1069] This invention can provide more personalized educational resources by generating stories and illustrations that reflect the user's emotions, which is expected to promote communication between parents and children and contribute to reducing educational disparities.

[1070] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1071] Step 1:

[1072] The user opens the device interface and inputs the theme (e.g., "Friendship") and character (e.g., "Rio the Rabbit"), and when the user confirms the input data, the device sends the data to the server.

[1073] Input: Input data for the theme "Friendship" and the character "Rabbit Rio"

[1074] Output: The input data is sent to the server

[1075] Step 2:

[1076] The device acquires the user's facial expressions and voice and sends them to the emotion engine. The emotion engine analyzes the facial and voice data to recognize the user's current emotion. The result is then sent to the server.

[1077] Input: User's facial expression data and voice data

[1078] Output: Recognized emotion data

[1079] Step 3:

[1080] The server receives the input data (theme and character settings) and analyzed emotional data, and performs filtering. Specifically, it screens for prohibited words and extreme content. Once filtering is complete, it notifies the user of the results.

[1081] Input: Theme "Friendship", Character "Rabbit Rio", Emotion data

[1082] Output: Filtering results and a message requesting corrections if necessary

[1083] Step 4:

[1084] The server invokes the generative AI model based on the verified input data and emotion data to generate a story. For example, if the user's emotion is "fun," it generates a story with a happy ending.

[1085] Input: Filtered data and sentiment data

[1086] Output: Generated narrative text

[1087] Specific operation example:

[1088] The server sends a prompt to the AI ​​model, such as "Generate a story about friendship, with Rio the rabbit and Ken the turtle going on a fun adventure. The user's emotion is 'fun'."

[1089] Step 5:

[1090] The server converts the generated story text into a text format and calls an image generation AI to generate illustrations corresponding to the story, which reflect the user's emotions.

[1091] Input: Generated narrative text

[1092] Output: Generated illustration

[1093] Specific operation example:

[1094] The prompt text "Generate an illustration that reflects the emotion of enjoyment based on the following story: 'Rio the rabbit and Ken the turtle go on a fun adventure'" is sent to the image generation AI.

[1095] Step 6:

[1096] The server invokes a translation engine that translates the generated story text into the language specified by the user (e.g., English, French). The translated text is returned to the server, which reviews it and makes corrections if necessary.

[1097] Input: Japanese story text

[1098] Output: Translated story text

[1099] Step 7:

[1100] The server integrates the translated story and illustrations to create digital picture book data, which is then sent to the user's device.

[1101] Input: translated story text, generated illustrations

[1102] Output: Digital picture book data

[1103] Step 8:

[1104] An interface for displaying the digital picture book received by the terminal is provided, allowing the user to view it.

[1105] Input: Digital picture book data

[1106] Output: A displayed digital picture book

[1107] Step 9:

[1108] Users can share the digital picture book they have created with other users using the sharing function of their device. Sharing methods include social networking sites, email, and cloud services.

[1109] Input: Digital picture book data

[1110] Output: Digital picture books shared via social media, email, and the cloud

[1111] Step 10:

[1112] The terminal transmits picture book data to a sharing device so that other users can view it.

[1113] Input: Share request

[1114] Output: Digital picture book data transferred to the sharing device

[1115] (Application example 2)

[1116] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1117] Conventional content generation and sharing systems were unable to generate and provide individual stories that took the user's emotions into consideration. This made it difficult to provide personalized content that responded to the user's emotions, and the educational and entertainment benefits were not fully realized. Furthermore, the inability to translate and share content in multiple languages ​​limited global sharing. This created challenges that made it difficult to improve the user experience and correct educational disparities.

[1118] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1119] In this invention, the server includes means for receiving input data and emotional data, means for filtering the received input data and emotional data, means for generating a story based on the filtered data and emotional data, means for generating images based on the generated story, means for translating the generated story into multiple languages, means for distributing the translated story to a terminal, means for sharing the story distributed to the terminal with other users, means for accepting theme and character settings and emotional data from a user, means for automatically generating a story based on the accepted theme and characters, means for personalizing the story based on the accepted emotional data, and means for simultaneously generating a story and images, integrating them, evaluating them based on the user's emotional data, and ensuring quality. This enables the generation of personalized stories and illustrations that take user emotions into consideration, realizes multilingual translation and easy sharing, and enables an improved user experience and mitigation of educational disparities.

[1120] "Input Data" refers to the theme, characters, and other setting information that a user provides to the system.

[1121] "Emotion data" refers to information about the emotional state of a user obtained from facial expressions, voice, etc.

[1122] "Filtering" refers to the process of removing inappropriate content or prohibited words from received data.

[1123] "Story" refers to a document or text that is generated based on themes, characters, and user emotional data.

[1124] "Images" refers to illustrations and visual content generated based on the story content and user emotional data.

[1125] "Translation" refers to the process of converting a generated story into multiple languages.

[1126] "Terminal" refers to an electronic device used by a user, such as a smartphone, tablet, or computer.

[1127] "Sharing" refers to the process of providing or distributing the generated stories and images to other users and platforms.

[1128] "Theme" refers to a concept that indicates the main point or main point of a story.

[1129] "Characters" refers to characters that appear in stories or animated characters.

[1130] "Personalization" refers to the process of adapting content based on a user's individual emotional data and preferences.

[1131] "Synthesis" refers to the process of bringing together the generated stories and images into a single piece of content.

[1132] "Evaluation" refers to the process used to judge the quality and appropriateness of the stories and images produced.

[1133] "Quality assurance" refers to the process of verifying that the generated content is appropriate and guaranteeing its quality.

[1134] The present invention is a system that combines an emotion engine that recognizes a user's emotions, generates safe and educational stories, and translates and shares them in multiple languages. An embodiment of this system will be described in detail below.

[1135] Description of system programs and processes

[1136] User theme and character settings, and emotion recognition

[1137] The user opens the device's interface and inputs the theme and character. For example, the user may select the theme "friendship" and "Rio the rabbit" as the character. Emotional data is also acquired from the user's facial expressions and voice. The device sends this data to an emotion engine, which analyzes it to recognize the user's current emotion. The results of this analysis are reflected in the theme and character settings.

[1138] Filtering input data

[1139] The server analyzes and filters all input data, including the theme and character settings received from the device and the emotion data obtained from the emotion engine. Specifically, a screening process is carried out for a list of prohibited words and explicit content. If inappropriate content is included, the server generates a message requesting the user to correct it and sends it to the device. If there is no inappropriate content, the process proceeds to the next step.

[1140] Story and illustration generation

[1141] The server generates a story by calling a generative AI model based on verified themes, character settings, and emotional data. For example, if the user is happy, a story with a happy ending is generated. This story is converted into text format, and then an image generation AI is used to generate illustrations corresponding to the story. Based on the emotional data, the illustrations also reflect emotional elements.

[1142] Multilingual Translation

[1143] The server invokes a translation engine to translate the generated story into a language specified by the user, such as Japanese, English, French, etc. The translated story is then reviewed by the server and any necessary corrections are made.

[1144] Story Distribution

[1145] The server generates digital picture book data that integrates the translated story and illustrations and transmits it to the user's terminal, which provides an interface for displaying the received digital picture book so that the user can view it.

[1146] Share your story

[1147] Users can share the digital picture book they have created with other users using the sharing function of their device. Sharing methods include social networking sites, email, and cloud services. The device sends the picture book data to the destination device, allowing other users to view it.

[1148] Hardware and software used

[1149] The server runs software such as generative AI models, emotion recognition engines, image generation AI, translation engines, and data filtering systems. Specific software and hardware used include the Google Translate API, EmotionRecognizer, StoryGenerator, and ImageGenerator. These work together to enable the generation, translation, and sharing of stories and illustrations based on user requests.

[1150] Specific examples

[1151] For example, if a parent or guardian selects characters Rio the Rabbit and Ken the Turtle as characters with the theme of "friendship" and the emotion at that time is "fun," a story about the fun adventures of Rio the Rabbit and Ken the Turtle is generated based on this emotion. The story is filtered and appropriate illustrations are generated, and then a digital picture book translated into Japanese, English, and French is delivered to the parent or guardian's smartphone. The parent or guardian can share this picture book with friends via social media.

[1152] Prompt Sentence Examples

[1153] Theme: Friendship

[1154] Character: Rio the Rabbit

[1155] Emotion data: {'emotion': 'joy'}

[1156] prompt:

[1157] "Rio the rabbit was having a great time playing with his friends. They decided to go on a new adventure together. Along the way, Rio..."

[1158] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1159] Step 1:

[1160] The user opens the device interface and inputs the theme and character. For example, the theme is set to "friendship" and the character is set to "Rio the rabbit." The user's facial expressions and voice data are also input. This input data is sent to the device (input data and emotion data).

[1161] Step 2:

[1162] The device processes the received theme, character, and emotion data and sends it to the server. The emotion data includes the results of the user's facial expression analysis and voice analysis. The device calls the emotion recognition engine to identify the user's emotion and sends the result as data (input data: theme, character, emotion data / output data: theme, character, emotion analysis result).

[1163] Step 3:

[1164] The server analyzes the theme, character, and emotion data received from the device and performs a filtering process. The filtering process screens the data based on a list of prohibited words and extreme content to check for inappropriate content. If inappropriate content is found, a message is generated requesting the user to correct it (input data: theme, character, emotion analysis results / output data: filtered data or correction request message).

[1165] Step 4:

[1166] Based on the filtered data, the server calls a generative AI model to generate a story. The generative AI model adjusts the tone and ending of the story based on the emotional data. For example, if the emotional data indicates joy, a story with a happy ending will be generated (input data: filtered data / output data: generated story).

[1167] Step 5:

[1168] Next, the server calls an image generation AI based on the generated story, which generates illustrations that correspond to the content of the story. Emotional data is also taken into account, and emotional elements are reflected in the illustrations (input data: generated story, emotion data / output data: generated illustrations).

[1169] Step 6:

[1170] The server invokes a translation engine to translate the generated story into multiple languages, for example, from Japanese to English and French. The translation is done automatically and can be manually corrected if necessary (input data: generated story / output data: translated story).

[1171] Step 7:

[1172] The server generates digital picture book data that integrates the translated story and generated illustrations and sends it to the user's device. The device provides an interface for displaying the received digital picture book, allowing the user to easily browse it (input data: translated story, generated illustrations / output data: digital picture book data).

[1173] Step 8:

[1174] Users can share the digital picture book they have created with other users using the sharing function of their device. Social networking sites, email, cloud services, and other methods of sharing are available. The device sends the picture book data to the destination device, allowing other users to view it (input data: digital picture book data; output data: data sent to the destination device).

[1175] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1176] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1177] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1178] [Fourth embodiment]

[1179] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1180] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1181] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1182] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1183] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1184] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1185] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1186] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1187] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1188] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1189] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1190] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1191] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1192] This system allows users to create safe and educational stories by setting themes and characters, and then translates and shares them in multiple languages. The program and processing of this system are described in detail below. The system mainly works in conjunction with three elements: the server, the terminal, and the user.

[1193] Description of system programs and processes

[1194] User Theme and Character Settings

[1195] 1. A user accesses the system through a terminal and inputs a theme (e.g., "friendship") and a character (e.g., "Rio the rabbit") using the input interface.

[1196] 2. The device sends the entered theme and character settings to the server.

[1197] Filtering input data

[1198] 1. The server analyzes and filters the theme and character settings received from the device, specifically by screening for prohibited words and explicit content.

[1199] 2. If the server determines that the input data is appropriate based on the filtering results, it proceeds to the next step using that data. If the data contains inappropriate data, it sends a message to the terminal requesting the user to correct it.

[1200] Narrative Generation

[1201] 1. The server calls a generative AI model based on the verified theme and character settings to generate a story. For example, it generates a story in which "Rio the Rabbit" cooperates with his friend "Ken the Turtle" to cross a large bridge.

[1202] 2. The server converts the generated story into a specified format, such as text format.

[1203] Illustration generation

[1204] 1. The server uses image generation AI to create illustrations that correspond to the generated story. Multiple illustrations that match the scenes in the story are generated.

[1205] 2. The server integrates the generated illustrations and story into a coherent picture book format.

[1206] Multilingual Translation

[1207] 1. The server invokes a translation engine to translate the generated story into the language specified by the user (e.g., English, French).

[1208] 2. The server reviews the translated story and makes any necessary corrections.

[1209] Story Distribution

[1210] 1. The server generates digital picture book data that integrates the translated story and illustrations and sends it to the user's device.

[1211] 2. The terminal provides an interface that displays the received picture book data to the user, allowing the user to view the picture book.

[1212] Share your story

[1213] 1. The user shares the created picture book with other users using the sharing function of the device. Sharing methods include social networking sites, email, and cloud services.

[1214] 2. The device sends the picture book data to the sharing device so that other users can view it.

[1215] For example, a parent or guardian can create a story with the theme of "friendship" and featuring characters Rio the Rabbit and Ken the Turtle. The story is then generated and filtered by the server. Illustrations corresponding to the story are then generated, and a digital picture book translated into Japanese, English, and French is finally generated. The picture book is then delivered to the parent's smartphone, and the parent can share it with other parents via social networking sites.

[1216] As described above, the system of the present invention can effectively generate and provide safe and educational stories to children, which is expected to promote communication between parents and children and contribute to reducing educational disparities.

[1217] The processing flow will be explained below.

[1218] Step 1:

[1219] The user opens the terminal interface and inputs the theme (e.g., "friendship") and character (e.g., "Rio the Rabbit").

[1220] Step 2:

[1221] The terminal receives the entered theme and character settings and transmits this data to the server.

[1222] Step 3:

[1223] The server analyzes the theme and character settings received from the device and begins the filtering process, specifically running a screening process for banned words and explicit content.

[1224] Step 4:

[1225] The server checks the filtering results, and if any inappropriate content is found, it generates a message requesting the user to correct the content and sends it to the terminal. If no inappropriate content is found, it proceeds to the next step.

[1226] Step 5:

[1227] The server calls the generative AI model based on the verified theme and character settings to generate a story. For example, it generates a story in which Rio the rabbit cooperates with his friend Ken the turtle to cross a large bridge.

[1228] Step 6:

[1229] The server converts the generated story into text format and then uses image generation AI to begin generating illustrations that correspond to the story.

[1230] Step 7:

[1231] The server integrates the generated illustrations with the story to create a coherent digital picture book.

[1232] Step 8:

[1233] The server executes a translation process to translate the generated digital picture book into a user-specified language (e.g., English, French).

[1234] Step 9:

[1235] The server checks the quality of the translated story, makes corrections if necessary, and finally generates the finished digital picture book.

[1236] Step 10:

[1237] The server transmits the completed digital picture book data to the user's terminal.

[1238] Step 11:

[1239] The terminal provides an interface for displaying the received digital picture book, allowing the user to view it.

[1240] Step 12:

[1241] The user shares the created digital picture book with other users using the sharing function of the device. For example, the picture book data is sent via social networking sites, email, or cloud services.

[1242] Step 13:

[1243] The terminal transmits picture book data to a sharing terminal, so that other users can view the picture book data.

[1244] The above is the specific processing flow of the system of the present invention, which enables safe and educational stories to be created, distributed, and shared.

[1245] Example 1

[1246] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1247] Previous story generation systems were complicated in the process of generating safe and educational stories based on themes and characters set by the user, and lacked a means to translate and share the generated stories in multiple languages. Furthermore, there were no systems that automatically generated illustrations along with the story and integrated them. As a result, users had to manually create the story, illustrations, and translation, which required a great deal of effort and time.

[1248] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1249] In this invention, the server includes an input means for a user to set a theme and characters, a means for receiving input data of the set theme and characters and analyzing and filtering the input data, a means for generating a story using a generative AI model based on the filtered data, a means for converting the generated story into a text format, a means for generating illustrations using an image generation AI model based on the generated story, a means for integrating the generated story and illustrations into a consistent format, a means for translating the generated story into multiple languages, a means for delivering the translated story to a terminal, and a means for sharing the story delivered to the terminal with other users. This makes it possible to effectively generate safe and educational stories based on themes and characters set by users, translate them into multiple languages, and share them.

[1250] A "user" is a person or entity that accesses the system and configures themes and characters.

[1251] A "terminal" is an electronic device used by a user, such as a computer, smartphone, or tablet.

[1252] "Server" means the central control unit that manages and executes the functions of the entire system.

[1253] A "theme" is the main theme or central concept of a story.

[1254] A "character" is a person, animal, or fictional being that appears in a story.

[1255] "Input means" refers to an interface that allows a user to input themes and characters into the system.

[1256] A "generative AI model" is an artificial intelligence model that generates stories based on themes and characters set by the user.

[1257] "Filtering" is the process of analyzing input data and eliminating inappropriate content.

[1258] "Text format" is a data format that expresses the generated story in a sentence format.

[1259] An "image generation AI model" is an artificial intelligence model for generating illustrations that correspond to a story.

[1260] A "translation engine" is software or a service that translates generated stories into other languages.

[1261] A "consistent format" means that the story and illustrations are in harmony and have a sense of unity.

[1262] "Distribution means" refers to a mechanism or method for transmitting the generated content to the user's terminal.

[1263] "Sharing means" refers to a function or method for sharing the generated content with other users.

[1264] This system allows users to create safe and educational stories by setting themes and characters, and then translates and shares them in multiple languages. This system operates in cooperation with three main elements: a server, a terminal, and users.

[1265] Users access the system using a terminal and set a theme (e.g., "Friendship") and a character (e.g., "Rio the Rabbit"). The terminal then sends the theme and character settings entered by the user to the server. The server analyzes the received data and filters it using natural language processing technology (NLP). Specifically, it screens for prohibited words and extreme content, and if inappropriate data is included, it sends a message to the terminal requesting the user to correct it.

[1266] The server then uses a generative AI model (e.g., GPT-3) to generate a story based on the filtered data. The generated story is converted into text format and formatted as needed. The server then uses an image generation AI (e.g., DALL-E) to create illustrations corresponding to each scene in the generated story. The prompt text also includes a description of the specific scene. The generated illustrations are integrated into the story and compiled into a coherent picture book format.

[1267] The server then uses a translation engine (e.g., Google Translate API) to translate the generated story into the user's specified language (e.g., English, French). The translation result undergoes grammar checks and style guide application, correcting any necessary parts. The server then sends the translated story and illustrations integrated into the digital picture book data to the user's device. The device then provides an interface for displaying the received picture book data, allowing the user to view the picture book.

[1268] As a concrete example, a user sets the characters "Rio the Rabbit" and "Ken the Turtle" as characters with the theme of "friendship." The server filters this and uses a generative AI model to generate a story in which "Rio the Rabbit" and "Ken the Turtle" cooperate to cross a bridge. Next, DALL-E is used to generate illustrations of the scene crossing the bridge. Finally, the generated story is translated into Japanese, English, and French and delivered to the user's smartphone as a digital picture book.

[1269] Example prompt sentence:

[1270] Theme: "Friendship"

[1271] Characters: "Rio the Rabbit" and "Ken the Turtle"

[1272] Objective: To generate stories that have educational value for children.

[1273] The system of the present invention can effectively generate safe and educational stories based on themes and characters set by the user, and can translate and share them in multiple languages, which is expected to promote communication between parents and children and contribute to reducing educational disparities.

[1274] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1275] Step 1:

[1276] A user accesses the system using a terminal and inputs a theme (e.g., "Friendship") and a character (e.g., "Rio the Rabbit"). The input theme and character settings are sent to the server by the terminal. The input data is sent in the form of a theme and a character.

[1277] Step 2:

[1278] The server analyzes the theme and character settings received from the device. It uses natural language processing (NLP) technology to parse the input data and understand its meaning. Based on the results of this analysis, it checks for inappropriate data by screening it against a list of prohibited words and for extreme content. As an output, it generates data that is deemed appropriate or that requires correction.

[1279] Step 3:

[1280] If the server determines that the filtered data is appropriate, it calls a generative AI model (e.g., GPT-3) based on the data to generate a story. At this time, the prompt text includes the theme and character information entered by the user. The input data is the theme and character settings, and the output data is the generated story.

[1281] Step 4:

[1282] The server converts the generated story into a text format and formats it, for example dividing the story into paragraphs, applying a particular font size and style, and adjusting line breaks and punctuation as needed. The input data is the generated story, and the output data is the formatted story text.

[1283] Step 5:

[1284] The server calls an image generation AI (e.g., DALL-E) to create illustrations corresponding to the generated story. By including a specific description of a particular scene in the story in the prompt, an illustration appropriate for that scene is generated. The input data is a description of the story scene, and the output data is the generated illustration.

[1285] Step 6:

[1286] The server integrates the generated illustrations and story into a coherent picture book format. This is the process of adjusting the placement of illustrations and text and determining the page layout. The input data is the story in text format and the generated illustrations, and the output data is the integrated digital picture book.

[1287] Step 7:

[1288] The server uses a translation engine (e.g., Google Translate API) to translate the generated story into the language specified by the user (e.g., English, French). It checks the translated content for grammar and makes any necessary corrections. The input data is the story in text format, and the output data is the translated story.

[1289] Step 8:

[1290] The server generates digital picture book data that integrates the translated story and illustrations and sends it to the user's terminal. The terminal provides an interface for displaying the received picture book data, allowing the user to view the picture book. The input data is the translated story and the generated illustrations, and the output data is the user's viewing interface.

[1291] Step 9:

[1292] Users can share the created picture book with other users using the sharing function of their device. Sharing methods include social networking sites, email, and cloud services. The input data is digital picture book data, and the output data is in a format that can be viewed by the recipient users.

[1293] (Application example 1)

[1294] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1295] In conventional story generation systems, it is common for users to set a theme and characters, and then generate a story based on those. However, they lack the functionality to compile the generated story and illustrations into a digital booklet in a consistent format, distribute it in multiple languages, or easily share it with other users. For this reason, there is a need for a system that increases users' freedom of expression and provides convenience in international multilingual support.

[1296] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1297] In this invention, the server includes means for receiving input data, means for filtering the received input data, means for generating a story based on the filtered data, means for generating images based on the generated story, means for translating the generated story into multiple languages, means for distributing the translated story to a terminal, means for sharing the story distributed to the terminal with other users, means for receiving theme and character settings from a user, means for automatically generating a story based on the received theme and characters, and means for integrating the story and images into a digital booklet format. This makes it possible to translate, distribute, and share educational and safe stories and corresponding illustrations into multiple languages ​​in a consistent digital booklet format based on the themes and characters set by the user.

[1298] The "means for receiving input data" is a function for receiving the theme and character information provided by the user via the terminal.

[1299] "Means for filtering received input data" refers to a function for filtering out inappropriate content from input data and narrowing it down to safe and appropriate data.

[1300] The "means for generating a story based on filtered data" is a function for automatically generating a story based on verified themes and character settings.

[1301] The "means for generating images based on the generated story" is a function for automatically generating illustrations according to the content of the story.

[1302] The "means for translating the generated story into multiple languages" is a function for automatically converting the generated story into multiple specified languages.

[1303] The "means for delivering the translated story to the terminal" is a function for transferring the translated story and corresponding illustrations to the user's terminal.

[1304] The "means for sharing a story delivered to the terminal with other users" is a function that enables a user to share a received story with other users via a social networking service or email.

[1305] The "means for accepting theme and character settings from the user" is a function that provides an input interface for the user to set the theme and character.

[1306] "Means for automatically generating a story based on the accepted theme and characters" is a function for automatically generating a story using a generative AI model based on the theme and character information set by the user.

[1307] "Means for integrating stories and images into a digital booklet format" is a function for combining the generated stories and illustrations and editing them into a coherent digital picture book format.

[1308] A "social networking service" is a platform for sharing information and interacting with other users over the Internet.

[1309] "Email" is a means of communication for sending and receiving text messages and files over the Internet.

[1310] The following describes an embodiment of the present invention. The system mainly consists of three elements: a server, a terminal, and a user. The specific roles and processing procedures of each element are described below.

[1311] System programs and their processing

[1312] User Theme and Character Settings

[1313] Users access the system through a device such as a smartphone or tablet and use an interface to input the theme (e.g., "Friendship") and character (e.g., "Rio the Rabbit") to set up the system. This information is then sent from the device to the server.

[1314] Filtering input data

[1315] The server analyzes the received themes and character settings, and screens them for banned words and explicit content. Scripts written in Python and Flask perform the filtering. If inappropriate data is detected, a message is sent to the user's device, prompting them to correct it.

[1316] Narrative Generation

[1317] Based on the filtered themes and character settings, the server invokes a generative AI model (e.g., GPT-4) to generate a story, which is then converted into text format and temporarily stored on the server.

[1318] Illustration generation

[1319] The server then uses image generation AI (e.g., Stable Diffusion or DALL-E) to generate illustrations that correspond to the generated story. These illustrations are aligned with the scenes in the story.

[1320] Multilingual Translation

[1321] The generated story is translated into the specified language (e.g., English, French) using the Google Translate API or DeepL API. The translated story is stored on the server and corrected as needed.

[1322] Story Distribution

[1323] The server then aggregates the translated stories and illustrations into a digital booklet and delivers it to users' devices, where they can view it through an application on their smartphones or tablets.

[1324] Share your story

[1325] Users can use the sharing function of their terminals to share the generated digital booklet with other users via social networking services or email.

[1326] Overview of the technologies and programs used

[1327] Server side: Python, Flask (or Django), a generative AI model (e.g. GPT-4), a translation engine (Google Translate API or DeepL).

[1328] Frontend: React Native (for mobile apps) or React.js (for web).

[1329] AI model: OpenAI's GPT-4, Stable Diffusion or DALL-E (illustration generation).

[1330] Database: PostgreSQL (storage of user data, generated stories, and illustrations).

[1331] Specific examples

[1332] For example, a child might create a story with the theme of "friendship" and featuring characters Rio the rabbit and Ken the turtle. The story is then generated and filtered by the server. Illustrations corresponding to the story are then generated, and a digital booklet translated into Japanese, English, and French is finally generated. This picture book is then delivered to the user's smartphone, and the user can share it with other users via social networking sites.

[1333] Generative AI model prompt example

[1334] prompt:

[1335] Theme: "Friendship"

[1336] Characters: "Rio the Rabbit" and "Ken the Turtle"

[1337] Instructions: Create an educational and safe story for children. The story should be based on the adventures of Rio the rabbit and Ken the turtle who team up to cross a big bridge. Write in a friendly, easy-to-understand voice.

[1338] In this way, the system of the present invention can effectively generate and provide safe and educational stories to users, which is expected to promote communication between users and contribute to the spread of education.

[1339] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1340] Step 1:

[1341] Users access the system through a device such as a smartphone or tablet and use the interface to input the theme (e.g., "Friendship") and character (e.g., "Rio the Rabbit") to set the game. The set information is sent from the device to the server as input data. The input is the theme and character setting data from the user's interface, and the output is the input data sent to the server.

[1342] Step 2:

[1343] The server filters the input data for themes and character settings received from the device. Specifically, it uses scripts using Python and Flask to screen for banned words and explicit content. The input is the theme and character setting data, and filtering is performed as data processing, with the output being the appropriate filtered data.

[1344] Step 3:

[1345] The server generates a story using a generative AI model (e.g., GPT-4) based on the filtered theme and character settings. The input is the filtered data, and a story is generated using the generative AI model as data calculation. The output is the text data of the generated story. The generated story is temporarily stored on the server.

[1346] Step 4:

[1347] The server uses an image generation AI (e.g., Stable Diffusion or DALL-E) based on the text data of the generated story to generate illustrations corresponding to each scene in the story. The input is the text data of the generated story, and the image generation AI is used as data calculation to generate illustrations, and the output is the generated illustration.

[1348] Step 5:

[1349] The server translates the generated story into multiple languages. Specifically, it uses the Google Translate API or DeepL API to translate the input data, which is the generated story text. The translation API is called as data processing, and the output is the translated multilingual story text data.

[1350] Step 6:

[1351] The server integrates the translated story and the generated illustrations into a digital booklet. The input is the text data of the translated story and the generated illustration data, which are integrated as data processing, and the output is the digital booklet data.

[1352] Step 7:

[1353] The server distributes the digital booklet data to the user's terminal. The input is the digital booklet data, which is distributed to the terminal as data transfer, and the output is the digital booklet displayed on the user's terminal.

[1354] Step 8:

[1355] The user uses the sharing function of the terminal to share the generated digital booklet with other users via social networking services or email. The input is the distributed digital booklet data, and the sharing function is used as data transfer, and the output is the digital booklet shared with other users.

[1356] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1357] This invention is a system that combines an emotion engine that recognizes the user's emotions, generates safe and educational stories, and translates and shares them in multiple languages. The program and processing of this system are explained in detail below. The system mainly works in cooperation with three elements: the server, the terminal, and the user.

[1358] Description of system programs and processes

[1359] User theme and character settings, and emotion recognition

[1360] 1. The user opens the terminal interface and inputs the theme (e.g., "Friendship") and character (e.g., "Rio the Rabbit").

[1361] 2. The device acquires emotion data from the user's facial expressions and voice and sends that data to the emotion engine.

[1362] 3. The emotion engine analyzes the acquired emotion data and recognizes the user's current emotion. The analysis results are reflected in the theme and character settings.

[1363] Filtering input data

[1364] 1. The server analyzes all input data, including the theme and character settings received from the device and the emotion data obtained from the emotion engine, and starts the filtering process. Specifically, it runs a screening process for banned words and explicit content.

[1365] 2. The server checks the filtering results and if any inappropriate content is found, it generates a message asking the user to correct it and sends it to the terminal. If there is no inappropriate content, it proceeds to the next step.

[1366] Story and illustration generation

[1367] 1. The server calls the generative AI model based on the verified theme, character settings, and emotional data to generate a story. For example, if the user is happy, it generates a story with a happy ending, depending on the user's emotions.

[1368] 2. The server converts the generated story into text format, and then uses image generation AI to generate illustrations corresponding to the story. Based on the emotional data, the illustrations also reflect emotional elements.

[1369] Multilingual Translation

[1370] 1. The server invokes a translation engine that translates the generated story into the language specified by the user (e.g., English, French).

[1371] 2. The server reviews the translated story and makes corrections if necessary.

[1372] Story Distribution

[1373] 1. The server generates digital picture book data that integrates the translated story and illustrations and sends it to the user's device.

[1374] 2. The terminal provides an interface for displaying the received digital picture book, allowing the user to view it.

[1375] Share your story

[1376] 1. The user shares the created digital picture book with other users using the sharing function of their device. Sharing methods include social networking sites, email, and cloud services.

[1377] 2. The device sends the picture book data to the sharing device so that other users can view it.

[1378] For example, if a parent or guardian selects characters Rio the Rabbit and Ken the Turtle as characters with the theme of "friendship" and the emotion they are feeling is "fun," a story about the fun adventures of Rio the Rabbit and Ken the Turtle is generated based on this emotion. The story is filtered, appropriate illustrations are generated, and a digital picture book translated into Japanese, English, and French is delivered to the parent or guardian's smartphone. The parent or guardian can then share the picture book with friends via social media.

[1379] As described above, the system of the present invention can provide more personalized educational resources by generating stories and illustrations that reflect the user's emotions, which is expected to promote communication between parents and children and contribute to reducing educational disparities.

[1380] The processing flow will be explained below.

[1381] Step 1:

[1382] The user opens the terminal interface and inputs the theme (e.g., "friendship") and character (e.g., "Rio the Rabbit").

[1383] Step 2:

[1384] The device captures the user's facial expressions and voice through the user's camera and microphone, and obtains emotional data in real time.

[1385] Step 3:

[1386] The device sends the acquired emotion data to the emotion engine, which analyzes the user's current emotion and identifies emotion categories such as "joy," "sadness," and "surprise."

[1387] Step 4:

[1388] The emotion engine returns the analysis results to the device, which then transmits the information to the server, thereby transmitting the user's emotion data to the server.

[1389] Step 5:

[1390] The server performs filtering based on the theme, character settings, and emotion data received from the device, specifically checking for taboo words and explicit content.

[1391] Step 6:

[1392] The server checks the filtering results, and if any inappropriate content is found, it generates a message requesting the user to correct the content and sends it to the terminal. If no inappropriate content is found, it proceeds to the next step.

[1393] Step 7:

[1394] The server then invokes a generative AI model based on verified themes, character settings, and emotional data to generate a story. For example, if the user expresses the emotion of "enjoyment," a story containing positive and enjoyable episodes will be generated.

[1395] Step 8:

[1396] The server converts the generated story into a text format, which can be modified as needed.

[1397] Step 9:

[1398] The server uses image generation AI to generate a corresponding illustration based on the story. The illustration also reflects the user's emotional data. For example, if the emotion of "fun" is conveyed, an illustration of a bright and cheerful scene is generated.

[1399] Step 10:

[1400] The server integrates the generated illustrations with the story to create a coherent digital picture book.

[1401] Step 11:

[1402] The server invokes a translation engine to translate the generated digital picture book into a language specified by the user (e.g., English, French).

[1403] Step 12:

[1404] The server checks the quality of the translated stories and makes corrections if necessary.

[1405] Step 13:

[1406] The server completes the digital picture book integrating the translation and illustrations and transmits the data to the user's terminal.

[1407] Step 14:

[1408] The terminal provides an interface for displaying the received digital picture book, allowing the user to view it.

[1409] Step 15:

[1410] The user can share the created digital picture book with other users using the sharing function of the device. Sharing methods include social networking sites, email, and cloud services.

[1411] Step 16:

[1412] The terminal transmits picture book data to a sharing terminal, so that other users can view the picture book data.

[1413] The above is the specific processing flow of the system of the present invention, which combines an emotion engine. This flow makes it possible to generate, distribute, and share personalized stories that reflect the user's emotions.

[1414] Example 2

[1415] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1416] Current story generation systems struggle to generate content that reflects the user's emotions, resulting in a lack of personalized educational resources. Furthermore, they lack the ability to automatically translate the generated stories and images into multiple languages ​​to match the user's emotions and themes, making them unable to meet the needs of global users. Furthermore, they have yet to provide digital content in a format that can be easily shared among users.

[1417] The specification process by the specification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving input data, means for filtering the received input data, means for generating a story based on the filtered data, means for generating images based on the generated story, means for translating the generated story into multiple languages, means for delivering the translated story to the terminal, means for sharing the story delivered to the terminal with other users, means for acquiring and analyzing user emotion data, and means for generating a story and images based on the acquired emotion data. This makes it possible to generate a personalized story and images based on the user's emotions, translate the story into multiple languages, and provide digital content that can be easily shared among users.

[1418] The "means for receiving input data" is a method by which a user inputs theme and character settings and transmits them to the system.

[1419] A "means for filtering received input data" is a method for analyzing received data and removing inappropriate content.

[1420] A "means for generating a story" is a method for automatically creating a story based on the filtered data and the user's emotional data.

[1421] An "image generation means" is a method for creating related visual content based on the generated narrative.

[1422] A "means for translating stories" is a method for translating the generated stories into multiple languages.

[1423] A "means for delivering a story to a terminal" is a method for transferring a translated story to a user's device.

[1424] "Means for sharing stories with other users" refers to methods that provide an interface or protocol for sharing distributed stories.

[1425] "Means for acquiring and analyzing emotion data" refers to a method for collecting and analyzing emotions from the user's facial expressions, voice, etc.

[1426] The "means for generating a story and images based on the acquired emotion data" is a method for generating a story and images using the analyzed emotion data.

[1427] The present invention is a system that combines an emotion engine that recognizes a user's emotions, generates safe and educational stories, and translates and shares them in multiple languages. Specific embodiments of this system and their operation are described in detail below.

[1428] System configuration

[1429] The system includes the following elements:

[1430] 1. Device: A device on which users input theme and character settings and acquire facial and voice data. This includes smartphones, tablets, and PCs.

[1431] 2. Server: A central processing unit that filters input data, generates stories and images, translates, and distributes data.

[1432] 3. Emotion engine: Software for acquiring and analyzing emotional data from the user's facial expressions and voice. Uses an API for emotion analysis.

[1433] 4. Generative AI model: An artificial intelligence model for generating stories and images based on user input and emotional data. Specifically, it uses large-scale language models and image generation models such as GPT (Generative Pre-trained Transformer).

[1434] How it works

[1435] User theme and character settings, and emotion recognition

[1436] The user opens the device's interface and inputs a theme (e.g., "Friendship") and a character (e.g., "Rio the Rabbit"). The device captures the user's facial expressions and voice through a camera and microphone, and sends the data to the emotion engine. The emotion engine analyzes the acquired emotion data and recognizes the user's current emotion. The results of this analysis are sent to the server and reflected in the theme and character settings.

[1437] Filtering input data

[1438] The server analyzes all input data, including the theme and character settings received from the device and the emotional data obtained from the emotion engine, and then begins the filtering process, specifically running a screening process for banned words and explicit content.

[1439] Story and illustration generation

[1440] The server generates a story by calling a generative AI model based on verified themes, character settings, and emotional data. For example, if the user is happy, it generates a story with a happy ending.

[1441] Example prompt sentence:

[1442] "Generate a story about friendship, featuring Rio the rabbit and Ken the turtle going on a fun adventure. The user emotion is 'fun'."

[1443] The server then converts the generated story into text format and uses image generation AI to generate illustrations that correspond to the story. The generated illustrations also reflect emotional elements based on the emotional data.

[1444] Example prompt sentence:

[1445] "Generate an illustration that reflects the emotion of enjoyment based on the following story: 'Rio the rabbit and Ken the turtle go on a fun adventure.'"

[1446] Multilingual Translation

[1447] The server calls a translation engine to translate the generated story into the language specified by the user (e.g., English, French). It checks the translated text obtained from the translation engine and makes corrections if necessary.

[1448] Story Distribution

[1449] The server generates digital picture book data that integrates the translated story and illustrations and sends it to the user's device, which provides an interface for displaying the received digital picture book so that the user can view it.

[1450] Share your story

[1451] Users can share the digital picture book they have created with other users using the sharing function of their device. Sharing methods include social networking sites, email, and cloud services. The device sends the picture book data to the destination device, allowing other users to view it.

[1452] Specific examples

[1453] For example, if a parent or guardian selects characters Rio the Rabbit and Ken the Turtle as characters with the theme of "friendship" and the emotion at that time is "fun," a story about the fun adventures of Rio the Rabbit and Ken the Turtle is generated based on this emotion. This story is filtered, appropriate illustrations are generated, and a digital picture book translated into Japanese, English, and French is delivered to the parent or guardian's smartphone. The parent or guardian can then share this picture book with friends via social media.

[1454] This invention can provide more personalized educational resources by generating stories and illustrations that reflect the user's emotions, which is expected to promote communication between parents and children and contribute to reducing educational disparities.

[1455] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1456] Step 1:

[1457] The user opens the device interface and inputs the theme (e.g., "Friendship") and character (e.g., "Rio the Rabbit"), and when the user confirms the input data, the device sends the data to the server.

[1458] Input: Input data for the theme "Friendship" and the character "Rabbit Rio"

[1459] Output: The input data is sent to the server

[1460] Step 2:

[1461] The device acquires the user's facial expressions and voice and sends them to the emotion engine. The emotion engine analyzes the facial and voice data to recognize the user's current emotion. The result is then sent to the server.

[1462] Input: User's facial expression data and voice data

[1463] Output: Recognized emotion data

[1464] Step 3:

[1465] The server receives the input data (theme and character settings) and analyzed emotional data, and performs filtering. Specifically, it screens for prohibited words and extreme content. Once filtering is complete, it notifies the user of the results.

[1466] Input: Theme "Friendship", Character "Rabbit Rio", Emotion data

[1467] Output: Filtering results and a message requesting corrections if necessary

[1468] Step 4:

[1469] The server invokes the generative AI model based on the verified input data and emotion data to generate a story. For example, if the user's emotion is "fun," it generates a story with a happy ending.

[1470] Input: Filtered data and sentiment data

[1471] Output: Generated narrative text

[1472] Specific operation example:

[1473] The server sends a prompt to the AI ​​model, such as "Generate a story about friendship, with Rio the rabbit and Ken the turtle going on a fun adventure. The user's emotion is 'fun'."

[1474] Step 5:

[1475] The server converts the generated story text into a text format and calls an image generation AI to generate illustrations corresponding to the story, which reflect the user's emotions.

[1476] Input: Generated narrative text

[1477] Output: Generated illustration

[1478] Specific operation example:

[1479] The prompt text "Generate an illustration that reflects the emotion of enjoyment based on the following story: 'Rio the rabbit and Ken the turtle go on a fun adventure'" is sent to the image generation AI.

[1480] Step 6:

[1481] The server invokes a translation engine that translates the generated story text into the language specified by the user (e.g., English, French). The translated text is returned to the server, which reviews it and makes corrections if necessary.

[1482] Input: Japanese story text

[1483] Output: Translated story text

[1484] Step 7:

[1485] The server integrates the translated story and illustrations to create digital picture book data, which is then sent to the user's device.

[1486] Input: translated story text, generated illustrations

[1487] Output: Digital picture book data

[1488] Step 8:

[1489] An interface for displaying the digital picture book received by the terminal is provided, allowing the user to view it.

[1490] Input: Digital picture book data

[1491] Output: A displayed digital picture book

[1492] Step 9:

[1493] Users can share the digital picture book they have created with other users using the sharing function of their device. Sharing methods include social networking sites, email, and cloud services.

[1494] Input: Digital picture book data

[1495] Output: Digital picture books shared via social media, email, and the cloud

[1496] Step 10:

[1497] The terminal transmits picture book data to a sharing device so that other users can view it.

[1498] Input: Share request

[1499] Output: Digital picture book data transferred to the sharing device

[1500] (Application example 2)

[1501] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1502] Conventional content generation and sharing systems were unable to generate and provide individual stories that took the user's emotions into consideration. This made it difficult to provide personalized content that responded to the user's emotions, and the educational and entertainment benefits were not fully realized. Furthermore, the inability to translate and share content in multiple languages ​​limited global sharing. This created challenges that made it difficult to improve the user experience and correct educational disparities.

[1503] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1504] In this invention, the server includes means for receiving input data and emotional data, means for filtering the received input data and emotional data, means for generating a story based on the filtered data and emotional data, means for generating images based on the generated story, means for translating the generated story into multiple languages, means for distributing the translated story to a terminal, means for sharing the story distributed to the terminal with other users, means for accepting theme and character settings and emotional data from a user, means for automatically generating a story based on the accepted theme and characters, means for personalizing the story based on the accepted emotional data, and means for simultaneously generating a story and images, integrating them, evaluating them based on the user's emotional data, and ensuring quality. This enables the generation of personalized stories and illustrations that take user emotions into consideration, realizes multilingual translation and easy sharing, and enables an improved user experience and mitigation of educational disparities.

[1505] "Input Data" refers to the theme, characters, and other setting information that a user provides to the system.

[1506] "Emotion data" refers to information about the emotional state of a user obtained from facial expressions, voice, etc.

[1507] "Filtering" refers to the process of removing inappropriate content or prohibited words from received data.

[1508] "Story" refers to a document or text that is generated based on themes, characters, and user emotional data.

[1509] "Images" refers to illustrations and visual content generated based on the story content and user emotional data.

[1510] "Translation" refers to the process of converting a generated story into multiple languages.

[1511] "Terminal" refers to an electronic device used by a user, such as a smartphone, tablet, or computer.

[1512] "Sharing" refers to the process of providing or distributing the generated stories and images to other users and platforms.

[1513] "Theme" refers to a concept that indicates the main point or main point of a story.

[1514] "Characters" refers to characters that appear in stories or animated characters.

[1515] "Personalization" refers to the process of adapting content based on a user's individual emotional data and preferences.

[1516] "Synthesis" refers to the process of bringing together the generated stories and images into a single piece of content.

[1517] "Evaluation" refers to the process used to judge the quality and appropriateness of the stories and images produced.

[1518] "Quality assurance" refers to the process of verifying that the generated content is appropriate and guaranteeing its quality.

[1519] The present invention is a system that combines an emotion engine that recognizes a user's emotions, generates safe and educational stories, and translates and shares them in multiple languages. An embodiment of this system will be described in detail below.

[1520] Description of system programs and processes

[1521] User theme and character settings, and emotion recognition

[1522] The user opens the device's interface and inputs the theme and character. For example, the user may select the theme "friendship" and "Rio the rabbit" as the character. Emotional data is also acquired from the user's facial expressions and voice. The device sends this data to an emotion engine, which analyzes it to recognize the user's current emotion. The results of this analysis are reflected in the theme and character settings.

[1523] Filtering input data

[1524] The server analyzes and filters all input data, including the theme and character settings received from the device and the emotion data obtained from the emotion engine. Specifically, a screening process is carried out for a list of prohibited words and explicit content. If inappropriate content is included, the server generates a message requesting the user to correct it and sends it to the device. If there is no inappropriate content, the process proceeds to the next step.

[1525] Story and illustration generation

[1526] The server generates a story by calling a generative AI model based on verified themes, character settings, and emotional data. For example, if the user is happy, a story with a happy ending is generated. This story is converted into text format, and then an image generation AI is used to generate illustrations corresponding to the story. Based on the emotional data, the illustrations also reflect emotional elements.

[1527] Multilingual Translation

[1528] The server invokes a translation engine to translate the generated story into a language specified by the user, such as Japanese, English, French, etc. The translated story is then reviewed by the server and any necessary corrections are made.

[1529] Story Distribution

[1530] The server generates digital picture book data that integrates the translated story and illustrations and transmits it to the user's terminal, which provides an interface for displaying the received digital picture book so that the user can view it.

[1531] Share your story

[1532] Users can share the digital picture book they have created with other users using the sharing function of their device. Sharing methods include social networking sites, email, and cloud services. The device sends the picture book data to the destination device, allowing other users to view it.

[1533] Hardware and software used

[1534] The server runs software such as generative AI models, emotion recognition engines, image generation AI, translation engines, and data filtering systems. Specific software and hardware used include the Google Translate API, EmotionRecognizer, StoryGenerator, and ImageGenerator. These work together to enable the generation, translation, and sharing of stories and illustrations based on user requests.

[1535] Specific examples

[1536] For example, if a parent or guardian selects characters Rio the Rabbit and Ken the Turtle as characters with the theme of "friendship" and the emotion at that time is "fun," a story about the fun adventures of Rio the Rabbit and Ken the Turtle is generated based on this emotion. The story is filtered and appropriate illustrations are generated, and then a digital picture book translated into Japanese, English, and French is delivered to the parent or guardian's smartphone. The parent or guardian can share this picture book with friends via social media.

[1537] Prompt Sentence Examples

[1538] Theme: Friendship

[1539] Character: Rio the Rabbit

[1540] Emotion data: {'emotion': 'joy'}

[1541] prompt:

[1542] "Rio the rabbit was having a great time playing with his friends. They decided to go on a new adventure together. Along the way, Rio..."

[1543] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1544] Step 1:

[1545] The user opens the device interface and inputs the theme and character. For example, the theme is set to "friendship" and the character is set to "Rio the rabbit." The user's facial expressions and voice data are also input. This input data is sent to the device (input data and emotion data).

[1546] Step 2:

[1547] The device processes the received theme, character, and emotion data and sends it to the server. The emotion data includes the results of the user's facial expression analysis and voice analysis. The device calls the emotion recognition engine to identify the user's emotion and sends the result as data (input data: theme, character, emotion data / output data: theme, character, emotion analysis result).

[1548] Step 3:

[1549] The server analyzes the theme, character, and emotion data received from the device and performs a filtering process. The filtering process screens the data based on a list of prohibited words and extreme content to check for inappropriate content. If inappropriate content is found, a message is generated requesting the user to correct it (input data: theme, character, emotion analysis results / output data: filtered data or correction request message).

[1550] Step 4:

[1551] Based on the filtered data, the server calls a generative AI model to generate a story. The generative AI model adjusts the tone and ending of the story based on the emotional data. For example, if the emotional data indicates joy, a story with a happy ending will be generated (input data: filtered data / output data: generated story).

[1552] Step 5:

[1553] Next, the server calls an image generation AI based on the generated story, which generates illustrations that correspond to the content of the story. Emotional data is also taken into account, and emotional elements are reflected in the illustrations (input data: generated story, emotion data / output data: generated illustrations).

[1554] Step 6:

[1555] The server invokes a translation engine to translate the generated story into multiple languages, for example, from Japanese to English and French. The translation is done automatically and can be manually corrected if necessary (input data: generated story / output data: translated story).

[1556] Step 7:

[1557] The server generates digital picture book data that integrates the translated story and generated illustrations and sends it to the user's device. The device provides an interface for displaying the received digital picture book, allowing the user to easily browse it (input data: translated story, generated illustrations / output data: digital picture book data).

[1558] Step 8:

[1559] Users can share the digital picture book they have created with other users using the sharing function of their device. Social networking sites, email, cloud services, and other methods of sharing are available. The device sends the picture book data to the destination device, allowing other users to view it (input data: digital picture book data; output data: data sent to the destination device).

[1560] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1561] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1562] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1563] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1564] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1565] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1566] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1567] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1568] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1569] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1570] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1571] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1572] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1573] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1574] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1575] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1576] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1577] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1578] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1579] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1580] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1581] The following is further disclosed regarding the above embodiment.

[1582] (Claim 1)

[1583] means for receiving input data;

[1584] means for filtering received input data;

[1585] A means of generating a narrative based on the filtered data;

[1586] means for generating images based on the generated narrative;

[1587] a means for translating the generated stories into multiple languages;

[1588] a means for delivering the translated story to the device;

[1589] A means for sharing the story delivered to the device with other users;

[1590] A system including:

[1591] (Claim 2)

[1592] means for accepting theme and character settings from a user;

[1593] means for automatically generating a story based on accepted themes and characters;

[1594] 10. The system of claim 1, comprising:

[1595] (Claim 3)

[1596] a means of simultaneously generating and integrating narrative and imagery;

[1597] A means of evaluating and ensuring the quality of the generated stories and images;

[1598] 10. The system of claim 1, comprising:

[1599] "Example 1"

[1600] (Claim 1)

[1601] an input means for allowing a user to set a theme and a character;

[1602] means for receiving, parsing, and filtering input data for the set theme and characters;

[1603] A means for generating a story using a generative AI model based on the filtered data; and

[1604] means for converting the generated narrative into a text format;

[1605] A means for generating illustrations using an image generation AI model based on the generated story;

[1606] a means of integrating the generated stories and illustrations into a coherent format;

[1607] a means for translating the generated stories into multiple languages;

[1608] a means for delivering the translated story to the device;

[1609] A means for sharing the story delivered to the device with other users;

[1610] A system including:

[1611] (Claim 2)

[1612] means for accepting theme and character settings from a user;

[1613] A means for automatically generating a story by generating prompt sentences using a generative AI model based on the received theme and characters;

[1614] 10. The system of claim 1, comprising:

[1615] (Claim 3)

[1616] a means of simultaneously generating and integrating narrative and imagery;

[1617] A means of evaluating and ensuring the quality of the generated stories and images;

[1618] 10. The system of claim 1, comprising:

[1619] "Application Example 1"

[1620] (Claim 1)

[1621] means for receiving input data;

[1622] means for filtering received input data;

[1623] A means of generating a narrative based on the filtered data;

[1624] means for generating images based on the generated narrative;

[1625] a means for translating the generated stories into multiple languages;

[1626] a means for delivering the translated story to the device;

[1627] A means for sharing the story delivered to the device with other users;

[1628] means for accepting theme and character settings from a user;

[1629] means for automatically generating a story based on accepted themes and characters;

[1630] A means of integrating the stories and images into a digital booklet format;

[1631] A system including:

[1632] (Claim 2)

[1633] a means of simultaneously generating and integrating narrative and imagery;

[1634] A means of evaluating and ensuring the quality of the generated stories and images;

[1635] 10. The system of claim 1, comprising:

[1636] (Claim 3)

[1637] A means for sharing the generated digital booklet data via social networking services and email;

[1638] 10. The system of claim 1, comprising:

[1639] "Example 2: Combining Emotion Engines"

[1640] The following is a rewritten version of the claims, including the characteristic features of the system.

[1641] (Claim 1)

[1642] means for receiving input data;

[1643] means for filtering received input data;

[1644] A means of generating a narrative based on the filtered data;

[1645] means for generating images based on the generated narrative;

[1646] a means for translating the generated stories into multiple languages;

[1647] a means for delivering the translated story to the device;

[1648] A means for sharing the story delivered to the device with other users;

[1649] A means for acquiring and analyzing user emotion data;

[1650] means for generating a narrative and images based on the acquired emotion data;

[1651] A system including:

[1652] (Claim 2)

[1653] means for accepting theme and character settings from a user;

[1654] means for automatically generating a story based on accepted themes and characters;

[1655] means for invoking a generative AI model based on the filtered data and the emotion data;

[1656] 10. The system of claim 1, comprising:

[1657] (Claim 3)

[1658] a means of simultaneously generating and integrating narrative and imagery;

[1659] A means of evaluating and ensuring the quality of the generated stories and images;

[1660] means for translating the narrative and images generated based on the emotion data into multiple languages;

[1661] A means to check and correct the quality of the translated data;

[1662] 10. The system of claim 1, comprising:

[1663] "Application example 2 when combining emotion engines"

[1664] (Claim 1)

[1665] means for receiving input data and emotion data;

[1666] means for filtering the received input data and emotion data;

[1667] means for generating a narrative based on the filtered data and the sentiment data;

[1668] means for generating images based on the generated narrative;

[1669] a means for translating the generated stories into multiple languages;

[1670] a means for delivering the translated story to the device;

[1671] A means for sharing the story delivered to the device with other users;

[1672] A system including:

[1673] (Claim 2)

[1674] means for accepting theme and character settings and emotion data from a user;

[1675] means for automatically generating a story based on accepted themes and characters;

[1676] a means for personalizing the story based on the received emotional data;

[1677] 10. The system of claim 1, comprising:

[1678] (Claim 3)

[1679] A means for simultaneously generating stories and images, integrating them, and evaluating them based on user emotion data to ensure quality;

[1680] 10. The system of claim 1, comprising: [Explanation of symbols]

[1681] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for receiving input data; means for filtering received input data; A means of generating a narrative based on the filtered data; means for generating images based on the generated narrative; a means for translating the generated stories into multiple languages; a means for delivering the translated story to the device; A means for sharing the story delivered to the device with other users; A system including:

2. means for accepting theme and character settings from a user; means for automatically generating a story based on accepted themes and characters; The system of claim 1 , comprising:

3. a means of simultaneously generating and integrating narrative and imagery; A means of evaluating and ensuring the quality of the generated stories and images; The system of claim 1 , comprising:

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A