System
The system addresses fragmented reading experiences by allowing users to input page numbers and questions for immediate answers and visual illustrations, improving comprehension and engagement.
Patent Information
- Application Number
- JP2024121466
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-26
- Publication Date
- 2026-02-05
AI Technical Summary
Existing reading systems lack efficient means for users to check story details, important hints, and character information in real-time, require separate searches for visual images, and fail to provide immediate answers to questions, leading to fragmented reading experiences.
A system that allows users to input the page number and questions, which are transmitted to a server for analysis, generating information on plot, foreshadowing, and characters, and enabling instant answers and visual illustrations using AI models.
Enables real-time checking of plot and character information, immediate answers to questions, and dynamic generation of visual images, enhancing the overall reading experience.
Smart Images

Figure 2026019718000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In today's world, readers often have many questions and requests while reading, creating demands for improved story comprehension and a better reading experience. These requests include grasping the plot as the story progresses, identifying foreshadowing, understanding the relationships between characters, and visualizing illustrations of specific scenes. Quick answers to questions are also required. The lack of efficient ways to satisfy these requests results in fragmented reading experiences, making it difficult to achieve a comprehensive understanding. [Means for solving the problem]
[0005] To solve the above problems, the present invention provides a system that includes: a means for a user to input the page number of a book they have finished reading; a means for transmitting the input page number to a server; a means for the server to analyze the content based on the page number and generate information about the plot, foreshadowing, and characters; a means for receiving and displaying information from the server; a means for a user to input a question and transmit it to the server; a means for the server to generate and transmit an answer based on the question; a means for a user to request the generation of illustrations; and a means for the server to generate and transmit illustrations using an AI model. This allows users to constantly check information about the plot, foreshadowing, and characters while reading, and quickly obtain answers to questions. Furthermore, the generation of illustrations aids visual comprehension, improving the overall reading experience.
[0006] A "user" is an entity that uses the system of the present invention to read books and perform operations such as inputting information, asking questions, and making requests.
[0007] The "read page number" is the number of the page that the user recognizes as having finished reading.
[0008] A "terminal" is a device that a user uses to read or operate something, including, for example, a smartphone, tablet, or computer.
[0009] A "server" is a device that has the function of generating and analyzing information in response to user input or requests and sending it to a terminal.
[0010] A "plot" is a summary of a story up to a certain point, a concise description of the main events.
[0011] A "foreshadowing" is an element that refers to important information or hints related to later developments in the story.
[0012] "Characters" refers to the characters or subjects who act or speak within a story.
[0013] A "question" is a question that a user enters to clarify something they are unsure about in the story or the system.
[0014] An "answer" is a response or explanatory information generated by a server in response to a question.
[0015] An "illustration" is an illustration or diagram that visually shows a specific scene or content of a story.
[0016] A "request" is a request that a user sends to a system for a particular action or service.
[0017] An "AI model" is a computational algorithm or program that uses artificial intelligence to accomplish a specific task. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram illustrating a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] The present invention is a system designed to enhance a user's reading experience, and is embodied in the following manner.
[0040] System Overview
[0041] This system has three main components: the user, the terminal, and the server. The user inputs the page number and question they have read into the terminal, which then sends it to the server. The server analyzes the received information, generates appropriate data, and sends it to the terminal to provide information to the user.
[0042] Enter the page number the user has finished reading
[0043] The user uses the terminal to input the page number they have just finished reading. This information is sent to the server via the terminal. At this stage, the user is ready to receive information from the system based on their progress in the story.
[0044] Server analysis and information generation
[0045] The server analyzes the received page number and generates a list of the plot, hints, and characters up to that page. This allows the user to review the main points of the story or notice hints. The information generated by the server is compiled in the form of plot, hints, and character information and sent to the terminal.
[0046] Terminal display
[0047] The device displays the information received from the server in a user-friendly interface, allowing users to easily understand the main points of the story and the relationships between characters. If users require specific information or detailed explanations, they can type their questions into the device.
[0048] User questions and responses
[0049] When a user inputs a question, the device sends the question to the server, which extracts relevant information from existing databases and content based on the question and generates an answer. This answer is then immediately sent to the device, which displays it to the user, allowing the user to deepen their understanding of the story.
[0050] Dynamic generation of illustrations
[0051] If a user wants a visual image of a particular scene, they can request the generation of an illustration by pressing the "Image" button. The device sends this request to the server, which uses an AI model to generate an illustration for the scene. The generated illustration is sent to the device and displayed to the user. This feature allows users to visually enjoy the story scenes.
[0052] Specific examples
[0053] For example, suppose a user is reading a particular novel. When the user inputs into the system that he or she has read up to page 56, the server receives and analyzes that information, generates information about the plot, foreshadowing, and characters up to that point, and sends it to the device. If the user asks, "Was there any information about the perpetrator's motive?", the server will provide an answer to that question. Furthermore, if the user wants to see an illustration of a particular scene, the illustration will be generated by pressing the "Image" button.
[0054] The above is an embodiment of the present invention, which provides a specific method for improving the user's reading experience.
[0055] The processing flow will be explained below.
[0056] Step 1:
[0057] The user inputs the page number of the page that has been read. The user inputs the page number of the page that has been read in the terminal interface. For example, the user inputs "56".
[0058] Step 2:
[0059] The terminal transmits the input page number to the server. The terminal generates an HTTP request for transmitting the input page number to the server and transmits it to the server.
[0060] Step 3:
[0061] The server analyzes the received page number and retrieves the data up to that page from the database or content management system based on the received page number.
[0062] Step 4:
[0063] The server analyzes the content and generates plot, foreshadowing, and character information. The server uses a data analysis algorithm to identify plot, foreshadowing, and characters based on page numbers and compiles each piece of information.
[0064] Step 5:
[0065] The server sends the generated information to the terminal. The server then sends the generated information about the plot, hints, and characters to the terminal as an HTTP response.
[0066] Step 6:
[0067] The terminal receives and displays information from the server. The terminal receives a response from the server and displays the plot, plot twists, and character information on the user interface.
[0068] Step 7:
[0069] The user enters a question. The user enters a question into the question input field on the device. For example, the user enters a question such as "What is the perpetrator's motive?"
[0070] Step 8:
[0071] The device sends the question to the server. The device sends the user's question to the server as an HTTP request.
[0072] Step 9:
[0073] The server generates an answer based on the question. The server searches for information related to the question from a database or existing content and generates an answer.
[0074] Step 10:
[0075] The server sends the generated answer to the terminal. The server sends the generated answer to the terminal as an HTTP response.
[0076] Step 11:
[0077] The terminal displays the answer from the server. The terminal receives the response from the server and displays the answer to the user on the user interface.
[0078] Step 12:
[0079] The user requests the creation of an illustration. If the user wants to see an illustration of a particular scene, they press the "Image" button.
[0080] Step 13:
[0081] The terminal sends a request to generate an illustration to the server. The terminal detects the button press event and sends a request to generate an illustration to the server.
[0082] Step 14:
[0083] The server uses the AI model to generate illustrations. The server uses the AI model for generating illustrations to generate illustrations based on the specified page content.
[0084] Step 15:
[0085] The server sends the generated illustration to the terminal. The server generates and sends an HTTP response to send the generated illustration to the terminal.
[0086] Step 16:
[0087] The terminal receives and displays the illustration. The terminal receives the illustration data from the server and displays it to the user on the user interface.
[0088] Example 1
[0089] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0090] Previous systems lacked sufficient means for users to check story details and important hints while reading, or to rely on character information to understand the story. Furthermore, obtaining visual images required separate searches, which was time-consuming. Furthermore, few systems could provide immediate answers to user questions or dynamically generate information based on the page being read. Therefore, an effective system to enhance users' reading experience was needed.
[0091] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0092] In this invention, the server includes means for inputting the page number of a page that a user has finished reading, means for transmitting the input page number to the server, means for the server to analyze the content based on the page number and generate information on the plot, foreshadowing, and characters, means for receiving and displaying information from the server, means for inputting a user's question and transmitting it to the server, means for the server to generate and transmit an answer based on the question, means for the user to request the generation of an illustration, means for the server to generate and transmit an illustration using an AI model, and means for the terminal to display information in a user-friendly interface. This enables the user to check the plot, foreshadowing, and information on characters in real time while reading, get instant answers to questions, and generate visual images.
[0093] "User" means a person who uses the system and takes action to enhance their reading experience.
[0094] A "terminal" is a device that allows a user to perform input operations and receive and display information from a server.
[0095] A "server" is a central computing device that analyzes information sent from users and terminals and generates and provides the necessary data.
[0096] The "page number read" is the number of the page that the user currently recognizes as having finished reading.
[0097] "Input means" refers to the interface or device that allows the user to input the page number they have read or questions into the terminal.
[0098] "Transmission means" refers to the function or device for transmitting information entered by the user from the terminal to the server.
[0099] The "analysis means" refers to a function that allows the server to process data based on the input page number and analyze the content.
[0100] The "synopsis generation means" refers to a function or method for summarizing the main points of the story up to the page that the server has finished reading.
[0101] The "foreshadowing generation means" refers to a function or method by which the server creates a list of important foreshadowings that appear in the story.
[0102] "Character information generation means" refers to the function or method by which the server analyzes and lists the characters in the story and their relationships.
[0103] "Display means" refers to the functions and methods by which the terminal presents the information received from the server in an easy-to-read format for the user.
[0104] "Question input means" refers to an interface or device that allows a user to input detailed information or specific questions into a terminal.
[0105] "Answer generation means" refers to the function or method by which the server generates related information based on the user's question and creates an answer.
[0106] The "illustration generation request means" refers to a function or method that allows a user to send a request from a terminal to a server in order to generate a visual image of a specific scene.
[0107] "AI model" refers to the artificial intelligence algorithms and functions that the server uses to generate illustrations and other data.
[0108] A "user-friendly interface" refers to the screen or method by which a device presents information to the user in an intuitive and easy-to-understand format.
[0109] This invention is a system designed to improve the user's reading experience, and it includes three main components: the terminal, the server, and the user. The following describes each flow in detail.
[0110] First, the user inputs the page number they have finished reading and a question into the terminal. The terminal is a mobile device such as a tablet or smartphone, which receives and processes the input information. The terminal is equipped with a user interface and is designed for intuitive operation. Specifically, the user might input "I have finished reading up to page 56" or a question such as "What is the perpetrator's motive?"
[0111] The device then sends the input information to the server, where it transmits the data via HTTP over Wi-Fi or mobile networks, and organizes the data in JSON format for efficient delivery to the server.
[0112] The server is a high-performance data server, such as AWS EC2, a widely used cloud service. The server searches the database based on the page number sent by the user and analyzes the content up to that page. This analysis uses an algorithm that utilizes natural language processing technology to dynamically generate the story's plot, hints, and character information. This allows the user to see the progress of the story and understand its key points.
[0113] When a user enters a question, the server generates an answer based on that question. Here, machine learning technology is used to extract relevant information from existing databases and derive appropriate answers. For example, in response to the question "What was the perpetrator's motive?", the server searches the database for relevant information, extracts the relevant parts, and generates an answer.
[0114] Furthermore, if a user wants a visual image of a specific scene, they can request the generation of an illustration by pressing the "Image" button on their device. This request is sent to the server, which then uses a generative AI model (e.g., DALL-E 2) to generate an illustration of the scene. The AI model generates an image of a specific scene based on a prompt. For example, a prompt might be, "Please generate an illustration of the scene where the protagonist confronts the criminal for the first time."
[0115] The information and illustrations generated by the server are then sent back to the device, which receives them and displays them in a user-friendly interface, allowing users to enjoy the story visually as well.
[0116] (Example of a prompt)
[0117] "I've finished reading page 56 so far. Please tell me the plot, foreshadowing, and characters."
[0118] "Did you have any information about the perpetrator's motive?"
[0119] "Generate an illustration of the scene where the protagonist confronts the culprit for the first time."
[0120] As described above, this system performs appropriate data processing and calculations based on the information entered by the user, providing the user with a fulfilling reading experience.
[0121] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0122] Step 1:
[0123] The user inputs the page number they have read and a question into the device. The user inputs the page number into the device's input field in the format of "I have now finished reading page 56." They also input questions such as "What was the perpetrator's motive?" The input is in text format, and the device receives it and stores it in its internal memory.
[0124] Step 2:
[0125] The device sends the input information to the server. The device converts the page number and question information entered by the user into JSON format and sends it to the server via an HTTP request. Specifically, the data is sent using Wi-Fi or a mobile network.
[0126] Input: A page number or question entered by the user
[0127] Output: Data converted to JSON format
[0128] Specific operation: The device converts the text data into JSON format and sends it to the server using an HTTP request.
[0129] Step 3:
[0130] The server analyzes the page number and generates a synopsis, foreshadowing, and character information. The server then analyzes the received JSON data and retrieves the content up to the relevant page from the database. Based on the retrieved data, the server uses natural language processing technology to analyze the content and generate information to provide to the user.
[0131] Input: JSON data of page numbers received from the terminal
[0132] Output: Generated plot, hints, and character information
[0133] Specific operation: The server searches the database, analyzes the retrieved data, and generates the required information.
[0134] Step 4:
[0135] The server sends the generated information to the terminal. The generated plot, foreshadowing, and character information are compiled in JSON format and sent to the terminal as an HTTP response.
[0136] Input: Generated plot, hints, character information
[0137] Output: Data in JSON format
[0138] Specific operation: The information generated by the server is converted into JSON format and sent to the terminal via an HTTP response.
[0139] Step 5:
[0140] The device displays the received information in a user-friendly manner. The device analyzes the received JSON data and presents the plot, plot twists, and character information to the user in a visually easy-to-understand format.
[0141] Input: JSON data received from the server
[0142] Output: Information displayed in a user-friendly format
[0143] Specific operation: The device parses the JSON data and displays it on the screen.
[0144] Step 6:
[0145] The user enters a follow-up question into the device, which then sends it to the server. The user enters a follow-up question into the device, such as "What is your specific motivation?" The device receives the question and sends it back to the server.
[0146] Input: Any additional questions entered by the user
[0147] Output: The question sent to the server
[0148] Specific operation: The device receives a new question, converts it into JSON format, and sends it to the server.
[0149] Step 7:
[0150] The server analyzes the question, generates relevant information, and sends it to the device. The server analyzes the question, extracts relevant information from an existing database, and generates an answer. The generated answer is compiled in JSON format and sent to the device.
[0151] Input: JSON data of the question received from the terminal
[0152] Output: JSON data of the generated answer
[0153] What it does: The server searches the database, extracts relevant information, and generates an answer.
[0154] Step 8:
[0155] The terminal displays the answer and, if necessary, sends a request to the server to generate an illustration. The terminal displays the answer to the user, and if necessary, the user presses the "Image" button to request the generation of an illustration.
[0156] Input: JSON data of the response received from the server
[0157] Output: Answers displayed to the user and illustration generation request
[0158] Specific operation: The device displays the answer on the screen, and when the user presses the "Image" button, it sends a request to the server to generate an illustration.
[0159] Step 9:
[0160] The server generates illustrations and sends them to the device. The server receives illustration generation requests and generates illustrations for the corresponding scenes using a generative AI model. The generated illustrations are then sent to the device.
[0161] Input: Illustration generation request received from the device
[0162] Output: Generated illustration data
[0163] Specific operation: The server sends a prompt to the AI model and sends the generated illustration data to the terminal.
[0164] Step 10:
[0165] The terminal displays the generated illustration. The terminal displays the illustration data received from the server to the user, allowing the user to visually enjoy a specific scene in the story.
[0166] Input: illustration data received from the server
[0167] Output: The illustration displayed to the user
[0168] Specific operation: The device converts the illustration data for display on the screen and shows it to the user.
[0169] These are the specific processing steps of the present invention, which allow users to check plot summaries, hints, and character information in real time while reading, get instant answers to questions, and generate visual images.
[0170] (Application example 1)
[0171] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0172] Conventional reading systems make it difficult for users to obtain summaries, plot twists, and character information according to their reading progress. Furthermore, they lack the ability to generate illustrations to visually enjoy specific scenes, limiting the reading experience. Furthermore, when users have questions about the story, there is a lack of a way to get immediate answers. It is necessary to solve these problems and improve the user's reading experience.
[0173] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0174] In this invention, the server includes means for analyzing the content based on the page number that the user has read and generating information on the plot, foreshadowing, and characters, means for generating answers based on the user's questions, and means for generating illustrations using a generative AI model. This allows the user to easily obtain a summary of the content, foreshadowing, and character information according to their reading progress, as well as enjoy visual images of specific scenes and receive immediate answers to questions that will help them better understand the story.
[0175] "User" refers to an individual who utilizes the system to enhance their reading experience.
[0176] "Page number read" refers to the number of the specific page of the book that the user has currently read.
[0177] "Server" refers to the central processing unit that analyzes information sent by users and generates and transmits necessary data.
[0178] A "plot" is a concise summary of the main developments and events of a story.
[0179] A "foreshadowing" is information that provides hints or elements that will become important later in the development of the story.
[0180] "Character information" refers to detailed information about the characters' names, backgrounds, and roles in the story.
[0181] "Question" refers to the act of a user inquiring of the server about specific matters regarding the content of the story.
[0182] "Answer" refers to information generated by the server based on the user's question.
[0183] An "illustration" is an image or painting that visually depicts a particular scene from a story.
[0184] A "generative AI model" refers to an algorithm or system that uses artificial intelligence technology to generate visual images or answers in response to user requests.
[0185] "Prompt sentence" refers to the text input used to prompt a generative AI model to generate an illustration.
[0186] "Display means" refers to an interface for displaying the generated information and illustrations in a form that can be visually recognized by the user.
[0187] This invention is a system designed to improve the user's reading experience, and it includes three main elements: a user, a terminal, and a server. The user uses the terminal to input the page number they have read and questions, and the terminal sends that information to the server. The server analyzes the received information, generates appropriate data, and sends it to the terminal to provide information to the user.
[0188] System Overview
[0189] The system includes the following elements:
[0190] 1. Enter page number:
[0191] The user inputs the page number they have read into the device, which allows the device to provide information based on the progress of the story.
[0192] 2. Information analysis and generation:
[0193] The server parses the page number it receives and generates a list of plot, plot twists, and characters up to that page, then compiles this information in a user-friendly format and sends it to the device.
[0194] 3. Questions and Answers:
[0195] When a user enters a question, the device sends it to the server, which extracts relevant information from an existing database and immediately generates an answer that is sent to the device.
[0196] 4. Dynamic generation of illustrations:
[0197] When a user wants a visual image of a particular scene, they press the "Image" button to request the generation of an illustration. The server uses a generative AI model to generate an illustration of the scene, which is then sent to the device for display.
[0198] Hardware and software used
[0199] Hardware:
[0200] User devices: smartphones, tablets
[0201] Server: Server equipment that performs high-performance calculations
[0202] software:
[0203] Server API: e.g. Flask or Django
[0204] AI model: Generative AI models such as Stable Diffusion and DALL-E are used to generate illustrations.
[0205] HTTP request library: Requests
[0206] Example of a system
[0207] 1. Enter and analyze page numbers:
[0208] For example, if a user inputs into the system that they have read a particular novel and have progressed to page 56, the server receives and analyzes that information, generating information about the plot, foreshadowing, and characters up to that point, and sending it to the terminal.
[0209] Example prompt: "Use case: I've read up to page 56. What's the plot and foreshadowing so far?"
[0210] 2. Questions and Answers:
[0211] If a user asks, "Is there any information about the perpetrator's motive?", the server searches for relevant information based on the question, generates an answer, and sends it to the device.
[0212] Example prompt: "Use case: I want to ask for information about the perpetrator's motive."
[0213] 3. Illustration generation:
[0214] If a user wants to see a visual image of a particular scene, they can press the "Visualize" button, which will generate an illustration. For example, a request to generate an illustration of a scene in which the protagonist is in conflict is sent to the server, and the illustration is generated and displayed using a generative AI model.
[0215] Example prompt: "Use case: Visualize a scene where the protagonist is in conflict."
[0216] The above is an embodiment of the present invention, which provides a specific method for improving the user's reading experience.
[0217] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0218] Step 1:
[0219] The user uses the terminal to enter the page number that they have read.
[0220] Input: page number read
[0221] Specific behavior: The user enters the page number into the input field on the device and presses the "Submit" button.
[0222] Step 2:
[0223] The terminal transmits the input page number to the server.
[0224] Input: The page number entered by the user
[0225] Output: Page number data sent to the server
[0226] Specific operation: The terminal sends the entered page number to the server via an HTTP request.
[0227] Step 3:
[0228] The server analyzes the page number received and generates information on the plot, plot twists, and characters up to that page.
[0229] Input: Submitted page number data
[0230] Output: Synopsis, foreshadowing, character information
[0231] Specific operation: The server searches the database, analyzes the content up to the relevant page, and generates information on the plot, foreshadowing, and characters.
[0232] Step 4:
[0233] The server transmits the generated information to the terminal.
[0234] Input: Synopsis, foreshadowing, character information
[0235] Output: Information data sent to the terminal
[0236] Specific operation: The server sends the generated information to the terminal in a data format such as JSON.
[0237] Step 5:
[0238] The terminal displays the information received from the server in a user-friendly interface.
[0239] Input: Information data sent from the server
[0240] Output: Plot, plot twists, and character information displayed on the screen
[0241] Specific operation: The device analyzes the received information and generates and displays HTML and CSS to display on the interface.
[0242] Step 6:
[0243] The user enters a specific question, and the device sends the question to the server.
[0244] Input: The question entered by the user
[0245] Output: Question data sent to the server
[0246] Specific operation: The user enters a question in the question input field and presses the "Send" button. The device sends the question to the server via an HTTP request.
[0247] Step 7:
[0248] The server extracts relevant information from existing databases based on the question and generates an answer.
[0249] Input: Submitted question data
[0250] Output: relevant information and generated answers
[0251] What it does: The server analyzes the question, extracts relevant information from the database, and generates an answer.
[0252] Step 8:
[0253] The server generates a response and sends it to the terminal.
[0254] Input: Generated answer
[0255] Output: Answer data sent to the device
[0256] Specific operation: The server converts the generated response into a data format and sends it to the terminal.
[0257] Step 9:
[0258] The device displays the response received from the server to the user.
[0259] Input: Response data sent from the server
[0260] Output: Answer displayed on the screen
[0261] Specific operation: The device analyzes the received response and generates and displays an interface to display to the user.
[0262] Step 10:
[0263] A user requests a visual image of a particular scene, and the device sends the request to the server.
[0264] Input: The prompt for the scene the user wants
[0265] Output: Request data sent to the server
[0266] Specific operation: The user inputs a prompt sentence about a specific scene and presses the "Image" button. The device sends a request to the server.
[0267] Step 11:
[0268] The server uses a generative AI model to generate illustrations of the requested scene.
[0269] Input: Request data (prompt statement)
[0270] Output: Generated illustration data
[0271] Specific operation: The server inputs a prompt sentence into the generative AI model and receives the generated illustration.
[0272] Step 12:
[0273] The server sends the generated illustration to the terminal.
[0274] Input: Generated illustration data
[0275] Output: Illustration data sent to the device
[0276] Specific operation: The server converts the illustration data into an appropriate format and sends it to the terminal.
[0277] Step 13:
[0278] The terminal displays the illustrations received from the server in a user-friendly interface.
[0279] Input: Illustration data sent from the server
[0280] Output: Illustration displayed on the screen
[0281] Specific operation: The terminal analyzes the received illustration data and generates and displays an interface for visual display.
[0282] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0283] The present invention is an advanced system for improving a user's reading experience, incorporating an emotion engine for recognizing a user's emotions and providing corresponding information. Specific embodiments for implementing the present invention will be described below.
[0284] System Overview
[0285] This system includes four main components: the user, the device, the server, and the emotion engine. The user inputs the page number they have read and a question into the device, which then sends it to the server. The server analyzes the received information, generates appropriate data, and sends it to the device. During this process, the emotion engine analyzes the user's emotions and adjusts the information display and response content.
[0286] Enter the page number the user has finished reading
[0287] The user uses the terminal to input the page number they have just finished reading. This information is sent to the server via the terminal. The server also recognizes the emotion (e.g., excitement, sadness, doubt, etc.) with which the user seeks information.
[0288] Server analysis and information generation
[0289] The server analyzes the received page number and generates a list of the plot, foreshadowing, and characters up to that page. The information generated by the server is organized in the form of plot, foreshadowing, and character information and sent to the device. During this process, the emotion engine analyzes the user's emotions and adjusts the display method and content of the information.
[0290] Terminal display
[0291] The device receives information from the server and displays it in a user-friendly interface. The emotion engine customizes the interface based on the user's emotions. For example, if the user is excited, the display may change to a different color or larger font size.
[0292] User questions and responses
[0293] When a user enters a question, the device sends it to the server. Based on the question, the server extracts relevant information from existing databases and content to generate an answer. During this process, the emotion engine analyzes the user's emotions and adjusts the content and expression of the answer. For example, if the user is sad, it can incorporate words of encouragement. This answer is sent to the device and displayed to the user.
[0294] Dynamic generation of illustrations
[0295] If a user wants a visual image of a particular scene, they can request the generation of an illustration by pressing the "Image" button. The device sends this request to the server, which uses an AI model to generate an illustration of the scene. The generated illustration takes the user's emotions into account using an emotion engine, and is sent to the device and displayed to the user.
[0296] Specific examples
[0297] For example, suppose a user is reading a particular novel. The user enters into the system that they have read up to page 56, and the emotion engine recognizes that the user is currently excited. The server receives and analyzes this information, and generates information on the plot, foreshadowing, and characters accordingly. The emotion engine customizes this information to match the user's state of excitement and sends it to the device. If the user asks, "Was there any information on the perpetrator's motive?", the server generates an answer to the question that takes the emotion engine into consideration. Furthermore, if the user wants to see an illustration of a particular scene, they can press the "Visualize" button, and an illustration will be generated. The emotion engine customizes the illustration according to the user's emotion and displays it on the device.
[0298] The above is an embodiment of the present invention, which provides a specific method for taking user emotions into consideration to improve the reading experience.
[0299] The processing flow will be explained below.
[0300] Step 1:
[0301] The user inputs the page number of the page that has been read. The user inputs the page number of the page that has been read in the terminal interface. For example, the user inputs "56".
[0302] Step 2:
[0303] The device sends the entered page number and the user's emotional state to the server. The device then generates an HTTP request to send the emotional data analyzed by the emotion engine and the page number to the server, and sends it to the server.
[0304] Step 3:
[0305] The server analyzes the received page number and emotional state. Based on the received page number, the server retrieves data up to that page from a database or content management system and analyzes the user's emotions.
[0306] Step 4:
[0307] The server generates information on the plot, plot twists, and characters, and then adjusts it based on the user's emotions using an emotion engine.The server uses a data analysis algorithm to identify the plot, plot twists, and characters based on page numbers, and the emotion engine customizes the information according to the user's emotions.
[0308] Step 5:
[0309] The server sends the generated information to the terminal. The server then sends the generated plot, foreshadowing, and character information to the terminal as emotion-adjusted information in an HTTP response.
[0310] Step 6:
[0311] The terminal receives and displays information from the server. The terminal receives a response from the server and displays customized plot, plot twists, and character information on the user interface.
[0312] Step 7:
[0313] The user enters a question. The user enters a question into the question input field on the device. For example, the user enters a question such as "What is the perpetrator's motive?"
[0314] Step 8:
[0315] The device sends the question and emotional state to the server. The device generates an HTTP request to send the question including the user's emotional data to the server and sends it to the server.
[0316] Step 9:
[0317] The server generates an answer based on the question and adjusts it based on the user's emotions. The server searches for information related to the question from a database or existing content and generates an answer. The generated answer is adjusted according to the user's emotions by an emotion engine.
[0318] Step 10:
[0319] The server sends the generated answer to the terminal, and the server sends the adjusted answer to the terminal as an HTTP response.
[0320] Step 11:
[0321] The device displays the answer from the server. The device receives the response from the server and displays the answer adjusted to match the user's emotions on the user interface.
[0322] Step 12:
[0323] The user requests the creation of an illustration. If the user wants to see an illustration of a particular scene, they press the "Image" button.
[0324] Step 13:
[0325] The device sends a request to generate an illustration and the user's emotional state to the server. The device detects the button press event and sends a request to generate an illustration and the user's emotional state to the server.
[0326] Step 14:
[0327] The server uses an AI model to generate illustrations and adjust them based on the user's emotions.The server uses an AI model for generating illustrations to generate illustrations based on the specified page content, and the emotion engine adjusts them according to the user's emotions.
[0328] Step 15:
[0329] The server sends the generated illustration to the device. The server then generates and sends an HTTP response to send the emotion-adjusted illustration to the device.
[0330] Step 16:
[0331] The terminal receives and displays the illustration. The terminal receives the illustration data from the server and displays the illustration adjusted based on the emotion on the user interface.
[0332] Example 2
[0333] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0334] Previous systems for improving the reading experience did not provide information that took the user's emotions into consideration, and they had the problem of being unable to display or respond appropriately to the user's emotional state. Furthermore, the generation of visual images was not tailored to the user's needs, and only a uniform response was possible. This limited the user's reading experience, making it difficult to provide a customized experience tailored to each individual's emotional state.
[0335] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for generating information on a summary, foreshadowing, and characters based on page numbers, a means for adjusting the information content and display method based on emotion analysis, and a means for generating visual images using a generative AI model. This makes it possible to provide and display information according to the user's emotional state, and to provide a reading experience customized for each user.
[0336] A "user" is someone who uses the system to enhance their reading experience.
[0337] A "terminal" is a device that allows a user to input the page number of a page that has been read, input a question, or request the generation of a visual image.
[0338] The "server" is a central processing unit that analyzes the information sent by the user, generates answers and visual images, and sends them to the terminal.
[0339] A "page number" is a number that indicates the specific page of a book that the user has finished reading.
[0340] "Emotion analysis means" refers to technology for recognizing and analyzing a user's emotional state.
[0341] A "generative AI model" is an algorithm or software that uses AI technology to generate visual images based on specific prompts.
[0342] "Visual images" are illustrations or drawings of specific scenes requested by users.
[0343] A "summary" is a short summary of the contents up to a specific page number.
[0344] A "foreshadowing" is an element in a story that contains information or hints that will be important for later developments.
[0345] "Characters" is a list and profiles of characters that appear in the book.
[0346] A "question" is a question about specific information that a user inputs into the system.
[0347] "Means for adjusting information content and display method" refers to technology that changes the display format and content of information provided based on the user's emotional state.
[0348] This invention is a system for improving a user's reading experience, and in particular, has the function of customizing information according to the user's emotional state. This system mainly consists of a user (terminal), a terminal, a server, and an emotion analysis engine.
[0349] Enter the page number the user has read
[0350] The user inputs the page number they have finished reading into the device. The device receives this information and also analyzes the user's emotional state using an emotion analysis engine. For example, the device uses emotion recognition technology such as Microsoft Azure's Emotion API. The analyzed data (page number and emotional state) is sent to the server.
[0351] Server analysis and information generation
[0352] The server analyzes the received page number and emotion data. Based on the page number, a natural language processing engine (such as OpenAI's GPT-4) is used to generate a summary, plot twists, and character information. The emotion analysis engine also analyzes the user's emotional state and adjusts the display and content of the information. For example, if the user is excited, the text color or font size is changed. This adjusted information is then sent to the device.
[0353] Terminal display
[0354] The device displays the information received from the server in a user-friendly interface, which can be provided as a web page or a dedicated application. Information customized by the sentiment analysis engine (such as text color and font size) is also reflected here.
[0355] User questions and responses
[0356] The user enters a specific question into the device and presses the send button. The device then sends this question to the server. The server receives the question, extracts relevant information from the database and content, and generates an answer. Again, the sentiment analysis engine analyzes the user's emotions and adjusts the content and expression of the answer. The generated answer is sent to the device and displayed to the user.
[0357] Dynamic generation of illustrations
[0358] When a user requests a visualization of a particular scene, they can press the "Image" button to send the request to the device. The device then sends this request to the server. The server uses a generative AI model (e.g., DALL-E or Stable Diffusion) to generate a visual image based on the prompt. This visual image is also customized to take into account the user's emotional state. The generated visual image is then sent to the device and displayed to the user.
[0359] Specific examples
[0360] For example, suppose a user is reading a particular novel and has read up to page 56. The user enters "I've read up to page 56" into the device and presses the send button. At this point, the emotion analysis engine recognizes that the user is excited. The server receives and analyzes this information, generating a summary, foreshadowing, and character information accordingly. The emotion analysis engine customizes the information to match the user's state of excitement and sends it to the device. If the user also enters a question such as "Was there any information about the perpetrator's motive?", the server will generate an answer to the question that takes the emotion analysis engine into consideration. Furthermore, if the user wishes to visualize a particular scene, a visual image can be generated by pressing the "Visualize" button. The emotion analysis engine customizes the visual image according to the user's emotion and displays it on the device.
[0361] Prompt Sentence Examples
[0362] "I'm 56 pages in. Can you give me a summary of the novel and what the characters are? If the sentiment analysis engine thinks I'm excited, please make sure the information is presented in a way that's enjoyable for the user."
[0363] The above is an embodiment of the present invention.
[0364] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0365] Step 1:
[0366] Enter the page number that the user has read.
[0367] Specific operation: The user enters "page 56" into the terminal and presses the send button. The terminal receives this input data and recognizes the user's emotional state (e.g., excitement) through an emotion analysis engine. The obtained data (page number and emotional state) is sent to the server.
[0368] Input: page number (56) and emotional state (excitement).
[0369] Data processing: Add emotional data using a sentiment analysis engine.
[0370] Output: Data including page number and emotional state.
[0371] Step 2:
[0372] The server analyzes the received page number and emotional state.
[0373] Specific operation: The server analyzes the received data and uses a natural language processing engine such as GPT-4 to generate a summary, foreshadowing, and character information up to the relevant page. In addition, an emotion analysis engine analyzes the user's emotional state and adjusts the method and content of information display.
[0374] Input: Data including page number and emotional state.
[0375] Data processing: Use a natural language processing engine to generate summaries, plot twists, and character information, and customize the information with a sentiment analysis engine.
[0376] Output: Customized summary, foreshadowing, and character information.
[0377] Step 3:
[0378] The terminal displays the information received from the server.
[0379] Specific operation: The device displays the information received from the server in a user-friendly interface (e.g., a web page or dedicated app). Information customized by the emotion analysis engine is also reflected here. For example, the text color may be brightened and the font size increased depending on the state of excitement.
[0380] Input: Customized summary, foreshadowing, and character information.
[0381] Data Calculation: Convert customized information into the appropriate display format.
[0382] Output: Display of adjusted information.
[0383] Step 4:
[0384] The user types in a question and sends it to the server.
[0385] Specific operation: The user enters a specific question into the terminal (e.g., "Is there any information about the perpetrator's motive?") and presses the send button. The terminal then sends this question to the server.
[0386] Input: A question entered by the user (e.g., "Was there any information on the perpetrator's motive?").
[0387] Data processing: The entered question is sent to the server.
[0388] Output: The query data sent to the server.
[0389] Step 5:
[0390] The server receives the query and generates the relevant information.
[0391] Specific operation: The server analyzes the question and extracts relevant information from the database and content. At this time, the emotion analysis engine analyzes the user's emotional state and adjusts the content and expression of the answer. The generated answer is then sent to the device.
[0392] Input: The query data sent to the server.
[0393] Data processing: Uses database search and sentiment analysis engines to generate relevant answers.
[0394] Output: The customized answer.
[0395] Step 6:
[0396] The device displays the received response to the user.
[0397] Specific operation: The device displays the answers received from the server in a user-friendly format. Customized answers are also applied using the sentiment analysis engine.
[0398] Input: Customized Answer.
[0399] Data Calculation: Convert customized answers into the appropriate display format.
[0400] Output: A display of the appropriately adjusted answer.
[0401] Step 7:
[0402] The user requests the generation of a visual image.
[0403] Specific operation: The user presses the "Image" button on the device to request the device to visualize a specific scene. The device then sends this request to the server.
[0404] Input: Visual image generation requested by the user.
[0405] Data processing: Send the request to the server.
[0406] Output: The visual image generation request sent to the server.
[0407] Step 8:
[0408] The server generates the visual image and sends it to the terminal.
[0409] How it works: The server uses a generative AI model (e.g., DALL-E or Stable Diffusion) to generate a visual image based on the prompt. This image is also customized to take into account the user's emotional state. The generated image is then sent to the device and displayed to the user.
[0410] Input: A visual image generation request.
[0411] Data processing: Visual images are generated using generative AI models and customized with a sentiment analysis engine.
[0412] Output: A customized visual image.
[0413] (Application example 2)
[0414] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0415] Conventional reading support systems have difficulty customizing the user's reading experience, and are particularly limited in providing information tailored to the user's emotions. As a result, users are unable to receive the information they desire quickly and appropriately, resulting in a poor quality reading experience. Furthermore, the generation of illustrations, which are visual representations, is fixed and not dynamically generated to match the user's emotions. There was a need for a system that could solve these issues and improve the user's reading experience.
[0416] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0417] In this invention, the server includes a means for analyzing the content based on the page number and generating information on the plot, foreshadowing, and characters, a means for generating and transmitting illustrations using a generative AI model, and a means for analyzing the user's emotions using an emotion engine and adjusting the information display method, thereby enabling information display and dynamic illustration generation according to the user's emotions.
[0418] "User" refers to an individual user of the system.
[0419] "Page number read" refers to the specific page number of the book that the user has currently read.
[0420] "Input means" refers to the interfaces and devices through which a user provides information to the system.
[0421] "Server" refers to a computer system that receives requests from users, analyzes them, and provides appropriate information.
[0422] "Means for transmitting to the server" refers to a communication means for transmitting the user's input information to the server.
[0423] "Means for analyzing the content and generating plot, plot twists, and character information" refers to the process by which the server analyzes the content of the book based on page numbers and generates summary information to provide to the user.
[0424] "Displaying means" refers to devices and interfaces for visually presenting information received from the server to a user.
[0425] "Means for inputting a question and sending it to the server" refers to an interface and communication means for a user to input a question to the system and send the question to the server.
[0426] "Means for generating and sending answers based on questions" refers to the process by which the server analyzes the user's question, derives an appropriate answer, and provides it to the user.
[0427] "Means for requesting the generation of illustrations" refers to an interface through which a user requests the system to create a visual image of a particular scene.
[0428] "Means for generating and transmitting illustrations using a generative AI model" refers to the process by which a server uses AI technology to dynamically generate illustrations and provide those illustrations to users.
[0429] An "emotion engine" is an engine that analyzes a user's emotions and adjusts the way information is displayed based on those emotions.
[0430] "Means of analyzing emotions and adjusting how information is displayed" refers to the process by which the emotion engine understands the user's current emotional state and changes the content and format of the display accordingly.
[0431] System Overview
[0432] This system includes four main components: the user, the device, the server, and the emotion engine. The user inputs the page number they have read and a question into the device, which then sends it to the server. The server analyzes the received information, generates appropriate data, and sends it to the device. During this process, the emotion engine analyzes the user's emotions and adjusts the information display and response content accordingly.
[0433] Enter the page number the user has finished reading
[0434] The user uses the terminal to input the page number they have just finished reading. This information is sent to the server via the terminal. The server also recognizes the emotion (e.g., excitement, sadness, doubt, etc.) with which the user seeks information.
[0435] Server analysis and information generation
[0436] The server analyzes the received page number and generates a list of the plot, foreshadowing, and characters up to that page. The information generated by the server is organized in the form of plot, foreshadowing, and character information and sent to the device. During this process, the emotion engine analyzes the user's emotions and adjusts the display method and content of the information.
[0437] Terminal display
[0438] The device displays the information received from the server in a user-friendly interface. The emotion engine customizes the interface based on the user's emotions. For example, if the user is excited, the display may change to a larger font size or change the color of the text. If the user is sad, the display may use warmer colors.
[0439] User questions and responses
[0440] When a user enters a question, the device sends it to the server. Based on the question, the server extracts relevant information from existing databases and content to generate an answer. During this process, the emotion engine analyzes the user's emotions and adjusts the content and expression of the answer. For example, if the user is sad, it can incorporate words of encouragement. This answer is sent to the device and displayed to the user.
[0441] Dynamic generation of illustrations
[0442] If a user wants a visual image of a particular scene, they can request the generation of an illustration by pressing the "Image" button. The device sends this request to the server, which then uses a generative AI model to generate an illustration of the scene. The generated illustration, which takes the user's emotions into account using an emotion engine, is sent to the device and displayed to the user.
[0443] Specific examples
[0444] For example, suppose a user is reading a particular novel. The user enters into the system that they have read up to page 56, and the emotion engine recognizes that the user is currently excited. The server receives and analyzes this information, and generates information on the plot, foreshadowing, and characters accordingly. The emotion engine customizes this information to match the user's state of excitement and sends it to the device. If the user asks, "Was there any information on the perpetrator's motive?", the server generates an answer to the question that takes the emotion engine into consideration. Furthermore, if the user wants to see an illustration of a particular scene, they can press the "Visualize" button, and an illustration will be generated. The emotion engine customizes the illustration according to the user's emotion and displays it on the device.
[0445] The specific hardware and software used
[0446] Hardware:
[0447] Smartphone (iOS or Android)
[0448] Server (equipped with a high-performance GPU)
[0449] Emotion engine (commonly known as "emotion analysis engine")
[0450] software:
[0451] Frontend: React Native
[0452] Backend: Node.js, Express
[0453] Database: MongoDB
[0454] Emotion analysis: Python, TensorFlow
[0455] Prompt Sentence Examples
[0456] "Page 56, Emotion: Excite, Question: Was there any information about the perpetrator's motive?"
[0457] The above is a specific embodiment for implementing this system.
[0458] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0459] Program processing steps
[0460] Step 1:
[0461] The user enters the page number they have read into the terminal.
[0462] Input: page number the user finished reading, emotional state (e.g., excited, sad)
[0463] Specific operation: The user enters the page number of the page they have just finished reading in the page number input interface of the application, selects their emotional state, and the result is input to the terminal.
[0464] Step 2:
[0465] The terminal transmits the input page number and emotional state to the server.
[0466] Input: User-entered page number and emotional state
[0467] Output: Data packet sent to the server
[0468] Specific operation: The device sends the input page number and emotional state as a data packet to the server using an HTTP POST request.
[0469] Step 3:
[0470] The server analyzes the content based on the page number and generates plot summary, foreshadowing, and character information.
[0471] Input: The page number received by the server
[0472] Output: Synopsis, foreshadowing, character information data
[0473] Specific operation: The server retrieves information up to the relevant page from the database, and based on this generates a plot summary, foreshadowing, and a list of characters. A Python script is executed.
[0474] Step 4:
[0475] The server uses an emotion engine to analyze the user's emotions and adjust how information is displayed.
[0476] Input: User's emotional state, generated information data
[0477] Output: Information data adjusted based on user sentiment
[0478] How it works: The emotion engine (a model using TensorFlow) analyzes the emotional state and adjusts the format and display style of the generated information data. For example, if the user is excited, the font size and color will be changed.
[0479] Step 5:
[0480] The server transmits the adjusted information data to the terminal.
[0481] Input: Adjusted information data
[0482] Output: Data packets sent to the device
[0483] Specific operation: The server assembles the adjusted information data into a data packet and sends it to the terminal. The communication method is an HTTP POST request.
[0484] Step 6:
[0485] The information received by the terminal is displayed in a user-friendly interface.
[0486] Input: Adjusted information data received from the server
[0487] Output: On-screen display
[0488] Specific operation: Based on the received information data, the device displays information through a customized interface according to the user's emotional state. The display process is performed using React Native.
[0489] Step 7:
[0490] The user enters a question and the device sends the question to the server.
[0491] Input: The question entered by the user
[0492] Output: Data packet sent to the server
[0493] Specific operation: The user inputs a question through the question input interface, and the device sends the question to the server via an HTTP POST request.
[0494] Step 8:
[0495] The server extracts relevant information from existing databases based on the question and generates an answer.
[0496] Input: User question
[0497] Output: Response data
[0498] What happens: The server searches a database for information related to the question and generates an answer. A Python script is executed to extract the requested information.
[0499] Step 9:
[0500] The emotion engine analyzes the user's emotions and adjusts the content and expression of the response.
[0501] Input: User's emotional state, generated answer data
[0502] Output: Answer data adjusted based on user sentiment
[0503] What it does: The emotion engine analyzes the user's emotional state and adjusts the wording of the generated response data accordingly, possibly adding words of encouragement.
[0504] Step 10:
[0505] The server transmits the adjusted response data to the terminal.
[0506] Input: Adjusted response data
[0507] Output: Data packets sent to the device
[0508] Specific operation: The server assembles the adjusted response data into a data packet and sends it to the terminal using an HTTP POST request.
[0509] Step 11:
[0510] The user requests the creation of an illustration, and the device sends the request to the server.
[0511] Input: Illustration generation request
[0512] Output: Data packet sent to the server
[0513] Specific operation: The user presses the "Image" button to request the creation of an illustration, and the device sends the request to the server via an HTTP POST request.
[0514] Step 12:
[0515] The server uses a generative AI model to generate and transmit illustrations.
[0516] Input: Illustration generation request
[0517] Output: Generated illustration data
[0518] Specific operation: The server uses a generative AI model (e.g., DALL-E) to generate illustrations according to the request. The generated illustration data is prepared.
[0519] Step 13:
[0520] The emotion engine customizes the generated illustrations according to the user's emotions.
[0521] Input: Generated illustration data, user's emotional state
[0522] Output: Adjusted illustration data
[0523] How it works: The emotion engine analyzes the user's emotional state and customizes the generated illustrations accordingly. For example, if the user is excited, the colors and details will be enhanced.
[0524] Step 14:
[0525] The server sends the adjusted illustration data to the terminal, which displays it to the user.
[0526] Input: Adjusted illustration data
[0527] Output: Illustration displayed on the terminal
[0528] Specific operation: The server assembles the adjusted illustration data into a data packet and sends it to the device. The device displays the received illustration data in a user-friendly interface. The display process is performed using React Native.
[0529] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0530] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0531] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0532] [Second embodiment]
[0533] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0534] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0535] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0536] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0537] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0538] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0539] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0540] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0541] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0542] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0543] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0544] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0545] The present invention is a system designed to enhance a user's reading experience, and is embodied in the following manner.
[0546] System Overview
[0547] This system has three main components: the user, the terminal, and the server. The user inputs the page number and question they have read into the terminal, which then sends it to the server. The server analyzes the received information, generates appropriate data, and sends it to the terminal to provide information to the user.
[0548] Enter the page number the user has finished reading
[0549] The user uses the terminal to input the page number they have just finished reading. This information is sent to the server via the terminal. At this stage, the user is ready to receive information from the system based on their progress in the story.
[0550] Server analysis and information generation
[0551] The server analyzes the received page number and generates a list of the plot, hints, and characters up to that page. This allows the user to review the main points of the story or notice hints. The information generated by the server is compiled in the form of plot, hints, and character information and sent to the terminal.
[0552] Terminal display
[0553] The device displays the information received from the server in a user-friendly interface, allowing users to easily understand the main points of the story and the relationships between characters. If users require specific information or detailed explanations, they can type their questions into the device.
[0554] User questions and responses
[0555] When a user inputs a question, the device sends the question to the server, which extracts relevant information from existing databases and content based on the question and generates an answer. This answer is then immediately sent to the device, which displays it to the user, allowing the user to deepen their understanding of the story.
[0556] Dynamic generation of illustrations
[0557] If a user wants a visual image of a particular scene, they can request the generation of an illustration by pressing the "Image" button. The device sends this request to the server, which uses an AI model to generate an illustration for the scene. The generated illustration is sent to the device and displayed to the user. This feature allows users to visually enjoy the story scenes.
[0558] Specific examples
[0559] For example, suppose a user is reading a particular novel. When the user inputs into the system that he or she has read up to page 56, the server receives and analyzes that information, generates information about the plot, foreshadowing, and characters up to that point, and sends it to the device. If the user asks, "Was there any information about the perpetrator's motive?", the server will provide an answer to that question. Furthermore, if the user wants to see an illustration of a particular scene, the illustration will be generated by pressing the "Image" button.
[0560] The above is an embodiment of the present invention, which provides a specific method for improving the user's reading experience.
[0561] The processing flow will be explained below.
[0562] Step 1:
[0563] The user inputs the page number of the page that has been read. The user inputs the page number of the page that has been read in the terminal interface. For example, the user inputs "56".
[0564] Step 2:
[0565] The terminal transmits the input page number to the server. The terminal generates an HTTP request for transmitting the input page number to the server and transmits it to the server.
[0566] Step 3:
[0567] The server analyzes the received page number and retrieves the data up to that page from the database or content management system based on the received page number.
[0568] Step 4:
[0569] The server analyzes the content and generates plot, foreshadowing, and character information. The server uses a data analysis algorithm to identify plot, foreshadowing, and characters based on page numbers and compiles each piece of information.
[0570] Step 5:
[0571] The server sends the generated information to the terminal. The server then sends the generated information about the plot, hints, and characters to the terminal as an HTTP response.
[0572] Step 6:
[0573] The terminal receives and displays information from the server. The terminal receives a response from the server and displays the plot, plot twists, and character information on the user interface.
[0574] Step 7:
[0575] The user enters a question. The user enters a question into the question input field on the device. For example, the user enters a question such as "What is the perpetrator's motive?"
[0576] Step 8:
[0577] The device sends the question to the server. The device sends the user's question to the server as an HTTP request.
[0578] Step 9:
[0579] The server generates an answer based on the question. The server searches for information related to the question from a database or existing content and generates an answer.
[0580] Step 10:
[0581] The server sends the generated answer to the terminal. The server sends the generated answer to the terminal as an HTTP response.
[0582] Step 11:
[0583] The terminal displays the answer from the server. The terminal receives the response from the server and displays the answer to the user on the user interface.
[0584] Step 12:
[0585] The user requests the creation of an illustration. If the user wants to see an illustration of a particular scene, they press the "Image" button.
[0586] Step 13:
[0587] The terminal sends a request to generate an illustration to the server. The terminal detects the button press event and sends a request to generate an illustration to the server.
[0588] Step 14:
[0589] The server uses the AI model to generate illustrations. The server uses the AI model for generating illustrations to generate illustrations based on the specified page content.
[0590] Step 15:
[0591] The server sends the generated illustration to the terminal. The server generates and sends an HTTP response to send the generated illustration to the terminal.
[0592] Step 16:
[0593] The terminal receives and displays the illustration. The terminal receives the illustration data from the server and displays it to the user on the user interface.
[0594] Example 1
[0595] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0596] Previous systems lacked sufficient means for users to check story details and important hints while reading, or to rely on character information to understand the story. Furthermore, obtaining visual images required separate searches, which was time-consuming. Furthermore, few systems could provide immediate answers to user questions or dynamically generate information based on the page being read. Therefore, an effective system to enhance users' reading experience was needed.
[0597] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0598] In this invention, the server includes means for inputting the page number of a page that a user has finished reading, means for transmitting the input page number to the server, means for the server to analyze the content based on the page number and generate information on the plot, foreshadowing, and characters, means for receiving and displaying information from the server, means for inputting a user's question and transmitting it to the server, means for the server to generate and transmit an answer based on the question, means for the user to request the generation of an illustration, means for the server to generate and transmit an illustration using an AI model, and means for the terminal to display information in a user-friendly interface. This enables the user to check the plot, foreshadowing, and information on characters in real time while reading, get instant answers to questions, and generate visual images.
[0599] "User" means a person who uses the system and takes action to enhance their reading experience.
[0600] A "terminal" is a device that allows a user to perform input operations and receive and display information from a server.
[0601] A "server" is a central computing device that analyzes information sent from users and terminals and generates and provides the necessary data.
[0602] The "page number read" is the number of the page that the user currently recognizes as having finished reading.
[0603] "Input means" refers to the interface or device that allows the user to input the page number they have read or questions into the terminal.
[0604] "Transmission means" refers to the function or device for transmitting information entered by the user from the terminal to the server.
[0605] The "analysis means" refers to a function that allows the server to process data based on the input page number and analyze the content.
[0606] The "synopsis generation means" refers to a function or method for summarizing the main points of the story up to the page that the server has finished reading.
[0607] The "foreshadowing generation means" refers to a function or method by which the server creates a list of important foreshadowings that appear in the story.
[0608] "Character information generation means" refers to the function or method by which the server analyzes and lists the characters in the story and their relationships.
[0609] "Display means" refers to the functions and methods by which the terminal presents the information received from the server in an easy-to-read format for the user.
[0610] "Question input means" refers to an interface or device that allows a user to input detailed information or specific questions into a terminal.
[0611] "Answer generation means" refers to the function or method by which the server generates related information based on the user's question and creates an answer.
[0612] The "illustration generation request means" refers to a function or method that allows a user to send a request from a terminal to a server in order to generate a visual image of a specific scene.
[0613] "AI model" refers to the artificial intelligence algorithms and functions that the server uses to generate illustrations and other data.
[0614] A "user-friendly interface" refers to the screen or method by which a device presents information to the user in an intuitive and easy-to-understand format.
[0615] This invention is a system designed to improve the user's reading experience, and it includes three main components: the terminal, the server, and the user. The following describes each flow in detail.
[0616] First, the user inputs the page number they have finished reading and a question into the terminal. The terminal is a mobile device such as a tablet or smartphone, which receives and processes the input information. The terminal is equipped with a user interface and is designed for intuitive operation. Specifically, the user might input "I have finished reading up to page 56" or a question such as "What is the perpetrator's motive?"
[0617] The device then sends the input information to the server, where it transmits the data via HTTP over Wi-Fi or mobile networks, and organizes the data in JSON format for efficient delivery to the server.
[0618] The server is a high-performance data server, such as AWS EC2, a widely used cloud service. The server searches the database based on the page number sent by the user and analyzes the content up to that page. This analysis uses an algorithm that utilizes natural language processing technology to dynamically generate the story's plot, hints, and character information. This allows the user to see the progress of the story and understand its key points.
[0619] When a user enters a question, the server generates an answer based on that question. Here, machine learning technology is used to extract relevant information from existing databases and derive appropriate answers. For example, in response to the question "What was the perpetrator's motive?", the server searches the database for relevant information, extracts the relevant parts, and generates an answer.
[0620] Furthermore, if a user wants a visual image of a specific scene, they can request the generation of an illustration by pressing the "Image" button on their device. This request is sent to the server, which then uses a generative AI model (e.g., DALL-E 2) to generate an illustration of the scene. The AI model generates an image of a specific scene based on a prompt. For example, a prompt might be, "Please generate an illustration of the scene where the protagonist confronts the criminal for the first time."
[0621] The information and illustrations generated by the server are then sent back to the device, which receives them and displays them in a user-friendly interface, allowing users to enjoy the story visually as well.
[0622] (Example of a prompt)
[0623] "I've finished reading page 56 so far. Please tell me the plot, foreshadowing, and characters."
[0624] "Did you have any information about the perpetrator's motive?"
[0625] "Generate an illustration of the scene where the protagonist confronts the culprit for the first time."
[0626] As described above, this system performs appropriate data processing and calculations based on the information entered by the user, providing the user with a fulfilling reading experience.
[0627] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0628] Step 1:
[0629] The user inputs the page number they have read and a question into the device. The user inputs the page number into the device's input field in the format of "I have now finished reading page 56." They also input questions such as "What was the perpetrator's motive?" The input is in text format, and the device receives it and stores it in its internal memory.
[0630] Step 2:
[0631] The device sends the input information to the server. The device converts the page number and question information entered by the user into JSON format and sends it to the server via an HTTP request. Specifically, the data is sent using Wi-Fi or a mobile network.
[0632] Input: A page number or question entered by the user
[0633] Output: Data converted to JSON format
[0634] Specific operation: The device converts the text data into JSON format and sends it to the server using an HTTP request.
[0635] Step 3:
[0636] The server analyzes the page number and generates a synopsis, foreshadowing, and character information. The server then analyzes the received JSON data and retrieves the content up to the relevant page from the database. Based on the retrieved data, the server uses natural language processing technology to analyze the content and generate information to provide to the user.
[0637] Input: JSON data of page numbers received from the terminal
[0638] Output: Generated plot, hints, and character information
[0639] Specific operation: The server searches the database, analyzes the retrieved data, and generates the required information.
[0640] Step 4:
[0641] The server sends the generated information to the terminal. The generated plot, foreshadowing, and character information are compiled in JSON format and sent to the terminal as an HTTP response.
[0642] Input: Generated plot, hints, character information
[0643] Output: Data in JSON format
[0644] Specific operation: The information generated by the server is converted into JSON format and sent to the terminal via an HTTP response.
[0645] Step 5:
[0646] The device displays the received information in a user-friendly manner. The device analyzes the received JSON data and presents the plot, plot twists, and character information to the user in a visually easy-to-understand format.
[0647] Input: JSON data received from the server
[0648] Output: Information displayed in a user-friendly format
[0649] Specific operation: The device parses the JSON data and displays it on the screen.
[0650] Step 6:
[0651] The user enters a follow-up question into the device, which then sends it to the server. The user enters a follow-up question into the device, such as "What is your specific motivation?" The device receives the question and sends it back to the server.
[0652] Input: Any additional questions entered by the user
[0653] Output: The question sent to the server
[0654] Specific operation: The device receives a new question, converts it into JSON format, and sends it to the server.
[0655] Step 7:
[0656] The server analyzes the question, generates relevant information, and sends it to the device. The server analyzes the question, extracts relevant information from an existing database, and generates an answer. The generated answer is compiled in JSON format and sent to the device.
[0657] Input: JSON data of the question received from the terminal
[0658] Output: JSON data of the generated answer
[0659] What it does: The server searches the database, extracts relevant information, and generates an answer.
[0660] Step 8:
[0661] The terminal displays the answer and, if necessary, sends a request to the server to generate an illustration. The terminal displays the answer to the user, and if necessary, the user presses the "Image" button to request the generation of an illustration.
[0662] Input: JSON data of the response received from the server
[0663] Output: Answers displayed to the user and illustration generation request
[0664] Specific operation: The device displays the answer on the screen, and when the user presses the "Image" button, it sends a request to the server to generate an illustration.
[0665] Step 9:
[0666] The server generates illustrations and sends them to the device. The server receives illustration generation requests and generates illustrations for the corresponding scenes using a generative AI model. The generated illustrations are then sent to the device.
[0667] Input: Illustration generation request received from the device
[0668] Output: Generated illustration data
[0669] Specific operation: The server sends a prompt to the AI model and sends the generated illustration data to the terminal.
[0670] Step 10:
[0671] The terminal displays the generated illustration. The terminal displays the illustration data received from the server to the user, allowing the user to visually enjoy a specific scene in the story.
[0672] Input: illustration data received from the server
[0673] Output: The illustration displayed to the user
[0674] Specific operation: The device converts the illustration data for display on the screen and shows it to the user.
[0675] These are the specific processing steps of the present invention, which allow users to check plot summaries, hints, and character information in real time while reading, get instant answers to questions, and generate visual images.
[0676] (Application example 1)
[0677] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0678] Conventional reading systems make it difficult for users to obtain summaries, plot twists, and character information according to their reading progress. Furthermore, they lack the ability to generate illustrations to visually enjoy specific scenes, limiting the reading experience. Furthermore, when users have questions about the story, there is a lack of a way to get immediate answers. It is necessary to solve these problems and improve the user's reading experience.
[0679] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0680] In this invention, the server includes means for analyzing the content based on the page number that the user has read and generating information on the plot, foreshadowing, and characters, means for generating answers based on the user's questions, and means for generating illustrations using a generative AI model. This allows the user to easily obtain a summary of the content, foreshadowing, and character information according to their reading progress, as well as enjoy visual images of specific scenes and receive immediate answers to questions that will help them better understand the story.
[0681] "User" refers to an individual who utilizes the system to enhance their reading experience.
[0682] "Page number read" refers to the number of the specific page of the book that the user has currently read.
[0683] "Server" refers to the central processing unit that analyzes information sent by users and generates and transmits necessary data.
[0684] A "plot" is a concise summary of the main developments and events of a story.
[0685] A "foreshadowing" is information that provides hints or elements that will become important later in the development of the story.
[0686] "Character information" refers to detailed information about the characters' names, backgrounds, and roles in the story.
[0687] "Question" refers to the act of a user inquiring of the server about specific matters regarding the content of the story.
[0688] "Answer" refers to information generated by the server based on the user's question.
[0689] An "illustration" is an image or painting that visually depicts a particular scene from a story.
[0690] A "generative AI model" refers to an algorithm or system that uses artificial intelligence technology to generate visual images or answers in response to user requests.
[0691] "Prompt sentence" refers to the text input used to prompt a generative AI model to generate an illustration.
[0692] "Display means" refers to an interface for displaying the generated information and illustrations in a form that can be visually recognized by the user.
[0693] This invention is a system designed to improve the user's reading experience, and it includes three main elements: a user, a terminal, and a server. The user uses the terminal to input the page number they have read and questions, and the terminal sends that information to the server. The server analyzes the received information, generates appropriate data, and sends it to the terminal to provide information to the user.
[0694] System Overview
[0695] The system includes the following elements:
[0696] 1. Enter page number:
[0697] The user inputs the page number they have read into the device, which allows the device to provide information based on the progress of the story.
[0698] 2. Information analysis and generation:
[0699] The server parses the page number it receives and generates a list of plot, plot twists, and characters up to that page, then compiles this information in a user-friendly format and sends it to the device.
[0700] 3. Questions and Answers:
[0701] When a user enters a question, the device sends it to the server, which extracts relevant information from an existing database and immediately generates an answer that is sent to the device.
[0702] 4. Dynamic generation of illustrations:
[0703] When a user wants a visual image of a particular scene, they press the "Image" button to request the generation of an illustration. The server uses a generative AI model to generate an illustration of the scene, which is then sent to the device for display.
[0704] Hardware and software used
[0705] Hardware:
[0706] User devices: smartphones, tablets
[0707] Server: Server equipment that performs high-performance calculations
[0708] software:
[0709] Server API: e.g. Flask or Django
[0710] AI model: Generative AI models such as Stable Diffusion and DALL-E are used to generate illustrations.
[0711] HTTP request library: Requests
[0712] Example of a system
[0713] 1. Enter and analyze page numbers:
[0714] For example, if a user inputs into the system that they have read a particular novel and have progressed to page 56, the server receives and analyzes that information, generating information about the plot, foreshadowing, and characters up to that point, and sending it to the terminal.
[0715] Example prompt: "Use case: I've read up to page 56. What's the plot and foreshadowing so far?"
[0716] 2. Questions and Answers:
[0717] If a user asks, "Is there any information about the perpetrator's motive?", the server searches for relevant information based on the question, generates an answer, and sends it to the device.
[0718] Example prompt: "Use case: I want to ask for information about the perpetrator's motive."
[0719] 3. Illustration generation:
[0720] If a user wants to see a visual image of a particular scene, they can press the "Visualize" button, which will generate an illustration. For example, a request to generate an illustration of a scene in which the protagonist is in conflict is sent to the server, and the illustration is generated and displayed using a generative AI model.
[0721] Example prompt: "Use case: Visualize a scene where the protagonist is in conflict."
[0722] The above is an embodiment of the present invention, which provides a specific method for improving the user's reading experience.
[0723] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0724] Step 1:
[0725] The user uses the terminal to enter the page number that they have read.
[0726] Input: page number read
[0727] Specific behavior: The user enters the page number into the input field on the device and presses the "Submit" button.
[0728] Step 2:
[0729] The terminal transmits the input page number to the server.
[0730] Input: The page number entered by the user
[0731] Output: Page number data sent to the server
[0732] Specific operation: The terminal sends the entered page number to the server via an HTTP request.
[0733] Step 3:
[0734] The server analyzes the page number received and generates information on the plot, plot twists, and characters up to that page.
[0735] Input: Submitted page number data
[0736] Output: Synopsis, foreshadowing, character information
[0737] Specific operation: The server searches the database, analyzes the content up to the relevant page, and generates information on the plot, foreshadowing, and characters.
[0738] Step 4:
[0739] The server transmits the generated information to the terminal.
[0740] Input: Synopsis, foreshadowing, character information
[0741] Output: Information data sent to the terminal
[0742] Specific operation: The server sends the generated information to the terminal in a data format such as JSON.
[0743] Step 5:
[0744] The terminal displays the information received from the server in a user-friendly interface.
[0745] Input: Information data sent from the server
[0746] Output: Plot, plot twists, and character information displayed on the screen
[0747] Specific operation: The device analyzes the received information and generates and displays HTML and CSS to display on the interface.
[0748] Step 6:
[0749] The user enters a specific question, and the device sends the question to the server.
[0750] Input: The question entered by the user
[0751] Output: Question data sent to the server
[0752] Specific operation: The user enters a question in the question input field and presses the "Send" button. The device sends the question to the server via an HTTP request.
[0753] Step 7:
[0754] The server extracts relevant information from existing databases based on the question and generates an answer.
[0755] Input: Submitted question data
[0756] Output: relevant information and generated answers
[0757] What it does: The server analyzes the question, extracts relevant information from the database, and generates an answer.
[0758] Step 8:
[0759] The server generates a response and sends it to the terminal.
[0760] Input: Generated answer
[0761] Output: Answer data sent to the device
[0762] Specific operation: The server converts the generated response into a data format and sends it to the terminal.
[0763] Step 9:
[0764] The device displays the response received from the server to the user.
[0765] Input: Response data sent from the server
[0766] Output: Answer displayed on the screen
[0767] Specific operation: The device analyzes the received response and generates and displays an interface to display to the user.
[0768] Step 10:
[0769] A user requests a visual image of a particular scene, and the device sends the request to the server.
[0770] Input: The prompt for the scene the user wants
[0771] Output: Request data sent to the server
[0772] Specific operation: The user inputs a prompt sentence about a specific scene and presses the "Image" button. The device sends a request to the server.
[0773] Step 11:
[0774] The server uses a generative AI model to generate illustrations of the requested scene.
[0775] Input: Request data (prompt statement)
[0776] Output: Generated illustration data
[0777] Specific operation: The server inputs a prompt sentence into the generative AI model and receives the generated illustration.
[0778] Step 12:
[0779] The server sends the generated illustration to the terminal.
[0780] Input: Generated illustration data
[0781] Output: Illustration data sent to the device
[0782] Specific operation: The server converts the illustration data into an appropriate format and sends it to the terminal.
[0783] Step 13:
[0784] The terminal displays the illustrations received from the server in a user-friendly interface.
[0785] Input: Illustration data sent from the server
[0786] Output: Illustration displayed on the screen
[0787] Specific operation: The terminal analyzes the received illustration data and generates and displays an interface for visual display.
[0788] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0789] The present invention is an advanced system for improving a user's reading experience, incorporating an emotion engine for recognizing a user's emotions and providing corresponding information. Specific embodiments for implementing the present invention will be described below.
[0790] System Overview
[0791] This system includes four main components: the user, the device, the server, and the emotion engine. The user inputs the page number they have read and a question into the device, which then sends it to the server. The server analyzes the received information, generates appropriate data, and sends it to the device. During this process, the emotion engine analyzes the user's emotions and adjusts the information display and response content.
[0792] Enter the page number the user has finished reading
[0793] The user uses the terminal to input the page number they have just finished reading. This information is sent to the server via the terminal. The server also recognizes the emotion (e.g., excitement, sadness, doubt, etc.) with which the user seeks information.
[0794] Server analysis and information generation
[0795] The server analyzes the received page number and generates a list of the plot, foreshadowing, and characters up to that page. The information generated by the server is organized in the form of plot, foreshadowing, and character information and sent to the device. During this process, the emotion engine analyzes the user's emotions and adjusts the display method and content of the information.
[0796] Terminal display
[0797] The device receives information from the server and displays it in a user-friendly interface. The emotion engine customizes the interface based on the user's emotions. For example, if the user is excited, the display may change to a different color or larger font size.
[0798] User questions and responses
[0799] When a user enters a question, the device sends it to the server. Based on the question, the server extracts relevant information from existing databases and content to generate an answer. During this process, the emotion engine analyzes the user's emotions and adjusts the content and expression of the answer. For example, if the user is sad, it can incorporate words of encouragement. This answer is sent to the device and displayed to the user.
[0800] Dynamic generation of illustrations
[0801] If a user wants a visual image of a particular scene, they can request the generation of an illustration by pressing the "Image" button. The device sends this request to the server, which uses an AI model to generate an illustration of the scene. The generated illustration takes the user's emotions into account using an emotion engine, and is sent to the device and displayed to the user.
[0802] Specific examples
[0803] For example, suppose a user is reading a particular novel. The user enters into the system that they have read up to page 56, and the emotion engine recognizes that the user is currently excited. The server receives and analyzes this information, and generates information on the plot, foreshadowing, and characters accordingly. The emotion engine customizes this information to match the user's state of excitement and sends it to the device. If the user asks, "Was there any information on the perpetrator's motive?", the server generates an answer to the question that takes the emotion engine into consideration. Furthermore, if the user wants to see an illustration of a particular scene, they can press the "Visualize" button, and an illustration will be generated. The emotion engine customizes the illustration according to the user's emotion and displays it on the device.
[0804] The above is an embodiment of the present invention, which provides a specific method for taking user emotions into consideration to improve the reading experience.
[0805] The processing flow will be explained below.
[0806] Step 1:
[0807] The user inputs the page number of the page that has been read. The user inputs the page number of the page that has been read in the terminal interface. For example, the user inputs "56".
[0808] Step 2:
[0809] The device sends the entered page number and the user's emotional state to the server. The device then generates an HTTP request to send the emotional data analyzed by the emotion engine and the page number to the server, and sends it to the server.
[0810] Step 3:
[0811] The server analyzes the received page number and emotional state. Based on the received page number, the server retrieves data up to that page from a database or content management system and analyzes the user's emotions.
[0812] Step 4:
[0813] The server generates information on the plot, plot twists, and characters, and then adjusts it based on the user's emotions using an emotion engine.The server uses a data analysis algorithm to identify the plot, plot twists, and characters based on page numbers, and the emotion engine customizes the information according to the user's emotions.
[0814] Step 5:
[0815] The server sends the generated information to the terminal. The server then sends the generated plot, foreshadowing, and character information to the terminal as emotion-adjusted information in an HTTP response.
[0816] Step 6:
[0817] The terminal receives and displays information from the server. The terminal receives a response from the server and displays customized plot, plot twists, and character information on the user interface.
[0818] Step 7:
[0819] The user enters a question. The user enters a question into the question input field on the device. For example, the user enters a question such as "What is the perpetrator's motive?"
[0820] Step 8:
[0821] The device sends the question and emotional state to the server. The device generates an HTTP request to send the question including the user's emotional data to the server and sends it to the server.
[0822] Step 9:
[0823] The server generates an answer based on the question and adjusts it based on the user's emotions. The server searches for information related to the question from a database or existing content and generates an answer. The generated answer is adjusted according to the user's emotions by an emotion engine.
[0824] Step 10:
[0825] The server sends the generated answer to the terminal, and the server sends the adjusted answer to the terminal as an HTTP response.
[0826] Step 11:
[0827] The device displays the answer from the server. The device receives the response from the server and displays the answer adjusted to match the user's emotions on the user interface.
[0828] Step 12:
[0829] The user requests the creation of an illustration. If the user wants to see an illustration of a particular scene, they press the "Image" button.
[0830] Step 13:
[0831] The device sends a request to generate an illustration and the user's emotional state to the server. The device detects the button press event and sends a request to generate an illustration and the user's emotional state to the server.
[0832] Step 14:
[0833] The server uses an AI model to generate illustrations and adjust them based on the user's emotions.The server uses an AI model for generating illustrations to generate illustrations based on the specified page content, and the emotion engine adjusts them according to the user's emotions.
[0834] Step 15:
[0835] The server sends the generated illustration to the device. The server then generates and sends an HTTP response to send the emotion-adjusted illustration to the device.
[0836] Step 16:
[0837] The terminal receives and displays the illustration. The terminal receives the illustration data from the server and displays the illustration adjusted based on the emotion on the user interface.
[0838] Example 2
[0839] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0840] Previous systems for improving the reading experience did not provide information that took the user's emotions into consideration, and they had the problem of being unable to display or respond appropriately to the user's emotional state. Furthermore, the generation of visual images was not tailored to the user's needs, and only a uniform response was possible. This limited the user's reading experience, making it difficult to provide a customized experience tailored to each individual's emotional state.
[0841] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for generating information on a summary, foreshadowing, and characters based on page numbers, a means for adjusting the information content and display method based on emotion analysis, and a means for generating visual images using a generative AI model. This makes it possible to provide and display information according to the user's emotional state, and to provide a reading experience customized for each user.
[0842] A "user" is someone who uses the system to enhance their reading experience.
[0843] A "terminal" is a device that allows a user to input the page number of a page that has been read, input a question, or request the generation of a visual image.
[0844] The "server" is a central processing unit that analyzes the information sent by the user, generates answers and visual images, and sends them to the terminal.
[0845] A "page number" is a number that indicates the specific page of a book that the user has finished reading.
[0846] "Emotion analysis means" refers to technology for recognizing and analyzing a user's emotional state.
[0847] A "generative AI model" is an algorithm or software that uses AI technology to generate visual images based on specific prompts.
[0848] "Visual images" are illustrations or drawings of specific scenes requested by users.
[0849] A "summary" is a short summary of the contents up to a specific page number.
[0850] A "foreshadowing" is an element in a story that contains information or hints that will be important for later developments.
[0851] "Characters" is a list and profiles of characters that appear in the book.
[0852] A "question" is a question about specific information that a user inputs into the system.
[0853] "Means for adjusting information content and display method" refers to technology that changes the display format and content of information provided based on the user's emotional state.
[0854] This invention is a system for improving a user's reading experience, and in particular, has the function of customizing information according to the user's emotional state. This system mainly consists of a user (terminal), a terminal, a server, and an emotion analysis engine.
[0855] Enter the page number the user has read
[0856] The user inputs the page number they have finished reading into the device. The device receives this information and also analyzes the user's emotional state using an emotion analysis engine. For example, the device uses emotion recognition technology such as Microsoft Azure's Emotion API. The analyzed data (page number and emotional state) is sent to the server.
[0857] Server analysis and information generation
[0858] The server analyzes the received page number and emotion data. Based on the page number, a natural language processing engine (such as OpenAI's GPT-4) is used to generate a summary, plot twists, and character information. The emotion analysis engine also analyzes the user's emotional state and adjusts the display and content of the information. For example, if the user is excited, the text color or font size is changed. This adjusted information is then sent to the device.
[0859] Terminal display
[0860] The device displays the information received from the server in a user-friendly interface, which can be provided as a web page or a dedicated application. Information customized by the sentiment analysis engine (such as text color and font size) is also reflected here.
[0861] User questions and responses
[0862] The user enters a specific question into the device and presses the send button. The device then sends this question to the server. The server receives the question, extracts relevant information from the database and content, and generates an answer. Again, the sentiment analysis engine analyzes the user's emotions and adjusts the content and expression of the answer. The generated answer is sent to the device and displayed to the user.
[0863] Dynamic generation of illustrations
[0864] When a user requests a visualization of a particular scene, they can press the "Image" button to send the request to the device. The device then sends this request to the server. The server uses a generative AI model (e.g., DALL-E or Stable Diffusion) to generate a visual image based on the prompt. This visual image is also customized to take into account the user's emotional state. The generated visual image is then sent to the device and displayed to the user.
[0865] Specific examples
[0866] For example, suppose a user is reading a particular novel and has read up to page 56. The user enters "I've read up to page 56" into the device and presses the send button. At this point, the emotion analysis engine recognizes that the user is excited. The server receives and analyzes this information, generating a summary, foreshadowing, and character information accordingly. The emotion analysis engine customizes the information to match the user's state of excitement and sends it to the device. If the user also enters a question such as "Was there any information about the perpetrator's motive?", the server will generate an answer to the question that takes the emotion analysis engine into consideration. Furthermore, if the user wishes to visualize a particular scene, a visual image can be generated by pressing the "Visualize" button. The emotion analysis engine customizes the visual image according to the user's emotion and displays it on the device.
[0867] Prompt Sentence Examples
[0868] "I'm 56 pages in. Can you give me a summary of the novel and what the characters are? If the sentiment analysis engine thinks I'm excited, please make sure the information is presented in a way that's enjoyable for the user."
[0869] The above is an embodiment of the present invention.
[0870] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0871] Step 1:
[0872] Enter the page number that the user has read.
[0873] Specific operation: The user enters "page 56" into the terminal and presses the send button. The terminal receives this input data and recognizes the user's emotional state (e.g., excitement) through an emotion analysis engine. The obtained data (page number and emotional state) is sent to the server.
[0874] Input: page number (56) and emotional state (excitement).
[0875] Data processing: Add emotional data using a sentiment analysis engine.
[0876] Output: Data including page number and emotional state.
[0877] Step 2:
[0878] The server analyzes the received page number and emotional state.
[0879] Specific operation: The server analyzes the received data and uses a natural language processing engine such as GPT-4 to generate a summary, foreshadowing, and character information up to the relevant page. In addition, an emotion analysis engine analyzes the user's emotional state and adjusts the method and content of information display.
[0880] Input: Data including page number and emotional state.
[0881] Data processing: Use a natural language processing engine to generate summaries, plot twists, and character information, and customize the information with a sentiment analysis engine.
[0882] Output: Customized summary, foreshadowing, and character information.
[0883] Step 3:
[0884] The terminal displays the information received from the server.
[0885] Specific operation: The device displays the information received from the server in a user-friendly interface (e.g., a web page or dedicated app). Information customized by the emotion analysis engine is also reflected here. For example, the text color may be brightened and the font size increased depending on the state of excitement.
[0886] Input: Customized summary, foreshadowing, and character information.
[0887] Data Calculation: Convert customized information into the appropriate display format.
[0888] Output: Display of adjusted information.
[0889] Step 4:
[0890] The user types in a question and sends it to the server.
[0891] Specific operation: The user enters a specific question into the terminal (e.g., "Is there any information about the perpetrator's motive?") and presses the send button. The terminal then sends this question to the server.
[0892] Input: A question entered by the user (e.g., "Was there any information on the perpetrator's motive?").
[0893] Data processing: The entered question is sent to the server.
[0894] Output: The query data sent to the server.
[0895] Step 5:
[0896] The server receives the query and generates the relevant information.
[0897] Specific operation: The server analyzes the question and extracts relevant information from the database and content. At this time, the emotion analysis engine analyzes the user's emotional state and adjusts the content and expression of the answer. The generated answer is then sent to the device.
[0898] Input: The query data sent to the server.
[0899] Data processing: Uses database search and sentiment analysis engines to generate relevant answers.
[0900] Output: The customized answer.
[0901] Step 6:
[0902] The device displays the received response to the user.
[0903] Specific operation: The device displays the answers received from the server in a user-friendly format. Customized answers are also applied using the sentiment analysis engine.
[0904] Input: Customized Answer.
[0905] Data Calculation: Convert customized answers into the appropriate display format.
[0906] Output: A display of the appropriately adjusted answer.
[0907] Step 7:
[0908] The user requests the generation of a visual image.
[0909] Specific operation: The user presses the "Image" button on the device to request the device to visualize a specific scene. The device then sends this request to the server.
[0910] Input: Visual image generation requested by the user.
[0911] Data processing: Send the request to the server.
[0912] Output: The visual image generation request sent to the server.
[0913] Step 8:
[0914] The server generates the visual image and sends it to the terminal.
[0915] How it works: The server uses a generative AI model (e.g., DALL-E or Stable Diffusion) to generate a visual image based on the prompt. This image is also customized to take into account the user's emotional state. The generated image is then sent to the device and displayed to the user.
[0916] Input: A visual image generation request.
[0917] Data processing: Visual images are generated using generative AI models and customized with a sentiment analysis engine.
[0918] Output: A customized visual image.
[0919] (Application example 2)
[0920] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0921] Conventional reading support systems have difficulty customizing the user's reading experience, and are particularly limited in providing information tailored to the user's emotions. As a result, users are unable to receive the information they desire quickly and appropriately, resulting in a poor quality reading experience. Furthermore, the generation of illustrations, which are visual representations, is fixed and not dynamically generated to match the user's emotions. There was a need for a system that could solve these issues and improve the user's reading experience.
[0922] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0923] In this invention, the server includes a means for analyzing the content based on the page number and generating information on the plot, foreshadowing, and characters, a means for generating and transmitting illustrations using a generative AI model, and a means for analyzing the user's emotions using an emotion engine and adjusting the information display method, thereby enabling information display and dynamic illustration generation according to the user's emotions.
[0924] "User" refers to an individual user of the system.
[0925] "Page number read" refers to the specific page number of the book that the user has currently read.
[0926] "Input means" refers to the interfaces and devices through which a user provides information to the system.
[0927] "Server" refers to a computer system that receives requests from users, analyzes them, and provides appropriate information.
[0928] "Means for transmitting to the server" refers to a communication means for transmitting the user's input information to the server.
[0929] "Means for analyzing the content and generating plot, plot twists, and character information" refers to the process by which the server analyzes the content of the book based on page numbers and generates summary information to provide to the user.
[0930] "Displaying means" refers to devices and interfaces for visually presenting information received from the server to a user.
[0931] "Means for inputting a question and sending it to the server" refers to an interface and communication means for a user to input a question to the system and send the question to the server.
[0932] "Means for generating and sending answers based on questions" refers to the process by which the server analyzes the user's question, derives an appropriate answer, and provides it to the user.
[0933] "Means for requesting the generation of illustrations" refers to an interface through which a user requests the system to create a visual image of a particular scene.
[0934] "Means for generating and transmitting illustrations using a generative AI model" refers to the process by which a server uses AI technology to dynamically generate illustrations and provide those illustrations to users.
[0935] An "emotion engine" is an engine that analyzes a user's emotions and adjusts the way information is displayed based on those emotions.
[0936] "Means of analyzing emotions and adjusting how information is displayed" refers to the process by which the emotion engine understands the user's current emotional state and changes the content and format of the display accordingly.
[0937] System Overview
[0938] This system includes four main components: the user, the device, the server, and the emotion engine. The user inputs the page number they have read and a question into the device, which then sends it to the server. The server analyzes the received information, generates appropriate data, and sends it to the device. During this process, the emotion engine analyzes the user's emotions and adjusts the information display and response content accordingly.
[0939] Enter the page number the user has finished reading
[0940] The user uses the terminal to input the page number they have just finished reading. This information is sent to the server via the terminal. The server also recognizes the emotion (e.g., excitement, sadness, doubt, etc.) with which the user seeks information.
[0941] Server analysis and information generation
[0942] The server analyzes the received page number and generates a list of the plot, foreshadowing, and characters up to that page. The information generated by the server is organized in the form of plot, foreshadowing, and character information and sent to the device. During this process, the emotion engine analyzes the user's emotions and adjusts the display method and content of the information.
[0943] Terminal display
[0944] The device displays the information received from the server in a user-friendly interface. The emotion engine customizes the interface based on the user's emotions. For example, if the user is excited, the display may change to a larger font size or change the color of the text. If the user is sad, the display may use warmer colors.
[0945] User questions and responses
[0946] When a user enters a question, the device sends it to the server. Based on the question, the server extracts relevant information from existing databases and content to generate an answer. During this process, the emotion engine analyzes the user's emotions and adjusts the content and expression of the answer. For example, if the user is sad, it can incorporate words of encouragement. This answer is sent to the device and displayed to the user.
[0947] Dynamic generation of illustrations
[0948] If a user wants a visual image of a particular scene, they can request the generation of an illustration by pressing the "Image" button. The device sends this request to the server, which then uses a generative AI model to generate an illustration of the scene. The generated illustration, which takes the user's emotions into account using an emotion engine, is sent to the device and displayed to the user.
[0949] Specific examples
[0950] For example, suppose a user is reading a particular novel. The user enters into the system that they have read up to page 56, and the emotion engine recognizes that the user is currently excited. The server receives and analyzes this information, and generates information on the plot, foreshadowing, and characters accordingly. The emotion engine customizes this information to match the user's state of excitement and sends it to the device. If the user asks, "Was there any information on the perpetrator's motive?", the server generates an answer to the question that takes the emotion engine into consideration. Furthermore, if the user wants to see an illustration of a particular scene, they can press the "Visualize" button, and an illustration will be generated. The emotion engine customizes the illustration according to the user's emotion and displays it on the device.
[0951] The specific hardware and software used
[0952] Hardware:
[0953] Smartphone (iOS or Android)
[0954] Server (equipped with a high-performance GPU)
[0955] Emotion engine (commonly known as "emotion analysis engine")
[0956] software:
[0957] Frontend: React Native
[0958] Backend: Node.js, Express
[0959] Database: MongoDB
[0960] Emotion analysis: Python, TensorFlow
[0961] Prompt Sentence Examples
[0962] "Page 56, Emotion: Excite, Question: Was there any information about the perpetrator's motive?"
[0963] The above is a specific embodiment for implementing this system.
[0964] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0965] Program processing steps
[0966] Step 1:
[0967] The user enters the page number they have read into the terminal.
[0968] Input: page number the user finished reading, emotional state (e.g., excited, sad)
[0969] Specific operation: The user enters the page number of the page they have just finished reading in the page number input interface of the application, selects their emotional state, and the result is input to the terminal.
[0970] Step 2:
[0971] The terminal transmits the input page number and emotional state to the server.
[0972] Input: User-entered page number and emotional state
[0973] Output: Data packet sent to the server
[0974] Specific operation: The device sends the input page number and emotional state as a data packet to the server using an HTTP POST request.
[0975] Step 3:
[0976] The server analyzes the content based on the page number and generates plot summary, foreshadowing, and character information.
[0977] Input: The page number received by the server
[0978] Output: Synopsis, foreshadowing, character information data
[0979] Specific operation: The server retrieves information up to the relevant page from the database, and based on this generates a plot summary, foreshadowing, and a list of characters. A Python script is executed.
[0980] Step 4:
[0981] The server uses an emotion engine to analyze the user's emotions and adjust how information is displayed.
[0982] Input: User's emotional state, generated information data
[0983] Output: Information data adjusted based on user sentiment
[0984] How it works: The emotion engine (a model using TensorFlow) analyzes the emotional state and adjusts the format and display style of the generated information data. For example, if the user is excited, the font size and color will be changed.
[0985] Step 5:
[0986] The server transmits the adjusted information data to the terminal.
[0987] Input: Adjusted information data
[0988] Output: Data packets sent to the device
[0989] Specific operation: The server assembles the adjusted information data into a data packet and sends it to the terminal. The communication method is an HTTP POST request.
[0990] Step 6:
[0991] The information received by the terminal is displayed in a user-friendly interface.
[0992] Input: Adjusted information data received from the server
[0993] Output: On-screen display
[0994] Specific operation: Based on the received information data, the device displays information through a customized interface according to the user's emotional state. The display process is performed using React Native.
[0995] Step 7:
[0996] The user enters a question and the device sends the question to the server.
[0997] Input: The question entered by the user
[0998] Output: Data packet sent to the server
[0999] Specific operation: The user inputs a question through the question input interface, and the device sends the question to the server via an HTTP POST request.
[1000] Step 8:
[1001] The server extracts relevant information from existing databases based on the question and generates an answer.
[1002] Input: User question
[1003] Output: Response data
[1004] What happens: The server searches a database for information related to the question and generates an answer. A Python script is executed to extract the requested information.
[1005] Step 9:
[1006] The emotion engine analyzes the user's emotions and adjusts the content and expression of the response.
[1007] Input: User's emotional state, generated answer data
[1008] Output: Answer data adjusted based on user sentiment
[1009] What it does: The emotion engine analyzes the user's emotional state and adjusts the wording of the generated response data accordingly, possibly adding words of encouragement.
[1010] Step 10:
[1011] The server transmits the adjusted response data to the terminal.
[1012] Input: Adjusted response data
[1013] Output: Data packets sent to the device
[1014] Specific operation: The server assembles the adjusted response data into a data packet and sends it to the terminal using an HTTP POST request.
[1015] Step 11:
[1016] The user requests the creation of an illustration, and the device sends the request to the server.
[1017] Input: Illustration generation request
[1018] Output: Data packet sent to the server
[1019] Specific operation: The user presses the "Image" button to request the creation of an illustration, and the device sends the request to the server via an HTTP POST request.
[1020] Step 12:
[1021] The server uses a generative AI model to generate and transmit illustrations.
[1022] Input: Illustration generation request
[1023] Output: Generated illustration data
[1024] Specific operation: The server uses a generative AI model (e.g., DALL-E) to generate illustrations according to the request. The generated illustration data is prepared.
[1025] Step 13:
[1026] The emotion engine customizes the generated illustrations according to the user's emotions.
[1027] Input: Generated illustration data, user's emotional state
[1028] Output: Adjusted illustration data
[1029] How it works: The emotion engine analyzes the user's emotional state and customizes the generated illustrations accordingly. For example, if the user is excited, the colors and details will be enhanced.
[1030] Step 14:
[1031] The server sends the adjusted illustration data to the terminal, which displays it to the user.
[1032] Input: Adjusted illustration data
[1033] Output: Illustration displayed on the terminal
[1034] Specific operation: The server assembles the adjusted illustration data into a data packet and sends it to the device. The device displays the received illustration data in a user-friendly interface. The display process is performed using React Native.
[1035] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1036] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1037] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1038] [Third embodiment]
[1039] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1040] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1041] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1042] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1043] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1044] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1045] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1046] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1047] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1048] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1049] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1050] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1051] The present invention is a system designed to enhance a user's reading experience, and is embodied in the following manner.
[1052] System Overview
[1053] This system has three main components: the user, the terminal, and the server. The user inputs the page number and question they have read into the terminal, which then sends it to the server. The server analyzes the received information, generates appropriate data, and sends it to the terminal to provide information to the user.
[1054] Enter the page number the user has finished reading
[1055] The user uses the terminal to input the page number they have just finished reading. This information is sent to the server via the terminal. At this stage, the user is ready to receive information from the system based on their progress in the story.
[1056] Server analysis and information generation
[1057] The server analyzes the received page number and generates a list of the plot, hints, and characters up to that page. This allows the user to review the main points of the story or notice hints. The information generated by the server is compiled in the form of plot, hints, and character information and sent to the terminal.
[1058] Terminal display
[1059] The device displays the information received from the server in a user-friendly interface, allowing users to easily understand the main points of the story and the relationships between characters. If users require specific information or detailed explanations, they can type their questions into the device.
[1060] User questions and responses
[1061] When a user inputs a question, the device sends the question to the server, which extracts relevant information from existing databases and content based on the question and generates an answer. This answer is then immediately sent to the device, which displays it to the user, allowing the user to deepen their understanding of the story.
[1062] Dynamic generation of illustrations
[1063] If a user wants a visual image of a particular scene, they can request the generation of an illustration by pressing the "Image" button. The device sends this request to the server, which uses an AI model to generate an illustration for the scene. The generated illustration is sent to the device and displayed to the user. This feature allows users to visually enjoy the story scenes.
[1064] Specific examples
[1065] For example, suppose a user is reading a particular novel. When the user inputs into the system that he or she has read up to page 56, the server receives and analyzes that information, generates information about the plot, foreshadowing, and characters up to that point, and sends it to the device. If the user asks, "Was there any information about the perpetrator's motive?", the server will provide an answer to that question. Furthermore, if the user wants to see an illustration of a particular scene, the illustration will be generated by pressing the "Image" button.
[1066] The above is an embodiment of the present invention, which provides a specific method for improving the user's reading experience.
[1067] The processing flow will be explained below.
[1068] Step 1:
[1069] The user inputs the page number of the page that has been read. The user inputs the page number of the page that has been read in the terminal interface. For example, the user inputs "56".
[1070] Step 2:
[1071] The terminal transmits the input page number to the server. The terminal generates an HTTP request for transmitting the input page number to the server and transmits it to the server.
[1072] Step 3:
[1073] The server analyzes the received page number and retrieves the data up to that page from the database or content management system based on the received page number.
[1074] Step 4:
[1075] The server analyzes the content and generates plot, foreshadowing, and character information. The server uses a data analysis algorithm to identify plot, foreshadowing, and characters based on page numbers and compiles each piece of information.
[1076] Step 5:
[1077] The server sends the generated information to the terminal. The server then sends the generated information about the plot, hints, and characters to the terminal as an HTTP response.
[1078] Step 6:
[1079] The terminal receives and displays information from the server. The terminal receives a response from the server and displays the plot, plot twists, and character information on the user interface.
[1080] Step 7:
[1081] The user enters a question. The user enters a question into the question input field on the device. For example, the user enters a question such as "What is the perpetrator's motive?"
[1082] Step 8:
[1083] The device sends the question to the server. The device sends the user's question to the server as an HTTP request.
[1084] Step 9:
[1085] The server generates an answer based on the question. The server searches for information related to the question from a database or existing content and generates an answer.
[1086] Step 10:
[1087] The server sends the generated answer to the terminal. The server sends the generated answer to the terminal as an HTTP response.
[1088] Step 11:
[1089] The terminal displays the answer from the server. The terminal receives the response from the server and displays the answer to the user on the user interface.
[1090] Step 12:
[1091] The user requests the creation of an illustration. If the user wants to see an illustration of a particular scene, they press the "Image" button.
[1092] Step 13:
[1093] The terminal sends a request to generate an illustration to the server. The terminal detects the button press event and sends a request to generate an illustration to the server.
[1094] Step 14:
[1095] The server uses the AI model to generate illustrations. The server uses the AI model for generating illustrations to generate illustrations based on the specified page content.
[1096] Step 15:
[1097] The server sends the generated illustration to the terminal. The server generates and sends an HTTP response to send the generated illustration to the terminal.
[1098] Step 16:
[1099] The terminal receives and displays the illustration. The terminal receives the illustration data from the server and displays it to the user on the user interface.
[1100] Example 1
[1101] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1102] Previous systems lacked sufficient means for users to check story details and important hints while reading, or to rely on character information to understand the story. Furthermore, obtaining visual images required separate searches, which was time-consuming. Furthermore, few systems could provide immediate answers to user questions or dynamically generate information based on the page being read. Therefore, an effective system to enhance users' reading experience was needed.
[1103] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1104] In this invention, the server includes means for inputting the page number of a page that a user has finished reading, means for transmitting the input page number to the server, means for the server to analyze the content based on the page number and generate information on the plot, foreshadowing, and characters, means for receiving and displaying information from the server, means for inputting a user's question and transmitting it to the server, means for the server to generate and transmit an answer based on the question, means for the user to request the generation of an illustration, means for the server to generate and transmit an illustration using an AI model, and means for the terminal to display information in a user-friendly interface. This enables the user to check the plot, foreshadowing, and information on characters in real time while reading, get instant answers to questions, and generate visual images.
[1105] "User" means a person who uses the system and takes action to enhance their reading experience.
[1106] A "terminal" is a device that allows a user to perform input operations and receive and display information from a server.
[1107] A "server" is a central computing device that analyzes information sent from users and terminals and generates and provides the necessary data.
[1108] The "page number read" is the number of the page that the user currently recognizes as having finished reading.
[1109] "Input means" refers to the interface or device that allows the user to input the page number they have read or questions into the terminal.
[1110] "Transmission means" refers to the function or device for transmitting information entered by the user from the terminal to the server.
[1111] The "analysis means" refers to a function that allows the server to process data based on the input page number and analyze the content.
[1112] The "synopsis generation means" refers to a function or method for summarizing the main points of the story up to the page that the server has finished reading.
[1113] The "foreshadowing generation means" refers to a function or method by which the server creates a list of important foreshadowings that appear in the story.
[1114] "Character information generation means" refers to the function or method by which the server analyzes and lists the characters in the story and their relationships.
[1115] "Display means" refers to the functions and methods by which the terminal presents the information received from the server in an easy-to-read format for the user.
[1116] "Question input means" refers to an interface or device that allows a user to input detailed information or specific questions into a terminal.
[1117] "Answer generation means" refers to the function or method by which the server generates related information based on the user's question and creates an answer.
[1118] The "illustration generation request means" refers to a function or method that allows a user to send a request from a terminal to a server in order to generate a visual image of a specific scene.
[1119] "AI model" refers to the artificial intelligence algorithms and functions that the server uses to generate illustrations and other data.
[1120] A "user-friendly interface" refers to the screen or method by which a device presents information to the user in an intuitive and easy-to-understand format.
[1121] This invention is a system designed to improve the user's reading experience, and it includes three main components: the terminal, the server, and the user. The following describes each flow in detail.
[1122] First, the user inputs the page number they have finished reading and a question into the terminal. The terminal is a mobile device such as a tablet or smartphone, which receives and processes the input information. The terminal is equipped with a user interface and is designed for intuitive operation. Specifically, the user might input "I have finished reading up to page 56" or a question such as "What is the perpetrator's motive?"
[1123] The device then sends the input information to the server, where it transmits the data via HTTP over Wi-Fi or mobile networks, and organizes the data in JSON format for efficient delivery to the server.
[1124] The server is a high-performance data server, such as AWS EC2, a widely used cloud service. The server searches the database based on the page number sent by the user and analyzes the content up to that page. This analysis uses an algorithm that utilizes natural language processing technology to dynamically generate the story's plot, hints, and character information. This allows the user to see the progress of the story and understand its key points.
[1125] When a user enters a question, the server generates an answer based on that question. Here, machine learning technology is used to extract relevant information from existing databases and derive appropriate answers. For example, in response to the question "What was the perpetrator's motive?", the server searches the database for relevant information, extracts the relevant parts, and generates an answer.
[1126] Furthermore, if a user wants a visual image of a specific scene, they can request the generation of an illustration by pressing the "Image" button on their device. This request is sent to the server, which then uses a generative AI model (e.g., DALL-E 2) to generate an illustration of the scene. The AI model generates an image of a specific scene based on a prompt. For example, a prompt might be, "Please generate an illustration of the scene where the protagonist confronts the criminal for the first time."
[1127] The information and illustrations generated by the server are then sent back to the device, which receives them and displays them in a user-friendly interface, allowing users to enjoy the story visually as well.
[1128] (Example of a prompt)
[1129] "I've finished reading page 56 so far. Please tell me the plot, foreshadowing, and characters."
[1130] "Did you have any information about the perpetrator's motive?"
[1131] "Generate an illustration of the scene where the protagonist confronts the culprit for the first time."
[1132] As described above, this system performs appropriate data processing and calculations based on the information entered by the user, providing the user with a fulfilling reading experience.
[1133] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1134] Step 1:
[1135] The user inputs the page number they have read and a question into the device. The user inputs the page number into the device's input field in the format of "I have now finished reading page 56." They also input questions such as "What was the perpetrator's motive?" The input is in text format, and the device receives it and stores it in its internal memory.
[1136] Step 2:
[1137] The device sends the input information to the server. The device converts the page number and question information entered by the user into JSON format and sends it to the server via an HTTP request. Specifically, the data is sent using Wi-Fi or a mobile network.
[1138] Input: A page number or question entered by the user
[1139] Output: Data converted to JSON format
[1140] Specific operation: The device converts the text data into JSON format and sends it to the server using an HTTP request.
[1141] Step 3:
[1142] The server analyzes the page number and generates a synopsis, foreshadowing, and character information. The server then analyzes the received JSON data and retrieves the content up to the relevant page from the database. Based on the retrieved data, the server uses natural language processing technology to analyze the content and generate information to provide to the user.
[1143] Input: JSON data of page numbers received from the terminal
[1144] Output: Generated plot, hints, and character information
[1145] Specific operation: The server searches the database, analyzes the retrieved data, and generates the required information.
[1146] Step 4:
[1147] The server sends the generated information to the terminal. The generated plot, foreshadowing, and character information are compiled in JSON format and sent to the terminal as an HTTP response.
[1148] Input: Generated plot, hints, character information
[1149] Output: Data in JSON format
[1150] Specific operation: The information generated by the server is converted into JSON format and sent to the terminal via an HTTP response.
[1151] Step 5:
[1152] The device displays the received information in a user-friendly manner. The device analyzes the received JSON data and presents the plot, plot twists, and character information to the user in a visually easy-to-understand format.
[1153] Input: JSON data received from the server
[1154] Output: Information displayed in a user-friendly format
[1155] Specific operation: The device parses the JSON data and displays it on the screen.
[1156] Step 6:
[1157] The user enters a follow-up question into the device, which then sends it to the server. The user enters a follow-up question into the device, such as "What is your specific motivation?" The device receives the question and sends it back to the server.
[1158] Input: Any additional questions entered by the user
[1159] Output: The question sent to the server
[1160] Specific operation: The device receives a new question, converts it into JSON format, and sends it to the server.
[1161] Step 7:
[1162] The server analyzes the question, generates relevant information, and sends it to the device. The server analyzes the question, extracts relevant information from an existing database, and generates an answer. The generated answer is compiled in JSON format and sent to the device.
[1163] Input: JSON data of the question received from the terminal
[1164] Output: JSON data of the generated answer
[1165] What it does: The server searches the database, extracts relevant information, and generates an answer.
[1166] Step 8:
[1167] The terminal displays the answer and, if necessary, sends a request to the server to generate an illustration. The terminal displays the answer to the user, and if necessary, the user presses the "Image" button to request the generation of an illustration.
[1168] Input: JSON data of the response received from the server
[1169] Output: Answers displayed to the user and illustration generation request
[1170] Specific operation: The device displays the answer on the screen, and when the user presses the "Image" button, it sends a request to the server to generate an illustration.
[1171] Step 9:
[1172] The server generates illustrations and sends them to the device. The server receives illustration generation requests and generates illustrations for the corresponding scenes using a generative AI model. The generated illustrations are then sent to the device.
[1173] Input: Illustration generation request received from the device
[1174] Output: Generated illustration data
[1175] Specific operation: The server sends a prompt to the AI model and sends the generated illustration data to the terminal.
[1176] Step 10:
[1177] The terminal displays the generated illustration. The terminal displays the illustration data received from the server to the user, allowing the user to visually enjoy a specific scene in the story.
[1178] Input: illustration data received from the server
[1179] Output: The illustration displayed to the user
[1180] Specific operation: The device converts the illustration data for display on the screen and shows it to the user.
[1181] These are the specific processing steps of the present invention, which allow users to check plot summaries, hints, and character information in real time while reading, get instant answers to questions, and generate visual images.
[1182] (Application example 1)
[1183] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1184] Conventional reading systems make it difficult for users to obtain summaries, plot twists, and character information according to their reading progress. Furthermore, they lack the ability to generate illustrations to visually enjoy specific scenes, limiting the reading experience. Furthermore, when users have questions about the story, there is a lack of a way to get immediate answers. It is necessary to solve these problems and improve the user's reading experience.
[1185] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1186] In this invention, the server includes means for analyzing the content based on the page number that the user has read and generating information on the plot, foreshadowing, and characters, means for generating answers based on the user's questions, and means for generating illustrations using a generative AI model. This allows the user to easily obtain a summary of the content, foreshadowing, and character information according to their reading progress, as well as enjoy visual images of specific scenes and receive immediate answers to questions that will help them better understand the story.
[1187] "User" refers to an individual who utilizes the system to enhance their reading experience.
[1188] "Page number read" refers to the number of the specific page of the book that the user has currently read.
[1189] "Server" refers to the central processing unit that analyzes information sent by users and generates and transmits necessary data.
[1190] A "plot" is a concise summary of the main developments and events of a story.
[1191] A "foreshadowing" is information that provides hints or elements that will become important later in the development of the story.
[1192] "Character information" refers to detailed information about the characters' names, backgrounds, and roles in the story.
[1193] "Question" refers to the act of a user inquiring of the server about specific matters regarding the content of the story.
[1194] "Answer" refers to information generated by the server based on the user's question.
[1195] An "illustration" is an image or painting that visually depicts a particular scene from a story.
[1196] A "generative AI model" refers to an algorithm or system that uses artificial intelligence technology to generate visual images or answers in response to user requests.
[1197] "Prompt sentence" refers to the text input used to prompt a generative AI model to generate an illustration.
[1198] "Display means" refers to an interface for displaying the generated information and illustrations in a form that can be visually recognized by the user.
[1199] This invention is a system designed to improve the user's reading experience, and it includes three main elements: a user, a terminal, and a server. The user uses the terminal to input the page number they have read and questions, and the terminal sends that information to the server. The server analyzes the received information, generates appropriate data, and sends it to the terminal to provide information to the user.
[1200] System Overview
[1201] The system includes the following elements:
[1202] 1. Enter page number:
[1203] The user inputs the page number they have read into the device, which allows the device to provide information based on the progress of the story.
[1204] 2. Information analysis and generation:
[1205] The server parses the page number it receives and generates a list of plot, plot twists, and characters up to that page, then compiles this information in a user-friendly format and sends it to the device.
[1206] 3. Questions and Answers:
[1207] When a user enters a question, the device sends it to the server, which extracts relevant information from an existing database and immediately generates an answer that is sent to the device.
[1208] 4. Dynamic generation of illustrations:
[1209] When a user wants a visual image of a particular scene, they press the "Image" button to request the generation of an illustration. The server uses a generative AI model to generate an illustration of the scene, which is then sent to the device for display.
[1210] Hardware and software used
[1211] Hardware:
[1212] User devices: smartphones, tablets
[1213] Server: Server equipment that performs high-performance calculations
[1214] software:
[1215] Server API: e.g. Flask or Django
[1216] AI model: Generative AI models such as Stable Diffusion and DALL-E are used to generate illustrations.
[1217] HTTP request library: Requests
[1218] Example of a system
[1219] 1. Enter and analyze page numbers:
[1220] For example, if a user inputs into the system that they have read a particular novel and have progressed to page 56, the server receives and analyzes that information, generating information about the plot, foreshadowing, and characters up to that point, and sending it to the terminal.
[1221] Example prompt: "Use case: I've read up to page 56. What's the plot and foreshadowing so far?"
[1222] 2. Questions and Answers:
[1223] If a user asks, "Is there any information about the perpetrator's motive?", the server searches for relevant information based on the question, generates an answer, and sends it to the device.
[1224] Example prompt: "Use case: I want to ask for information about the perpetrator's motive."
[1225] 3. Illustration generation:
[1226] If a user wants to see a visual image of a particular scene, they can press the "Visualize" button, which will generate an illustration. For example, a request to generate an illustration of a scene in which the protagonist is in conflict is sent to the server, and the illustration is generated and displayed using a generative AI model.
[1227] Example prompt: "Use case: Visualize a scene where the protagonist is in conflict."
[1228] The above is an embodiment of the present invention, which provides a specific method for improving the user's reading experience.
[1229] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1230] Step 1:
[1231] The user uses the terminal to enter the page number that they have read.
[1232] Input: page number read
[1233] Specific behavior: The user enters the page number into the input field on the device and presses the "Submit" button.
[1234] Step 2:
[1235] The terminal transmits the input page number to the server.
[1236] Input: The page number entered by the user
[1237] Output: Page number data sent to the server
[1238] Specific operation: The terminal sends the entered page number to the server via an HTTP request.
[1239] Step 3:
[1240] The server analyzes the page number received and generates information on the plot, plot twists, and characters up to that page.
[1241] Input: Submitted page number data
[1242] Output: Synopsis, foreshadowing, character information
[1243] Specific operation: The server searches the database, analyzes the content up to the relevant page, and generates information on the plot, foreshadowing, and characters.
[1244] Step 4:
[1245] The server transmits the generated information to the terminal.
[1246] Input: Synopsis, foreshadowing, character information
[1247] Output: Information data sent to the terminal
[1248] Specific operation: The server sends the generated information to the terminal in a data format such as JSON.
[1249] Step 5:
[1250] The terminal displays the information received from the server in a user-friendly interface.
[1251] Input: Information data sent from the server
[1252] Output: Plot, plot twists, and character information displayed on the screen
[1253] Specific operation: The device analyzes the received information and generates and displays HTML and CSS to display on the interface.
[1254] Step 6:
[1255] The user enters a specific question, and the device sends the question to the server.
[1256] Input: The question entered by the user
[1257] Output: Question data sent to the server
[1258] Specific operation: The user enters a question in the question input field and presses the "Send" button. The device sends the question to the server via an HTTP request.
[1259] Step 7:
[1260] The server extracts relevant information from existing databases based on the question and generates an answer.
[1261] Input: Submitted question data
[1262] Output: relevant information and generated answers
[1263] What it does: The server analyzes the question, extracts relevant information from the database, and generates an answer.
[1264] Step 8:
[1265] The server generates a response and sends it to the terminal.
[1266] Input: Generated answer
[1267] Output: Answer data sent to the device
[1268] Specific operation: The server converts the generated response into a data format and sends it to the terminal.
[1269] Step 9:
[1270] The device displays the response received from the server to the user.
[1271] Input: Response data sent from the server
[1272] Output: Answer displayed on the screen
[1273] Specific operation: The device analyzes the received response and generates and displays an interface to display to the user.
[1274] Step 10:
[1275] A user requests a visual image of a particular scene, and the device sends the request to the server.
[1276] Input: The prompt for the scene the user wants
[1277] Output: Request data sent to the server
[1278] Specific operation: The user inputs a prompt sentence about a specific scene and presses the "Image" button. The device sends a request to the server.
[1279] Step 11:
[1280] The server uses a generative AI model to generate illustrations of the requested scene.
[1281] Input: Request data (prompt statement)
[1282] Output: Generated illustration data
[1283] Specific operation: The server inputs a prompt sentence into the generative AI model and receives the generated illustration.
[1284] Step 12:
[1285] The server sends the generated illustration to the terminal.
[1286] Input: Generated illustration data
[1287] Output: Illustration data sent to the device
[1288] Specific operation: The server converts the illustration data into an appropriate format and sends it to the terminal.
[1289] Step 13:
[1290] The terminal displays the illustrations received from the server in a user-friendly interface.
[1291] Input: Illustration data sent from the server
[1292] Output: Illustration displayed on the screen
[1293] Specific operation: The terminal analyzes the received illustration data and generates and displays an interface for visual display.
[1294] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1295] The present invention is an advanced system for improving a user's reading experience, incorporating an emotion engine for recognizing a user's emotions and providing corresponding information. Specific embodiments for implementing the present invention will be described below.
[1296] System Overview
[1297] This system includes four main components: the user, the device, the server, and the emotion engine. The user inputs the page number they have read and a question into the device, which then sends it to the server. The server analyzes the received information, generates appropriate data, and sends it to the device. During this process, the emotion engine analyzes the user's emotions and adjusts the information display and response content.
[1298] Enter the page number the user has finished reading
[1299] The user uses the terminal to input the page number they have just finished reading. This information is sent to the server via the terminal. The server also recognizes the emotion (e.g., excitement, sadness, doubt, etc.) with which the user seeks information.
[1300] Server analysis and information generation
[1301] The server analyzes the received page number and generates a list of the plot, foreshadowing, and characters up to that page. The information generated by the server is organized in the form of plot, foreshadowing, and character information and sent to the device. During this process, the emotion engine analyzes the user's emotions and adjusts the display method and content of the information.
[1302] Terminal display
[1303] The device receives information from the server and displays it in a user-friendly interface. The emotion engine customizes the interface based on the user's emotions. For example, if the user is excited, the display may change to a different color or larger font size.
[1304] User questions and responses
[1305] When a user enters a question, the device sends it to the server. Based on the question, the server extracts relevant information from existing databases and content to generate an answer. During this process, the emotion engine analyzes the user's emotions and adjusts the content and expression of the answer. For example, if the user is sad, it can incorporate words of encouragement. This answer is sent to the device and displayed to the user.
[1306] Dynamic generation of illustrations
[1307] If a user wants a visual image of a particular scene, they can request the generation of an illustration by pressing the "Image" button. The device sends this request to the server, which uses an AI model to generate an illustration of the scene. The generated illustration takes the user's emotions into account using an emotion engine, and is sent to the device and displayed to the user.
[1308] Specific examples
[1309] For example, suppose a user is reading a particular novel. The user enters into the system that they have read up to page 56, and the emotion engine recognizes that the user is currently excited. The server receives and analyzes this information, and generates information on the plot, foreshadowing, and characters accordingly. The emotion engine customizes this information to match the user's state of excitement and sends it to the device. If the user asks, "Was there any information on the perpetrator's motive?", the server generates an answer to the question that takes the emotion engine into consideration. Furthermore, if the user wants to see an illustration of a particular scene, they can press the "Visualize" button, and an illustration will be generated. The emotion engine customizes the illustration according to the user's emotion and displays it on the device.
[1310] The above is an embodiment of the present invention, which provides a specific method for taking user emotions into consideration to improve the reading experience.
[1311] The processing flow will be explained below.
[1312] Step 1:
[1313] The user inputs the page number of the page that has been read. The user inputs the page number of the page that has been read in the terminal interface. For example, the user inputs "56".
[1314] Step 2:
[1315] The device sends the entered page number and the user's emotional state to the server. The device then generates an HTTP request to send the emotional data analyzed by the emotion engine and the page number to the server, and sends it to the server.
[1316] Step 3:
[1317] The server analyzes the received page number and emotional state. Based on the received page number, the server retrieves data up to that page from a database or content management system and analyzes the user's emotions.
[1318] Step 4:
[1319] The server generates information on the plot, plot twists, and characters, and then adjusts it based on the user's emotions using an emotion engine.The server uses a data analysis algorithm to identify the plot, plot twists, and characters based on page numbers, and the emotion engine customizes the information according to the user's emotions.
[1320] Step 5:
[1321] The server sends the generated information to the terminal. The server then sends the generated plot, foreshadowing, and character information to the terminal as emotion-adjusted information in an HTTP response.
[1322] Step 6:
[1323] The terminal receives and displays information from the server. The terminal receives a response from the server and displays customized plot, plot twists, and character information on the user interface.
[1324] Step 7:
[1325] The user enters a question. The user enters a question into the question input field on the device. For example, the user enters a question such as "What is the perpetrator's motive?"
[1326] Step 8:
[1327] The device sends the question and emotional state to the server. The device generates an HTTP request to send the question including the user's emotional data to the server and sends it to the server.
[1328] Step 9:
[1329] The server generates an answer based on the question and adjusts it based on the user's emotions. The server searches for information related to the question from a database or existing content and generates an answer. The generated answer is adjusted according to the user's emotions by an emotion engine.
[1330] Step 10:
[1331] The server sends the generated answer to the terminal, and the server sends the adjusted answer to the terminal as an HTTP response.
[1332] Step 11:
[1333] The device displays the answer from the server. The device receives the response from the server and displays the answer adjusted to match the user's emotions on the user interface.
[1334] Step 12:
[1335] The user requests the creation of an illustration. If the user wants to see an illustration of a particular scene, they press the "Image" button.
[1336] Step 13:
[1337] The device sends a request to generate an illustration and the user's emotional state to the server. The device detects the button press event and sends a request to generate an illustration and the user's emotional state to the server.
[1338] Step 14:
[1339] The server uses an AI model to generate illustrations and adjust them based on the user's emotions.The server uses an AI model for generating illustrations to generate illustrations based on the specified page content, and the emotion engine adjusts them according to the user's emotions.
[1340] Step 15:
[1341] The server sends the generated illustration to the device. The server then generates and sends an HTTP response to send the emotion-adjusted illustration to the device.
[1342] Step 16:
[1343] The terminal receives and displays the illustration. The terminal receives the illustration data from the server and displays the illustration adjusted based on the emotion on the user interface.
[1344] Example 2
[1345] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1346] Previous systems for improving the reading experience did not provide information that took the user's emotions into consideration, and they had the problem of being unable to display or respond appropriately to the user's emotional state. Furthermore, the generation of visual images was not tailored to the user's needs, and only a uniform response was possible. This limited the user's reading experience, making it difficult to provide a customized experience tailored to each individual's emotional state.
[1347] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for generating information on a summary, foreshadowing, and characters based on page numbers, a means for adjusting the information content and display method based on emotion analysis, and a means for generating visual images using a generative AI model. This makes it possible to provide and display information according to the user's emotional state, and to provide a reading experience customized for each user.
[1348] A "user" is someone who uses the system to enhance their reading experience.
[1349] A "terminal" is a device that allows a user to input the page number of a page that has been read, input a question, or request the generation of a visual image.
[1350] The "server" is a central processing unit that analyzes the information sent by the user, generates answers and visual images, and sends them to the terminal.
[1351] A "page number" is a number that indicates the specific page of a book that the user has finished reading.
[1352] "Emotion analysis means" refers to technology for recognizing and analyzing a user's emotional state.
[1353] A "generative AI model" is an algorithm or software that uses AI technology to generate visual images based on specific prompts.
[1354] "Visual images" are illustrations or drawings of specific scenes requested by users.
[1355] A "summary" is a short summary of the contents up to a specific page number.
[1356] A "foreshadowing" is an element in a story that contains information or hints that will be important for later developments.
[1357] "Characters" is a list and profiles of characters that appear in the book.
[1358] A "question" is a question about specific information that a user inputs into the system.
[1359] "Means for adjusting information content and display method" refers to technology that changes the display format and content of information provided based on the user's emotional state.
[1360] This invention is a system for improving a user's reading experience, and in particular, has the function of customizing information according to the user's emotional state. This system mainly consists of a user (terminal), a terminal, a server, and an emotion analysis engine.
[1361] Enter the page number the user has read
[1362] The user inputs the page number they have finished reading into the device. The device receives this information and also analyzes the user's emotional state using an emotion analysis engine. For example, the device uses emotion recognition technology such as Microsoft Azure's Emotion API. The analyzed data (page number and emotional state) is sent to the server.
[1363] Server analysis and information generation
[1364] The server analyzes the received page number and emotion data. Based on the page number, a natural language processing engine (such as OpenAI's GPT-4) is used to generate a summary, plot twists, and character information. The emotion analysis engine also analyzes the user's emotional state and adjusts the display and content of the information. For example, if the user is excited, the text color or font size is changed. This adjusted information is then sent to the device.
[1365] Terminal display
[1366] The device displays the information received from the server in a user-friendly interface, which can be provided as a web page or a dedicated application. Information customized by the sentiment analysis engine (such as text color and font size) is also reflected here.
[1367] User questions and responses
[1368] The user enters a specific question into the device and presses the send button. The device then sends this question to the server. The server receives the question, extracts relevant information from the database and content, and generates an answer. Again, the sentiment analysis engine analyzes the user's emotions and adjusts the content and expression of the answer. The generated answer is sent to the device and displayed to the user.
[1369] Dynamic generation of illustrations
[1370] When a user requests a visualization of a particular scene, they can press the "Image" button to send the request to the device. The device then sends this request to the server. The server uses a generative AI model (e.g., DALL-E or Stable Diffusion) to generate a visual image based on the prompt. This visual image is also customized to take into account the user's emotional state. The generated visual image is then sent to the device and displayed to the user.
[1371] Specific examples
[1372] For example, suppose a user is reading a particular novel and has read up to page 56. The user enters "I've read up to page 56" into the device and presses the send button. At this point, the emotion analysis engine recognizes that the user is excited. The server receives and analyzes this information, generating a summary, foreshadowing, and character information accordingly. The emotion analysis engine customizes the information to match the user's state of excitement and sends it to the device. If the user also enters a question such as "Was there any information about the perpetrator's motive?", the server will generate an answer to the question that takes the emotion analysis engine into consideration. Furthermore, if the user wishes to visualize a particular scene, a visual image can be generated by pressing the "Visualize" button. The emotion analysis engine customizes the visual image according to the user's emotion and displays it on the device.
[1373] Prompt Sentence Examples
[1374] "I'm 56 pages in. Can you give me a summary of the novel and what the characters are? If the sentiment analysis engine thinks I'm excited, please make sure the information is presented in a way that's enjoyable for the user."
[1375] The above is an embodiment of the present invention.
[1376] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1377] Step 1:
[1378] Enter the page number that the user has read.
[1379] Specific operation: The user enters "page 56" into the terminal and presses the send button. The terminal receives this input data and recognizes the user's emotional state (e.g., excitement) through an emotion analysis engine. The obtained data (page number and emotional state) is sent to the server.
[1380] Input: page number (56) and emotional state (excitement).
[1381] Data processing: Add emotional data using a sentiment analysis engine.
[1382] Output: Data including page number and emotional state.
[1383] Step 2:
[1384] The server analyzes the received page number and emotional state.
[1385] Specific operation: The server analyzes the received data and uses a natural language processing engine such as GPT-4 to generate a summary, foreshadowing, and character information up to the relevant page. In addition, an emotion analysis engine analyzes the user's emotional state and adjusts the method and content of information display.
[1386] Input: Data including page number and emotional state.
[1387] Data processing: Use a natural language processing engine to generate summaries, plot twists, and character information, and customize the information with a sentiment analysis engine.
[1388] Output: Customized summary, foreshadowing, and character information.
[1389] Step 3:
[1390] The terminal displays the information received from the server.
[1391] Specific operation: The device displays the information received from the server in a user-friendly interface (e.g., a web page or dedicated app). Information customized by the emotion analysis engine is also reflected here. For example, the text color may be brightened and the font size increased depending on the state of excitement.
[1392] Input: Customized summary, foreshadowing, and character information.
[1393] Data Calculation: Convert customized information into the appropriate display format.
[1394] Output: Display of adjusted information.
[1395] Step 4:
[1396] The user types in a question and sends it to the server.
[1397] Specific operation: The user enters a specific question into the terminal (e.g., "Is there any information about the perpetrator's motive?") and presses the send button. The terminal then sends this question to the server.
[1398] Input: A question entered by the user (e.g., "Was there any information on the perpetrator's motive?").
[1399] Data processing: The entered question is sent to the server.
[1400] Output: The query data sent to the server.
[1401] Step 5:
[1402] The server receives the query and generates the relevant information.
[1403] Specific operation: The server analyzes the question and extracts relevant information from the database and content. At this time, the emotion analysis engine analyzes the user's emotional state and adjusts the content and expression of the answer. The generated answer is then sent to the device.
[1404] Input: The query data sent to the server.
[1405] Data processing: Uses database search and sentiment analysis engines to generate relevant answers.
[1406] Output: The customized answer.
[1407] Step 6:
[1408] The device displays the received response to the user.
[1409] Specific operation: The device displays the answers received from the server in a user-friendly format. Customized answers are also applied using the sentiment analysis engine.
[1410] Input: Customized Answer.
[1411] Data Calculation: Convert customized answers into the appropriate display format.
[1412] Output: A display of the appropriately adjusted answer.
[1413] Step 7:
[1414] The user requests the generation of a visual image.
[1415] Specific operation: The user presses the "Image" button on the device to request the device to visualize a specific scene. The device then sends this request to the server.
[1416] Input: Visual image generation requested by the user.
[1417] Data processing: Send the request to the server.
[1418] Output: The visual image generation request sent to the server.
[1419] Step 8:
[1420] The server generates the visual image and sends it to the terminal.
[1421] How it works: The server uses a generative AI model (e.g., DALL-E or Stable Diffusion) to generate a visual image based on the prompt. This image is also customized to take into account the user's emotional state. The generated image is then sent to the device and displayed to the user.
[1422] Input: A visual image generation request.
[1423] Data processing: Visual images are generated using generative AI models and customized with a sentiment analysis engine.
[1424] Output: A customized visual image.
[1425] (Application example 2)
[1426] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1427] Conventional reading support systems have difficulty customizing the user's reading experience, and are particularly limited in providing information tailored to the user's emotions. As a result, users are unable to receive the information they desire quickly and appropriately, resulting in a poor quality reading experience. Furthermore, the generation of illustrations, which are visual representations, is fixed and not dynamically generated to match the user's emotions. There was a need for a system that could solve these issues and improve the user's reading experience.
[1428] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1429] In this invention, the server includes a means for analyzing the content based on the page number and generating information on the plot, foreshadowing, and characters, a means for generating and transmitting illustrations using a generative AI model, and a means for analyzing the user's emotions using an emotion engine and adjusting the information display method, thereby enabling information display and dynamic illustration generation according to the user's emotions.
[1430] "User" refers to an individual user of the system.
[1431] "Page number read" refers to the specific page number of the book that the user has currently read.
[1432] "Input means" refers to the interfaces and devices through which a user provides information to the system.
[1433] "Server" refers to a computer system that receives requests from users, analyzes them, and provides appropriate information.
[1434] "Means for transmitting to the server" refers to a communication means for transmitting the user's input information to the server.
[1435] "Means for analyzing the content and generating plot, plot twists, and character information" refers to the process by which the server analyzes the content of the book based on page numbers and generates summary information to provide to the user.
[1436] "Displaying means" refers to devices and interfaces for visually presenting information received from the server to a user.
[1437] "Means for inputting a question and sending it to the server" refers to an interface and communication means for a user to input a question to the system and send the question to the server.
[1438] "Means for generating and sending answers based on questions" refers to the process by which the server analyzes the user's question, derives an appropriate answer, and provides it to the user.
[1439] "Means for requesting the generation of illustrations" refers to an interface through which a user requests the system to create a visual image of a particular scene.
[1440] "Means for generating and transmitting illustrations using a generative AI model" refers to the process by which a server uses AI technology to dynamically generate illustrations and provide those illustrations to users.
[1441] An "emotion engine" is an engine that analyzes a user's emotions and adjusts the way information is displayed based on those emotions.
[1442] "Means of analyzing emotions and adjusting how information is displayed" refers to the process by which the emotion engine understands the user's current emotional state and changes the content and format of the display accordingly.
[1443] System Overview
[1444] This system includes four main components: the user, the device, the server, and the emotion engine. The user inputs the page number they have read and a question into the device, which then sends it to the server. The server analyzes the received information, generates appropriate data, and sends it to the device. During this process, the emotion engine analyzes the user's emotions and adjusts the information display and response content accordingly.
[1445] Enter the page number the user has finished reading
[1446] The user uses the terminal to input the page number they have just finished reading. This information is sent to the server via the terminal. The server also recognizes the emotion (e.g., excitement, sadness, doubt, etc.) with which the user seeks information.
[1447] Server analysis and information generation
[1448] The server analyzes the received page number and generates a list of the plot, foreshadowing, and characters up to that page. The information generated by the server is organized in the form of plot, foreshadowing, and character information and sent to the device. During this process, the emotion engine analyzes the user's emotions and adjusts the display method and content of the information.
[1449] Terminal display
[1450] The device displays the information received from the server in a user-friendly interface. The emotion engine customizes the interface based on the user's emotions. For example, if the user is excited, the display may change to a larger font size or change the color of the text. If the user is sad, the display may use warmer colors.
[1451] User questions and responses
[1452] When a user enters a question, the device sends it to the server. Based on the question, the server extracts relevant information from existing databases and content to generate an answer. During this process, the emotion engine analyzes the user's emotions and adjusts the content and expression of the answer. For example, if the user is sad, it can incorporate words of encouragement. This answer is sent to the device and displayed to the user.
[1453] Dynamic generation of illustrations
[1454] If a user wants a visual image of a particular scene, they can request the generation of an illustration by pressing the "Image" button. The device sends this request to the server, which then uses a generative AI model to generate an illustration of the scene. The generated illustration, which takes the user's emotions into account using an emotion engine, is sent to the device and displayed to the user.
[1455] Specific examples
[1456] For example, suppose a user is reading a particular novel. The user enters into the system that they have read up to page 56, and the emotion engine recognizes that the user is currently excited. The server receives and analyzes this information, and generates information on the plot, foreshadowing, and characters accordingly. The emotion engine customizes this information to match the user's state of excitement and sends it to the device. If the user asks, "Was there any information on the perpetrator's motive?", the server generates an answer to the question that takes the emotion engine into consideration. Furthermore, if the user wants to see an illustration of a particular scene, they can press the "Visualize" button, and an illustration will be generated. The emotion engine customizes the illustration according to the user's emotion and displays it on the device.
[1457] The specific hardware and software used
[1458] Hardware:
[1459] Smartphone (iOS or Android)
[1460] Server (equipped with a high-performance GPU)
[1461] Emotion engine (commonly known as "emotion analysis engine")
[1462] software:
[1463] Frontend: React Native
[1464] Backend: Node.js, Express
[1465] Database: MongoDB
[1466] Emotion analysis: Python, TensorFlow
[1467] Prompt Sentence Examples
[1468] "Page 56, Emotion: Excite, Question: Was there any information about the perpetrator's motive?"
[1469] The above is a specific embodiment for implementing this system.
[1470] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1471] Program processing steps
[1472] Step 1:
[1473] The user enters the page number they have read into the terminal.
[1474] Input: page number the user finished reading, emotional state (e.g., excited, sad)
[1475] Specific operation: The user enters the page number of the page they have just finished reading in the page number input interface of the application, selects their emotional state, and the result is input to the terminal.
[1476] Step 2:
[1477] The terminal transmits the input page number and emotional state to the server.
[1478] Input: User-entered page number and emotional state
[1479] Output: Data packet sent to the server
[1480] Specific operation: The device sends the input page number and emotional state as a data packet to the server using an HTTP POST request.
[1481] Step 3:
[1482] The server analyzes the content based on the page number and generates plot summary, foreshadowing, and character information.
[1483] Input: The page number received by the server
[1484] Output: Synopsis, foreshadowing, character information data
[1485] Specific operation: The server retrieves information up to the relevant page from the database, and based on this generates a plot summary, foreshadowing, and a list of characters. A Python script is executed.
[1486] Step 4:
[1487] The server uses an emotion engine to analyze the user's emotions and adjust how information is displayed.
[1488] Input: User's emotional state, generated information data
[1489] Output: Information data adjusted based on user sentiment
[1490] How it works: The emotion engine (a model using TensorFlow) analyzes the emotional state and adjusts the format and display style of the generated information data. For example, if the user is excited, the font size and color will be changed.
[1491] Step 5:
[1492] The server transmits the adjusted information data to the terminal.
[1493] Input: Adjusted information data
[1494] Output: Data packets sent to the device
[1495] Specific operation: The server assembles the adjusted information data into a data packet and sends it to the terminal. The communication method is an HTTP POST request.
[1496] Step 6:
[1497] The information received by the terminal is displayed in a user-friendly interface.
[1498] Input: Adjusted information data received from the server
[1499] Output: On-screen display
[1500] Specific operation: Based on the received information data, the device displays information through a customized interface according to the user's emotional state. The display process is performed using React Native.
[1501] Step 7:
[1502] The user enters a question and the device sends the question to the server.
[1503] Input: The question entered by the user
[1504] Output: Data packet sent to the server
[1505] Specific operation: The user inputs a question through the question input interface, and the device sends the question to the server via an HTTP POST request.
[1506] Step 8:
[1507] The server extracts relevant information from existing databases based on the question and generates an answer.
[1508] Input: User question
[1509] Output: Response data
[1510] What happens: The server searches a database for information related to the question and generates an answer. A Python script is executed to extract the requested information.
[1511] Step 9:
[1512] The emotion engine analyzes the user's emotions and adjusts the content and expression of the response.
[1513] Input: User's emotional state, generated answer data
[1514] Output: Answer data adjusted based on user sentiment
[1515] What it does: The emotion engine analyzes the user's emotional state and adjusts the wording of the generated response data accordingly, possibly adding words of encouragement.
[1516] Step 10:
[1517] The server transmits the adjusted response data to the terminal.
[1518] Input: Adjusted response data
[1519] Output: Data packets sent to the device
[1520] Specific operation: The server assembles the adjusted response data into a data packet and sends it to the terminal using an HTTP POST request.
[1521] Step 11:
[1522] The user requests the creation of an illustration, and the device sends the request to the server.
[1523] Input: Illustration generation request
[1524] Output: Data packet sent to the server
[1525] Specific operation: The user presses the "Image" button to request the creation of an illustration, and the device sends the request to the server via an HTTP POST request.
[1526] Step 12:
[1527] The server uses a generative AI model to generate and transmit illustrations.
[1528] Input: Illustration generation request
[1529] Output: Generated illustration data
[1530] Specific operation: The server uses a generative AI model (e.g., DALL-E) to generate illustrations according to the request. The generated illustration data is prepared.
[1531] Step 13:
[1532] The emotion engine customizes the generated illustrations according to the user's emotions.
[1533] Input: Generated illustration data, user's emotional state
[1534] Output: Adjusted illustration data
[1535] How it works: The emotion engine analyzes the user's emotional state and customizes the generated illustrations accordingly. For example, if the user is excited, the colors and details will be enhanced.
[1536] Step 14:
[1537] The server sends the adjusted illustration data to the terminal, which displays it to the user.
[1538] Input: Adjusted illustration data
[1539] Output: Illustration displayed on the terminal
[1540] Specific operation: The server assembles the adjusted illustration data into a data packet and sends it to the device. The device displays the received illustration data in a user-friendly interface. The display process is performed using React Native.
[1541] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1542] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1543] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1544] [Fourth embodiment]
[1545] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1546] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1547] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1548] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1549] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1550] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1551] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1552] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1553] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1554] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1555] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1556] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1557] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1558] The present invention is a system designed to enhance a user's reading experience, and is embodied in the following manner.
[1559] System Overview
[1560] This system has three main components: the user, the terminal, and the server. The user inputs the page number and question they have read into the terminal, which then sends it to the server. The server analyzes the received information, generates appropriate data, and sends it to the terminal to provide information to the user.
[1561] Enter the page number the user has finished reading
[1562] The user uses the terminal to input the page number they have just finished reading. This information is sent to the server via the terminal. At this stage, the user is ready to receive information from the system based on their progress in the story.
[1563] Server analysis and information generation
[1564] The server analyzes the received page number and generates a list of the plot, hints, and characters up to that page. This allows the user to review the main points of the story or notice hints. The information generated by the server is compiled in the form of plot, hints, and character information and sent to the terminal.
[1565] Terminal display
[1566] The device displays the information received from the server in a user-friendly interface, allowing users to easily understand the main points of the story and the relationships between characters. If users require specific information or detailed explanations, they can type their questions into the device.
[1567] User questions and responses
[1568] When a user inputs a question, the device sends the question to the server, which extracts relevant information from existing databases and content based on the question and generates an answer. This answer is then immediately sent to the device, which displays it to the user, allowing the user to deepen their understanding of the story.
[1569] Dynamic generation of illustrations
[1570] If a user wants a visual image of a particular scene, they can request the generation of an illustration by pressing the "Image" button. The device sends this request to the server, which uses an AI model to generate an illustration for the scene. The generated illustration is sent to the device and displayed to the user. This feature allows users to visually enjoy the story scenes.
[1571] Specific examples
[1572] For example, suppose a user is reading a particular novel. When the user inputs into the system that he or she has read up to page 56, the server receives and analyzes that information, generates information about the plot, foreshadowing, and characters up to that point, and sends it to the device. If the user asks, "Was there any information about the perpetrator's motive?", the server will provide an answer to that question. Furthermore, if the user wants to see an illustration of a particular scene, the illustration will be generated by pressing the "Image" button.
[1573] The above is an embodiment of the present invention, which provides a specific method for improving the user's reading experience.
[1574] The processing flow will be explained below.
[1575] Step 1:
[1576] The user inputs the page number of the page that has been read. The user inputs the page number of the page that has been read in the terminal interface. For example, the user inputs "56".
[1577] Step 2:
[1578] The terminal transmits the input page number to the server. The terminal generates an HTTP request for transmitting the input page number to the server and transmits it to the server.
[1579] Step 3:
[1580] The server analyzes the received page number and retrieves the data up to that page from the database or content management system based on the received page number.
[1581] Step 4:
[1582] The server analyzes the content and generates plot, foreshadowing, and character information. The server uses a data analysis algorithm to identify plot, foreshadowing, and characters based on page numbers and compiles each piece of information.
[1583] Step 5:
[1584] The server sends the generated information to the terminal. The server then sends the generated information about the plot, hints, and characters to the terminal as an HTTP response.
[1585] Step 6:
[1586] The terminal receives and displays information from the server. The terminal receives a response from the server and displays the plot, plot twists, and character information on the user interface.
[1587] Step 7:
[1588] The user enters a question. The user enters a question into the question input field on the device. For example, the user enters a question such as "What is the perpetrator's motive?"
[1589] Step 8:
[1590] The device sends the question to the server. The device sends the user's question to the server as an HTTP request.
[1591] Step 9:
[1592] The server generates an answer based on the question. The server searches for information related to the question from a database or existing content and generates an answer.
[1593] Step 10:
[1594] The server sends the generated answer to the terminal. The server sends the generated answer to the terminal as an HTTP response.
[1595] Step 11:
[1596] The terminal displays the answer from the server. The terminal receives the response from the server and displays the answer to the user on the user interface.
[1597] Step 12:
[1598] The user requests the creation of an illustration. If the user wants to see an illustration of a particular scene, they press the "Image" button.
[1599] Step 13:
[1600] The terminal sends a request to generate an illustration to the server. The terminal detects the button press event and sends a request to generate an illustration to the server.
[1601] Step 14:
[1602] The server uses the AI model to generate illustrations. The server uses the AI model for generating illustrations to generate illustrations based on the specified page content.
[1603] Step 15:
[1604] The server sends the generated illustration to the terminal. The server generates and sends an HTTP response to send the generated illustration to the terminal.
[1605] Step 16:
[1606] The terminal receives and displays the illustration. The terminal receives the illustration data from the server and displays it to the user on the user interface.
[1607] Example 1
[1608] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1609] Previous systems lacked sufficient means for users to check story details and important hints while reading, or to rely on character information to understand the story. Furthermore, obtaining visual images required separate searches, which was time-consuming. Furthermore, few systems could provide immediate answers to user questions or dynamically generate information based on the page being read. Therefore, an effective system to enhance users' reading experience was needed.
[1610] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1611] In this invention, the server includes means for inputting the page number of a page that a user has finished reading, means for transmitting the input page number to the server, means for the server to analyze the content based on the page number and generate information on the plot, foreshadowing, and characters, means for receiving and displaying information from the server, means for inputting a user's question and transmitting it to the server, means for the server to generate and transmit an answer based on the question, means for the user to request the generation of an illustration, means for the server to generate and transmit an illustration using an AI model, and means for the terminal to display information in a user-friendly interface. This enables the user to check the plot, foreshadowing, and information on characters in real time while reading, get instant answers to questions, and generate visual images.
[1612] "User" means a person who uses the system and takes action to enhance their reading experience.
[1613] A "terminal" is a device that allows a user to perform input operations and receive and display information from a server.
[1614] A "server" is a central computing device that analyzes information sent from users and terminals and generates and provides the necessary data.
[1615] The "page number read" is the number of the page that the user currently recognizes as having finished reading.
[1616] "Input means" refers to the interface or device that allows the user to input the page number they have read or questions into the terminal.
[1617] "Transmission means" refers to the function or device for transmitting information entered by the user from the terminal to the server.
[1618] The "analysis means" refers to a function that allows the server to process data based on the input page number and analyze the content.
[1619] The "synopsis generation means" refers to a function or method for summarizing the main points of the story up to the page that the server has finished reading.
[1620] The "foreshadowing generation means" refers to a function or method by which the server creates a list of important foreshadowings that appear in the story.
[1621] "Character information generation means" refers to the function or method by which the server analyzes and lists the characters in the story and their relationships.
[1622] "Display means" refers to the functions and methods by which the terminal presents the information received from the server in an easy-to-read format for the user.
[1623] "Question input means" refers to an interface or device that allows a user to input detailed information or specific questions into a terminal.
[1624] "Answer generation means" refers to the function or method by which the server generates related information based on the user's question and creates an answer.
[1625] The "illustration generation request means" refers to a function or method that allows a user to send a request from a terminal to a server in order to generate a visual image of a specific scene.
[1626] "AI model" refers to the artificial intelligence algorithms and functions that the server uses to generate illustrations and other data.
[1627] A "user-friendly interface" refers to the screen or method by which a device presents information to the user in an intuitive and easy-to-understand format.
[1628] This invention is a system designed to improve the user's reading experience, and it includes three main components: the terminal, the server, and the user. The following describes each flow in detail.
[1629] First, the user inputs the page number they have finished reading and a question into the terminal. The terminal is a mobile device such as a tablet or smartphone, which receives and processes the input information. The terminal is equipped with a user interface and is designed for intuitive operation. Specifically, the user might input "I have finished reading up to page 56" or a question such as "What is the perpetrator's motive?"
[1630] The device then sends the input information to the server, where it transmits the data via HTTP over Wi-Fi or mobile networks, and organizes the data in JSON format for efficient delivery to the server.
[1631] The server is a high-performance data server, such as AWS EC2, a widely used cloud service. The server searches the database based on the page number sent by the user and analyzes the content up to that page. This analysis uses an algorithm that utilizes natural language processing technology to dynamically generate the story's plot, hints, and character information. This allows the user to see the progress of the story and understand its key points.
[1632] When a user enters a question, the server generates an answer based on that question. Here, machine learning technology is used to extract relevant information from existing databases and derive appropriate answers. For example, in response to the question "What was the perpetrator's motive?", the server searches the database for relevant information, extracts the relevant parts, and generates an answer.
[1633] Furthermore, if a user wants a visual image of a specific scene, they can request the generation of an illustration by pressing the "Image" button on their device. This request is sent to the server, which then uses a generative AI model (e.g., DALL-E 2) to generate an illustration of the scene. The AI model generates an image of a specific scene based on a prompt. For example, a prompt might be, "Please generate an illustration of the scene where the protagonist confronts the criminal for the first time."
[1634] The information and illustrations generated by the server are then sent back to the device, which receives them and displays them in a user-friendly interface, allowing users to enjoy the story visually as well.
[1635] (Example of a prompt)
[1636] "I've finished reading page 56 so far. Please tell me the plot, foreshadowing, and characters."
[1637] "Did you have any information about the perpetrator's motive?"
[1638] "Generate an illustration of the scene where the protagonist confronts the culprit for the first time."
[1639] As described above, this system performs appropriate data processing and calculations based on the information entered by the user, providing the user with a fulfilling reading experience.
[1640] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1641] Step 1:
[1642] The user inputs the page number they have read and a question into the device. The user inputs the page number into the device's input field in the format of "I have now finished reading page 56." They also input questions such as "What was the perpetrator's motive?" The input is in text format, and the device receives it and stores it in its internal memory.
[1643] Step 2:
[1644] The device sends the input information to the server. The device converts the page number and question information entered by the user into JSON format and sends it to the server via an HTTP request. Specifically, the data is sent using Wi-Fi or a mobile network.
[1645] Input: A page number or question entered by the user
[1646] Output: Data converted to JSON format
[1647] Specific operation: The device converts the text data into JSON format and sends it to the server using an HTTP request.
[1648] Step 3:
[1649] The server analyzes the page number and generates a synopsis, foreshadowing, and character information. The server then analyzes the received JSON data and retrieves the content up to the relevant page from the database. Based on the retrieved data, the server uses natural language processing technology to analyze the content and generate information to provide to the user.
[1650] Input: JSON data of page numbers received from the terminal
[1651] Output: Generated plot, hints, and character information
[1652] Specific operation: The server searches the database, analyzes the retrieved data, and generates the required information.
[1653] Step 4:
[1654] The server sends the generated information to the terminal. The generated plot, foreshadowing, and character information are compiled in JSON format and sent to the terminal as an HTTP response.
[1655] Input: Generated plot, hints, character information
[1656] Output: Data in JSON format
[1657] Specific operation: The information generated by the server is converted into JSON format and sent to the terminal via an HTTP response.
[1658] Step 5:
[1659] The device displays the received information in a user-friendly manner. The device analyzes the received JSON data and presents the plot, plot twists, and character information to the user in a visually easy-to-understand format.
[1660] Input: JSON data received from the server
[1661] Output: Information displayed in a user-friendly format
[1662] Specific operation: The device parses the JSON data and displays it on the screen.
[1663] Step 6:
[1664] The user enters a follow-up question into the device, which then sends it to the server. The user enters a follow-up question into the device, such as "What is your specific motivation?" The device receives the question and sends it back to the server.
[1665] Input: Any additional questions entered by the user
[1666] Output: The question sent to the server
[1667] Specific operation: The device receives a new question, converts it into JSON format, and sends it to the server.
[1668] Step 7:
[1669] The server analyzes the question, generates relevant information, and sends it to the device. The server analyzes the question, extracts relevant information from an existing database, and generates an answer. The generated answer is compiled in JSON format and sent to the device.
[1670] Input: JSON data of the question received from the terminal
[1671] Output: JSON data of the generated answer
[1672] What it does: The server searches the database, extracts relevant information, and generates an answer.
[1673] Step 8:
[1674] The terminal displays the answer and, if necessary, sends a request to the server to generate an illustration. The terminal displays the answer to the user, and if necessary, the user presses the "Image" button to request the generation of an illustration.
[1675] Input: JSON data of the response received from the server
[1676] Output: Answers displayed to the user and illustration generation request
[1677] Specific operation: The device displays the answer on the screen, and when the user presses the "Image" button, it sends a request to the server to generate an illustration.
[1678] Step 9:
[1679] The server generates illustrations and sends them to the device. The server receives illustration generation requests and generates illustrations for the corresponding scenes using a generative AI model. The generated illustrations are then sent to the device.
[1680] Input: Illustration generation request received from the device
[1681] Output: Generated illustration data
[1682] Specific operation: The server sends a prompt to the AI model and sends the generated illustration data to the terminal.
[1683] Step 10:
[1684] The terminal displays the generated illustration. The terminal displays the illustration data received from the server to the user, allowing the user to visually enjoy a specific scene in the story.
[1685] Input: illustration data received from the server
[1686] Output: The illustration displayed to the user
[1687] Specific operation: The device converts the illustration data for display on the screen and shows it to the user.
[1688] These are the specific processing steps of the present invention, which allow users to check plot summaries, hints, and character information in real time while reading, get instant answers to questions, and generate visual images.
[1689] (Application example 1)
[1690] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1691] Conventional reading systems make it difficult for users to obtain summaries, plot twists, and character information according to their reading progress. Furthermore, they lack the ability to generate illustrations to visually enjoy specific scenes, limiting the reading experience. Furthermore, when users have questions about the story, there is a lack of a way to get immediate answers. It is necessary to solve these problems and improve the user's reading experience.
[1692] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1693] In this invention, the server includes means for analyzing the content based on the page number that the user has read and generating information on the plot, foreshadowing, and characters, means for generating answers based on the user's questions, and means for generating illustrations using a generative AI model. This allows the user to easily obtain a summary of the content, foreshadowing, and character information according to their reading progress, as well as enjoy visual images of specific scenes and receive immediate answers to questions that will help them better understand the story.
[1694] "User" refers to an individual who utilizes the system to enhance their reading experience.
[1695] "Page number read" refers to the number of the specific page of the book that the user has currently read.
[1696] "Server" refers to the central processing unit that analyzes information sent by users and generates and transmits necessary data.
[1697] A "plot" is a concise summary of the main developments and events of a story.
[1698] A "foreshadowing" is information that provides hints or elements that will become important later in the development of the story.
[1699] "Character information" refers to detailed information about the characters' names, backgrounds, and roles in the story.
[1700] "Question" refers to the act of a user inquiring of the server about specific matters regarding the content of the story.
[1701] "Answer" refers to information generated by the server based on the user's question.
[1702] An "illustration" is an image or painting that visually depicts a particular scene from a story.
[1703] A "generative AI model" refers to an algorithm or system that uses artificial intelligence technology to generate visual images or answers in response to user requests.
[1704] "Prompt sentence" refers to the text input used to prompt a generative AI model to generate an illustration.
[1705] "Display means" refers to an interface for displaying the generated information and illustrations in a form that can be visually recognized by the user.
[1706] This invention is a system designed to improve the user's reading experience, and it includes three main elements: a user, a terminal, and a server. The user uses the terminal to input the page number they have read and questions, and the terminal sends that information to the server. The server analyzes the received information, generates appropriate data, and sends it to the terminal to provide information to the user.
[1707] System Overview
[1708] The system includes the following elements:
[1709] 1. Enter page number:
[1710] The user inputs the page number they have read into the device, which allows the device to provide information based on the progress of the story.
[1711] 2. Information analysis and generation:
[1712] The server parses the page number it receives and generates a list of plot, plot twists, and characters up to that page, then compiles this information in a user-friendly format and sends it to the device.
[1713] 3. Questions and Answers:
[1714] When a user enters a question, the device sends it to the server, which extracts relevant information from an existing database and immediately generates an answer that is sent to the device.
[1715] 4. Dynamic generation of illustrations:
[1716] When a user wants a visual image of a particular scene, they press the "Image" button to request the generation of an illustration. The server uses a generative AI model to generate an illustration of the scene, which is then sent to the device for display.
[1717] Hardware and software used
[1718] Hardware:
[1719] User devices: smartphones, tablets
[1720] Server: Server equipment that performs high-performance calculations
[1721] software:
[1722] Server API: e.g. Flask or Django
[1723] AI model: Generative AI models such as Stable Diffusion and DALL-E are used to generate illustrations.
[1724] HTTP request library: Requests
[1725] Example of a system
[1726] 1. Enter and analyze page numbers:
[1727] For example, if a user inputs into the system that they have read a particular novel and have progressed to page 56, the server receives and analyzes that information, generating information about the plot, foreshadowing, and characters up to that point, and sending it to the terminal.
[1728] Example prompt: "Use case: I've read up to page 56. What's the plot and foreshadowing so far?"
[1729] 2. Questions and Answers:
[1730] If a user asks, "Is there any information about the perpetrator's motive?", the server searches for relevant information based on the question, generates an answer, and sends it to the device.
[1731] Example prompt: "Use case: I want to ask for information about the perpetrator's motive."
[1732] 3. Illustration generation:
[1733] If a user wants to see a visual image of a particular scene, they can press the "Visualize" button, which will generate an illustration. For example, a request to generate an illustration of a scene in which the protagonist is in conflict is sent to the server, and the illustration is generated and displayed using a generative AI model.
[1734] Example prompt: "Use case: Visualize a scene where the protagonist is in conflict."
[1735] The above is an embodiment of the present invention, which provides a specific method for improving the user's reading experience.
[1736] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1737] Step 1:
[1738] The user uses the terminal to enter the page number that they have read.
[1739] Input: page number read
[1740] Specific behavior: The user enters the page number into the input field on the device and presses the "Submit" button.
[1741] Step 2:
[1742] The terminal transmits the input page number to the server.
[1743] Input: The page number entered by the user
[1744] Output: Page number data sent to the server
[1745] Specific operation: The terminal sends the entered page number to the server via an HTTP request.
[1746] Step 3:
[1747] The server analyzes the page number received and generates information on the plot, plot twists, and characters up to that page.
[1748] Input: Submitted page number data
[1749] Output: Synopsis, foreshadowing, character information
[1750] Specific operation: The server searches the database, analyzes the content up to the relevant page, and generates information on the plot, foreshadowing, and characters.
[1751] Step 4:
[1752] The server transmits the generated information to the terminal.
[1753] Input: Synopsis, foreshadowing, character information
[1754] Output: Information data sent to the terminal
[1755] Specific operation: The server sends the generated information to the terminal in a data format such as JSON.
[1756] Step 5:
[1757] The terminal displays the information received from the server in a user-friendly interface.
[1758] Input: Information data sent from the server
[1759] Output: Plot, plot twists, and character information displayed on the screen
[1760] Specific operation: The device analyzes the received information and generates and displays HTML and CSS to display on the interface.
[1761] Step 6:
[1762] The user enters a specific question, and the device sends the question to the server.
[1763] Input: The question entered by the user
[1764] Output: Question data sent to the server
[1765] Specific operation: The user enters a question in the question input field and presses the "Send" button. The device sends the question to the server via an HTTP request.
[1766] Step 7:
[1767] The server extracts relevant information from existing databases based on the question and generates an answer.
[1768] Input: Submitted question data
[1769] Output: relevant information and generated answers
[1770] What it does: The server analyzes the question, extracts relevant information from the database, and generates an answer.
[1771] Step 8:
[1772] The server generates a response and sends it to the terminal.
[1773] Input: Generated answer
[1774] Output: Answer data sent to the device
[1775] Specific operation: The server converts the generated response into a data format and sends it to the terminal.
[1776] Step 9:
[1777] The device displays the response received from the server to the user.
[1778] Input: Response data sent from the server
[1779] Output: Answer displayed on the screen
[1780] Specific operation: The device analyzes the received response and generates and displays an interface to display to the user.
[1781] Step 10:
[1782] A user requests a visual image of a particular scene, and the device sends the request to the server.
[1783] Input: The prompt for the scene the user wants
[1784] Output: Request data sent to the server
[1785] Specific operation: The user inputs a prompt sentence about a specific scene and presses the "Image" button. The device sends a request to the server.
[1786] Step 11:
[1787] The server uses a generative AI model to generate illustrations of the requested scene.
[1788] Input: Request data (prompt statement)
[1789] Output: Generated illustration data
[1790] Specific operation: The server inputs a prompt sentence into the generative AI model and receives the generated illustration.
[1791] Step 12:
[1792] The server sends the generated illustration to the terminal.
[1793] Input: Generated illustration data
[1794] Output: Illustration data sent to the device
[1795] Specific operation: The server converts the illustration data into an appropriate format and sends it to the terminal.
[1796] Step 13:
[1797] The terminal displays the illustrations received from the server in a user-friendly interface.
[1798] Input: Illustration data sent from the server
[1799] Output: Illustration displayed on the screen
[1800] Specific operation: The terminal analyzes the received illustration data and generates and displays an interface for visual display.
[1801] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1802] The present invention is an advanced system for improving a user's reading experience, incorporating an emotion engine for recognizing a user's emotions and providing corresponding information. Specific embodiments for implementing the present invention will be described below.
[1803] System Overview
[1804] This system includes four main components: the user, the device, the server, and the emotion engine. The user inputs the page number they have read and a question into the device, which then sends it to the server. The server analyzes the received information, generates appropriate data, and sends it to the device. During this process, the emotion engine analyzes the user's emotions and adjusts the information display and response content.
[1805] Enter the page number the user has finished reading
[1806] The user uses the terminal to input the page number they have just finished reading. This information is sent to the server via the terminal. The server also recognizes the emotion (e.g., excitement, sadness, doubt, etc.) with which the user seeks information.
[1807] Server analysis and information generation
[1808] The server analyzes the received page number and generates a list of the plot, foreshadowing, and characters up to that page. The information generated by the server is organized in the form of plot, foreshadowing, and character information and sent to the device. During this process, the emotion engine analyzes the user's emotions and adjusts the display method and content of the information.
[1809] Terminal display
[1810] The device receives information from the server and displays it in a user-friendly interface. The emotion engine customizes the interface based on the user's emotions. For example, if the user is excited, the display may change to a different color or larger font size.
[1811] User questions and responses
[1812] When a user enters a question, the device sends it to the server. Based on the question, the server extracts relevant information from existing databases and content to generate an answer. During this process, the emotion engine analyzes the user's emotions and adjusts the content and expression of the answer. For example, if the user is sad, it can incorporate words of encouragement. This answer is sent to the device and displayed to the user.
[1813] Dynamic generation of illustrations
[1814] If a user wants a visual image of a particular scene, they can request the generation of an illustration by pressing the "Image" button. The device sends this request to the server, which uses an AI model to generate an illustration of the scene. The generated illustration takes the user's emotions into account using an emotion engine, and is sent to the device and displayed to the user.
[1815] Specific examples
[1816] For example, suppose a user is reading a particular novel. The user enters into the system that they have read up to page 56, and the emotion engine recognizes that the user is currently excited. The server receives and analyzes this information, and generates information on the plot, foreshadowing, and characters accordingly. The emotion engine customizes this information to match the user's state of excitement and sends it to the device. If the user asks, "Was there any information on the perpetrator's motive?", the server generates an answer to the question that takes the emotion engine into consideration. Furthermore, if the user wants to see an illustration of a particular scene, they can press the "Visualize" button, and an illustration will be generated. The emotion engine customizes the illustration according to the user's emotion and displays it on the device.
[1817] The above is an embodiment of the present invention, which provides a specific method for taking user emotions into consideration to improve the reading experience.
[1818] The processing flow will be explained below.
[1819] Step 1:
[1820] The user inputs the page number of the page that has been read. The user inputs the page number of the page that has been read in the terminal interface. For example, the user inputs "56".
[1821] Step 2:
[1822] The device sends the entered page number and the user's emotional state to the server. The device then generates an HTTP request to send the emotional data analyzed by the emotion engine and the page number to the server, and sends it to the server.
[1823] Step 3:
[1824] The server analyzes the received page number and emotional state. Based on the received page number, the server retrieves data up to that page from a database or content management system and analyzes the user's emotions.
[1825] Step 4:
[1826] The server generates information on the plot, plot twists, and characters, and then adjusts it based on the user's emotions using an emotion engine.The server uses a data analysis algorithm to identify the plot, plot twists, and characters based on page numbers, and the emotion engine customizes the information according to the user's emotions.
[1827] Step 5:
[1828] The server sends the generated information to the terminal. The server then sends the generated plot, foreshadowing, and character information to the terminal as emotion-adjusted information in an HTTP response.
[1829] Step 6:
[1830] The terminal receives and displays information from the server. The terminal receives a response from the server and displays customized plot, plot twists, and character information on the user interface.
[1831] Step 7:
[1832] The user enters a question. The user enters a question into the question input field on the device. For example, the user enters a question such as "What is the perpetrator's motive?"
[1833] Step 8:
[1834] The device sends the question and emotional state to the server. The device generates an HTTP request to send the question including the user's emotional data to the server and sends it to the server.
[1835] Step 9:
[1836] The server generates an answer based on the question and adjusts it based on the user's emotions. The server searches for information related to the question from a database or existing content and generates an answer. The generated answer is adjusted according to the user's emotions by an emotion engine.
[1837] Step 10:
[1838] The server sends the generated answer to the terminal, and the server sends the adjusted answer to the terminal as an HTTP response.
[1839] Step 11:
[1840] The device displays the answer from the server. The device receives the response from the server and displays the answer adjusted to match the user's emotions on the user interface.
[1841] Step 12:
[1842] The user requests the creation of an illustration. If the user wants to see an illustration of a particular scene, they press the "Image" button.
[1843] Step 13:
[1844] The device sends a request to generate an illustration and the user's emotional state to the server. The device detects the button press event and sends a request to generate an illustration and the user's emotional state to the server.
[1845] Step 14:
[1846] The server uses an AI model to generate illustrations and adjust them based on the user's emotions.The server uses an AI model for generating illustrations to generate illustrations based on the specified page content, and the emotion engine adjusts them according to the user's emotions.
[1847] Step 15:
[1848] The server sends the generated illustration to the device. The server then generates and sends an HTTP response to send the emotion-adjusted illustration to the device.
[1849] Step 16:
[1850] The terminal receives and displays the illustration. The terminal receives the illustration data from the server and displays the illustration adjusted based on the emotion on the user interface.
[1851] Example 2
[1852] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1853] Previous systems for improving the reading experience did not provide information that took the user's emotions into consideration, and they had the problem of being unable to display or respond appropriately to the user's emotional state. Furthermore, the generation of visual images was not tailored to the user's needs, and only a uniform response was possible. This limited the user's reading experience, making it difficult to provide a customized experience tailored to each individual's emotional state.
[1854] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for generating information on a summary, foreshadowing, and characters based on page numbers, a means for adjusting the information content and display method based on emotion analysis, and a means for generating visual images using a generative AI model. This makes it possible to provide and display information according to the user's emotional state, and to provide a reading experience customized for each user.
[1855] A "user" is someone who uses the system to enhance their reading experience.
[1856] A "terminal" is a device that allows a user to input the page number of a page that has been read, input a question, or request the generation of a visual image.
[1857] The "server" is a central processing unit that analyzes the information sent by the user, generates answers and visual images, and sends them to the terminal.
[1858] A "page number" is a number that indicates the specific page of a book that the user has finished reading.
[1859] "Emotion analysis means" refers to technology for recognizing and analyzing a user's emotional state.
[1860] A "generative AI model" is an algorithm or software that uses AI technology to generate visual images based on specific prompts.
[1861] "Visual images" are illustrations or drawings of specific scenes requested by users.
[1862] A "summary" is a short summary of the contents up to a specific page number.
[1863] A "foreshadowing" is an element in a story that contains information or hints that will be important for later developments.
[1864] "Characters" is a list and profiles of characters that appear in the book.
[1865] A "question" is a question about specific information that a user inputs into the system.
[1866] "Means for adjusting information content and display method" refers to technology that changes the display format and content of information provided based on the user's emotional state.
[1867] This invention is a system for improving a user's reading experience, and in particular, has the function of customizing information according to the user's emotional state. This system mainly consists of a user (terminal), a terminal, a server, and an emotion analysis engine.
[1868] Enter the page number the user has read
[1869] The user inputs the page number they have finished reading into the device. The device receives this information and also analyzes the user's emotional state using an emotion analysis engine. For example, the device uses emotion recognition technology such as Microsoft Azure's Emotion API. The analyzed data (page number and emotional state) is sent to the server.
[1870] Server analysis and information generation
[1871] The server analyzes the received page number and emotion data. Based on the page number, a natural language processing engine (such as OpenAI's GPT-4) is used to generate a summary, plot twists, and character information. The emotion analysis engine also analyzes the user's emotional state and adjusts the display and content of the information. For example, if the user is excited, the text color or font size is changed. This adjusted information is then sent to the device.
[1872] Terminal display
[1873] The device displays the information received from the server in a user-friendly interface, which can be provided as a web page or a dedicated application. Information customized by the sentiment analysis engine (such as text color and font size) is also reflected here.
[1874] User questions and responses
[1875] The user enters a specific question into the device and presses the send button. The device then sends this question to the server. The server receives the question, extracts relevant information from the database and content, and generates an answer. Again, the sentiment analysis engine analyzes the user's emotions and adjusts the content and expression of the answer. The generated answer is sent to the device and displayed to the user.
[1876] Dynamic generation of illustrations
[1877] When a user requests a visualization of a particular scene, they can press the "Image" button to send the request to the device. The device then sends this request to the server. The server uses a generative AI model (e.g., DALL-E or Stable Diffusion) to generate a visual image based on the prompt. This visual image is also customized to take into account the user's emotional state. The generated visual image is then sent to the device and displayed to the user.
[1878] Specific examples
[1879] For example, suppose a user is reading a particular novel and has read up to page 56. The user enters "I've read up to page 56" into the device and presses the send button. At this point, the emotion analysis engine recognizes that the user is excited. The server receives and analyzes this information, generating a summary, foreshadowing, and character information accordingly. The emotion analysis engine customizes the information to match the user's state of excitement and sends it to the device. If the user also enters a question such as "Was there any information about the perpetrator's motive?", the server will generate an answer to the question that takes the emotion analysis engine into consideration. Furthermore, if the user wishes to visualize a particular scene, a visual image can be generated by pressing the "Visualize" button. The emotion analysis engine customizes the visual image according to the user's emotion and displays it on the device.
[1880] Prompt Sentence Examples
[1881] "I'm 56 pages in. Can you give me a summary of the novel and what the characters are? If the sentiment analysis engine thinks I'm excited, please make sure the information is presented in a way that's enjoyable for the user."
[1882] The above is an embodiment of the present invention.
[1883] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1884] Step 1:
[1885] Enter the page number that the user has read.
[1886] Specific operation: The user enters "page 56" into the terminal and presses the send button. The terminal receives this input data and recognizes the user's emotional state (e.g., excitement) through an emotion analysis engine. The obtained data (page number and emotional state) is sent to the server.
[1887] Input: page number (56) and emotional state (excitement).
[1888] Data processing: Add emotional data using a sentiment analysis engine.
[1889] Output: Data including page number and emotional state.
[1890] Step 2:
[1891] The server analyzes the received page number and emotional state.
[1892] Specific operation: The server analyzes the received data and uses a natural language processing engine such as GPT-4 to generate a summary, foreshadowing, and character information up to the relevant page. In addition, an emotion analysis engine analyzes the user's emotional state and adjusts the method and content of information display.
[1893] Input: Data including page number and emotional state.
[1894] Data processing: Use a natural language processing engine to generate summaries, plot twists, and character information, and customize the information with a sentiment analysis engine.
[1895] Output: Customized summary, foreshadowing, and character information.
[1896] Step 3:
[1897] The terminal displays the information received from the server.
[1898] Specific operation: The device displays the information received from the server in a user-friendly interface (e.g., a web page or dedicated app). Information customized by the emotion analysis engine is also reflected here. For example, the text color may be brightened and the font size increased depending on the state of excitement.
[1899] Input: Customized summary, foreshadowing, and character information.
[1900] Data Calculation: Convert customized information into the appropriate display format.
[1901] Output: Display of adjusted information.
[1902] Step 4:
[1903] The user types in a question and sends it to the server.
[1904] Specific operation: The user enters a specific question into the terminal (e.g., "Is there any information about the perpetrator's motive?") and presses the send button. The terminal then sends this question to the server.
[1905] Input: A question entered by the user (e.g., "Was there any information on the perpetrator's motive?").
[1906] Data processing: The entered question is sent to the server.
[1907] Output: The query data sent to the server.
[1908] Step 5:
[1909] The server receives the query and generates the relevant information.
[1910] Specific operation: The server analyzes the question and extracts relevant information from the database and content. At this time, the emotion analysis engine analyzes the user's emotional state and adjusts the content and expression of the answer. The generated answer is then sent to the device.
[1911] Input: The query data sent to the server.
[1912] Data processing: Uses database search and sentiment analysis engines to generate relevant answers.
[1913] Output: The customized answer.
[1914] Step 6:
[1915] The device displays the received response to the user.
[1916] Specific operation: The device displays the answers received from the server in a user-friendly format. Customized answers are also applied using the sentiment analysis engine.
[1917] Input: Customized Answer.
[1918] Data Calculation: Convert customized answers into the appropriate display format.
[1919] Output: A display of the appropriately adjusted answer.
[1920] Step 7:
[1921] The user requests the generation of a visual image.
[1922] Specific operation: The user presses the "Image" button on the device to request the device to visualize a specific scene. The device then sends this request to the server.
[1923] Input: Visual image generation requested by the user.
[1924] Data processing: Send the request to the server.
[1925] Output: The visual image generation request sent to the server.
[1926] Step 8:
[1927] The server generates the visual image and sends it to the terminal.
[1928] How it works: The server uses a generative AI model (e.g., DALL-E or Stable Diffusion) to generate a visual image based on the prompt. This image is also customized to take into account the user's emotional state. The generated image is then sent to the device and displayed to the user.
[1929] Input: A visual image generation request.
[1930] Data processing: Visual images are generated using generative AI models and customized with a sentiment analysis engine.
[1931] Output: A customized visual image.
[1932] (Application example 2)
[1933] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1934] Conventional reading support systems have difficulty customizing the user's reading experience, and are particularly limited in providing information tailored to the user's emotions. As a result, users are unable to receive the information they desire quickly and appropriately, resulting in a poor quality reading experience. Furthermore, the generation of illustrations, which are visual representations, is fixed and not dynamically generated to match the user's emotions. There was a need for a system that could solve these issues and improve the user's reading experience.
[1935] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1936] In this invention, the server includes a means for analyzing the content based on the page number and generating information on the plot, foreshadowing, and characters, a means for generating and transmitting illustrations using a generative AI model, and a means for analyzing the user's emotions using an emotion engine and adjusting the information display method, thereby enabling information display and dynamic illustration generation according to the user's emotions.
[1937] "User" refers to an individual user of the system.
[1938] "Page number read" refers to the specific page number of the book that the user has currently read.
[1939] "Input means" refers to the interfaces and devices through which a user provides information to the system.
[1940] "Server" refers to a computer system that receives requests from users, analyzes them, and provides appropriate information.
[1941] "Means for transmitting to the server" refers to a communication means for transmitting the user's input information to the server.
[1942] "Means for analyzing the content and generating plot, plot twists, and character information" refers to the process by which the server analyzes the content of the book based on page numbers and generates summary information to provide to the user.
[1943] "Displaying means" refers to devices and interfaces for visually presenting information received from the server to a user.
[1944] "Means for inputting a question and sending it to the server" refers to an interface and communication means for a user to input a question to the system and send the question to the server.
[1945] "Means for generating and sending answers based on questions" refers to the process by which the server analyzes the user's question, derives an appropriate answer, and provides it to the user.
[1946] "Means for requesting the generation of illustrations" refers to an interface through which a user requests the system to create a visual image of a particular scene.
[1947] "Means for generating and transmitting illustrations using a generative AI model" refers to the process by which a server uses AI technology to dynamically generate illustrations and provide those illustrations to users.
[1948] An "emotion engine" is an engine that analyzes a user's emotions and adjusts the way information is displayed based on those emotions.
[1949] "Means of analyzing emotions and adjusting how information is displayed" refers to the process by which the emotion engine understands the user's current emotional state and changes the content and format of the display accordingly.
[1950] System Overview
[1951] This system includes four main components: the user, the device, the server, and the emotion engine. The user inputs the page number they have read and a question into the device, which then sends it to the server. The server analyzes the received information, generates appropriate data, and sends it to the device. During this process, the emotion engine analyzes the user's emotions and adjusts the information display and response content accordingly.
[1952] Enter the page number the user has finished reading
[1953] The user uses the terminal to input the page number they have just finished reading. This information is sent to the server via the terminal. The server also recognizes the emotion (e.g., excitement, sadness, doubt, etc.) with which the user seeks information.
[1954] Server analysis and information generation
[1955] The server analyzes the received page number and generates a list of the plot, foreshadowing, and characters up to that page. The information generated by the server is organized in the form of plot, foreshadowing, and character information and sent to the device. During this process, the emotion engine analyzes the user's emotions and adjusts the display method and content of the information.
[1956] Terminal display
[1957] The device displays the information received from the server in a user-friendly interface. The emotion engine customizes the interface based on the user's emotions. For example, if the user is excited, the display may change to a larger font size or change the color of the text. If the user is sad, the display may use warmer colors.
[1958] User questions and responses
[1959] When a user enters a question, the device sends it to the server. Based on the question, the server extracts relevant information from existing databases and content to generate an answer. During this process, the emotion engine analyzes the user's emotions and adjusts the content and expression of the answer. For example, if the user is sad, it can incorporate words of encouragement. This answer is sent to the device and displayed to the user.
[1960] Dynamic generation of illustrations
[1961] If a user wants a visual image of a particular scene, they can request the generation of an illustration by pressing the "Image" button. The device sends this request to the server, which then uses a generative AI model to generate an illustration of the scene. The generated illustration, which takes the user's emotions into account using an emotion engine, is sent to the device and displayed to the user.
[1962] Specific examples
[1963] For example, suppose a user is reading a particular novel. The user enters into the system that they have read up to page 56, and the emotion engine recognizes that the user is currently excited. The server receives and analyzes this information, and generates information on the plot, foreshadowing, and characters accordingly. The emotion engine customizes this information to match the user's state of excitement and sends it to the device. If the user asks, "Was there any information on the perpetrator's motive?", the server generates an answer to the question that takes the emotion engine into consideration. Furthermore, if the user wants to see an illustration of a particular scene, they can press the "Visualize" button, and an illustration will be generated. The emotion engine customizes the illustration according to the user's emotion and displays it on the device.
[1964] The specific hardware and software used
[1965] Hardware:
[1966] Smartphone (iOS or Android)
[1967] Server (equipped with a high-performance GPU)
[1968] Emotion engine (commonly known as "emotion analysis engine")
[1969] software:
[1970] Frontend: React Native
[1971] Backend: Node.js, Express
[1972] Database: MongoDB
[1973] Emotion analysis: Python, TensorFlow
[1974] Prompt Sentence Examples
[1975] "Page 56, Emotion: Excite, Question: Was there any information about the perpetrator's motive?"
[1976] The above is a specific embodiment for implementing this system.
[1977] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1978] Program processing steps
[1979] Step 1:
[1980] The user enters the page number they have read into the terminal.
[1981] Input: page number the user finished reading, emotional state (e.g., excited, sad)
[1982] Specific operation: The user enters the page number of the page they have just finished reading in the page number input interface of the application, selects their emotional state, and the result is input to the terminal.
[1983] Step 2:
[1984] The terminal transmits the input page number and emotional state to the server.
[1985] Input: User-entered page number and emotional state
[1986] Output: Data packet sent to the server
[1987] Specific operation: The device sends the input page number and emotional state as a data packet to the server using an HTTP POST request.
[1988] Step 3:
[1989] The server analyzes the content based on the page number and generates plot summary, foreshadowing, and character information.
[1990] Input: The page number received by the server
[1991] Output: Synopsis, foreshadowing, character information data
[1992] Specific operation: The server retrieves information up to the relevant page from the database, and based on this generates a plot summary, foreshadowing, and a list of characters. A Python script is executed.
[1993] Step 4:
[1994] The server uses an emotion engine to analyze the user's emotions and adjust how information is displayed.
[1995] Input: User's emotional state, generated information data
[1996] Output: Information data adjusted based on user sentiment
[1997] How it works: The emotion engine (a model using TensorFlow) analyzes the emotional state and adjusts the format and display style of the generated information data. For example, if the user is excited, the font size and color will be changed.
[1998] Step 5:
[1999] The server transmits the adjusted information data to the terminal.
[2000] Input: Adjusted information data
[2001] Output: Data packets sent to the device
[2002] Specific operation: The server assembles the adjusted information data into a data packet and sends it to the terminal. The communication method is an HTTP POST request.
[2003] Step 6:
[2004] The information received by the terminal is displayed in a user-friendly interface.
[2005] Input: Adjusted information data received from the server
[2006] Output: On-screen display
[2007] Specific operation: Based on the received information data, the device displays information through a customized interface according to the user's emotional state. The display process is performed using React Native.
[2008] Step 7:
[2009] The user enters a question and the device sends the question to the server.
[2010] Input: The question entered by the user
[2011] Output: Data packet sent to the server
[2012] Specific operation: The user inputs a question through the question input interface, and the device sends the question to the server via an HTTP POST request.
[2013] Step 8:
[2014] The server extracts relevant information from existing databases based on the question and generates an answer.
[2015] Input: User question
[2016] Output: Response data
[2017] What happens: The server searches a database for information related to the question and generates an answer. A Python script is executed to extract the requested information.
[2018] Step 9:
[2019] The emotion engine analyzes the user's emotions and adjusts the content and expression of the response.
[2020] Input: User's emotional state, generated answer data
[2021] Output: Answer data adjusted based on user sentiment
[2022] What it does: The emotion engine analyzes the user's emotional state and adjusts the wording of the generated response data accordingly, possibly adding words of encouragement.
[2023] Step 10:
[2024] The server transmits the adjusted response data to the terminal.
[2025] Input: Adjusted response data
[2026] Output: Data packets sent to the device
[2027] Specific operation: The server assembles the adjusted response data into a data packet and sends it to the terminal using an HTTP POST request.
[2028] Step 11:
[2029] The user requests the creation of an illustration, and the device sends the request to the server.
[2030] Input: Illustration generation request
[2031] Output: Data packet sent to the server
[2032] Specific operation: The user presses the "Image" button to request the creation of an illustration, and the device sends the request to the server via an HTTP POST request.
[2033] Step 12:
[2034] The server uses a generative AI model to generate and transmit illustrations.
[2035] Input: Illustration generation request
[2036] Output: Generated illustration data
[2037] Specific operation: The server uses a generative AI model (e.g., DALL-E) to generate illustrations according to the request. The generated illustration data is prepared.
[2038] Step 13:
[2039] The emotion engine customizes the generated illustrations according to the user's emotions.
[2040] Input: Generated illustration data, user's emotional state
[2041] Output: Adjusted illustration data
[2042] How it works: The emotion engine analyzes the user's emotional state and customizes the generated illustrations accordingly. For example, if the user is excited, the colors and details will be enhanced.
[2043] Step 14:
[2044] The server sends the adjusted illustration data to the terminal, which displays it to the user.
[2045] Input: Adjusted illustration data
[2046] Output: Illustration displayed on the terminal
[2047] Specific operation: The server assembles the adjusted illustration data into a data packet and sends it to the device. The device displays the received illustration data in a user-friendly interface. The display process is performed using React Native.
[2048] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2049] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2050] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2051] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2052] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2053] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulti...
Claims
1. a means for the user to input the page number that they have read; means for transmitting the input page number to a server; A means for the server to analyze the content based on the page number and generate information on the plot, hints, and characters; means for receiving and displaying information from the server; a means for inputting a user's question and transmitting it to a server; means for the server to generate and transmit an answer based on the question; A means for users to request the generation of illustrations; A server uses an AI model to generate and transmit illustrations; A system including:
2. A means for dynamically generating plot, plot twist, and character information based on page numbers analyzed by the server, The system of claim 1 .
3. means for the server to extract relevant information from an existing database based on the query and generate an answer; The system of claim 1 .
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A