system
The system efficiently generates and presents summaries to help users recall e-book content and resume tasks by recording progress and using AI to provide tailored summaries in text and audio formats.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-11
- Publication Date
- 2026-04-23
AI Technical Summary
Users face difficulties in efficiently recalling previous content without rereading when they interrupt reading an e-book, especially during limited time periods, hindering efficient learning.
A system that records the user's reading progress and uses AI to generate a summary up to the interruption point, providing it in both text and audio formats, allowing quick recall of key information.
Enables users to efficiently resume reading by quickly recalling important information, improving learning efficiency and ensuring smooth workflow in various tasks.
Smart Images

Figure 2026069138000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In modern busy lives, it is difficult for users to efficiently recall the previous content without rereading it when they interrupt reading an e-book. Also, when trying to acquire knowledge during limited time such as commuting time or spare time, there are few means to review the previous reading content in a short time, so the problem is that efficient learning is hindered.
Means for Solving the Problems
[0005] This invention provides a system that records the point at which a user interrupts reading an e-book and uses AI to automatically generate a summary up to the specified point based on the recorded information. The generated summary is presented to the user when they resume reading, allowing them to quickly recall the content. The summary is provided in both text and audio output formats, allowing the user to review it in a way that suits their situation. Furthermore, by using a specific algorithm in the summary, it becomes possible to extract important information, thereby improving learning efficiency.
[0006] A "user" is a person who uses ebooks to read books.
[0007] An "eBook" is a book that is distributed in digital format and can be viewed through an electronic device.
[0008] "Interrupting reading" means that the user temporarily stops reading the ebook.
[0009] "Means for recording reading progress" refers to a mechanism for users to save the last section they have read on a digital device.
[0010] A "summary" is a way of compressing a long text or content, extracting the most important points, and presenting them in a short format.
[0011] "AI" refers to artificially constructed systems and technologies used for data analysis, pattern recognition, and other similar tasks.
[0012] An "algorithm" is a set of procedures or rules designed for computation or problem-solving.
[0013] "Text display" refers to the visual display of textual information on the screen of a digital device.
[0014] "Audio output" refers to a method of generating sound using a digital device and providing information to the user audibly.
[0015] "Extraction of important information" is a process of selecting information that is particularly relevant and useful from data.
Brief Explanation of Drawings
[0016] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.
Embodiments for Carrying Out the Invention
[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), etc.
[0020] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0021] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0024] [First Embodiment]
[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0037] This invention is a system for supporting users in efficiently resuming reading ebooks after interrupting their reading. This system primarily consists of a server, a terminal, and the user.
[0038] System Configuration
[0039] Device: This is the device that a user uses to read ebooks. The device records the current page and position while reading, and saves the position when the user interrupts reading.
[0040] Server: Receives data sent from the terminal during an interruption and uses the AI engine to summarize the text within the specified range. The server stores the generated summary in a database and sends the summary to the terminal when the user resumes.
[0041] Program operation description
[0042] 1. How to handle interruptions while reading
[0043] When a user stops reading an ebook, the device sends its location to the server. This location information includes page numbers and paragraph information.
[0044] 2. Summary generation
[0045] The server retrieves the relevant text from its database based on the received location information and generates a summary using an AI engine. The summary is adjusted to include the main points and the core of the argument.
[0046] 3. Summarize, save, and present
[0047] After the summary is generated, the server saves it to a database. Later, when the user resumes reading, the summary is provided upon request from the terminal.
[0048] When the terminal receives a summary from the server, it displays it as text and also presents it to the user as audio data.
[0049] Specific example
[0050] For example, suppose a user is reading an ebook on self-improvement. The user reads during their daily commute, but today they needed to stop before finishing Chapter 2, "Goal-Setting Techniques." The device sends this interruption point to the server, and the server uses AI to summarize the content up to that point. The next day, when the user resumes reading during their commute, the device provides a summary in both audio and text format, stating that "basic steps for setting goals were explained, and specific approaches to self-actualization were introduced," allowing them to quickly resume reading.
[0051] This system allows users to efficiently utilize their valuable time and learn without spending long hours.
[0052] The following describes the processing flow.
[0053] Step 1:
[0054] When a user starts reading an ebook, the device acquires that information and begins recording the current page and position. This allows the device to track which part of the book the user is reading in real time.
[0055] Step 2:
[0056] When a user pauses reading, they press the pause button on their device. The device detects this action and determines the location information of the last text read (page number, paragraph number, etc.).
[0057] Step 3:
[0058] The device sends confirmed location information to the server. This information includes the page number and paragraph number at the time of interruption.
[0059] Step 4:
[0060] Based on the received location information, the server retrieves the content of ebooks up to that location from its database.
[0061] Step 5:
[0062] The server activates the AI engine and generates a summary of the acquired text. The AI engine extracts key points and condenses them into a short text.
[0063] Step 6:
[0064] The server saves the generated summary to a database and associates it with the user's ID. This allows for quick access to the summary later.
[0065] Step 7:
[0066] When the user resumes reading, the device sends a request for a summary to the server.
[0067] Step 8:
[0068] The server retrieves the stored summary from the database and sends it back to the terminal.
[0069] Step 9:
[0070] The device displays the received summary to the user. It offers the user the option of displaying the summary as text or playing it as audio.
[0071] This process allows users to efficiently interrupt and resume reading, enabling them to learn without wasting time.
[0072] (Example 1)
[0073] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0074] When users are forced to interrupt their information consumption, a problem arises in that it is difficult for them to remember where they left off and to efficiently understand the content from where they left off when they resume reading. In particular, with information that is read over a long period of time, reviewing and confirming key points is a time-consuming and laborious task. This invention aims to solve the above problems and provide a means for quickly and efficiently resuming information consumption.
[0075] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0076] In this invention, the server includes means for recording the location where the user interrupted information consumption, a computing device for automatically summarizing relevant information based on the recorded data, means for saving the summary and presenting it upon resumption, and a processing device for generating the summary based on a generation model and prompt statements. This enables the user to efficiently understand what was interrupted and quickly resume information consumption.
[0077] A "user" is an individual or group that consumes and manipulates information and utilizes a system.
[0078] "Information consumption" refers to the act of using and understanding content such as ebooks.
[0079] "The location where the information was stopped" refers to the specific location where the user stopped consuming information.
[0080] "Means of recording" refers to technical elements or methods for saving the location where the operation was stopped as data.
[0081] "Data" refers to recordable content, including information about user behavior and location.
[0082] An "automatic summarizing computing device" refers to hardware or software that uses a specified algorithm or model to extract and simplify the key points of information.
[0083] A "summary" refers to a simplified version of the original information, extracting only the main points.
[0084] "Means of preservation" refers to technical elements or methods for storing summarized information so that it can be used later.
[0085] "Means of presentation" refers to technical elements or methods for providing the saved summary to the user visually or aurally.
[0086] A "generative model" refers to an algorithm or machine learning framework that processes information and creates summaries.
[0087] A "prompt statement" is an instructional statement input into a generative model, referring to text that provides conditions and directions for generating a summary.
[0088] This invention is a system that efficiently supports the interruption and resumption of electronic information. The system mainly consists of a terminal, a server, and a generative AI model, and when a user interrupts the consumption of electronic information, it records the location and presents a summary when it resumes.
[0089] The terminal has a function to automatically record the location where the user interrupted information consumption. Specifically, the terminal sends information such as page number and paragraph position to the server. Hardware used for this includes information input devices (e.g., touchscreen, keyboard).
[0090] The server retrieves relevant content from a database based on location information received from the terminal. It then uses a generative AI model to summarize the relevant text. The generative AI model has the capability to generate summaries using specified prompt sentences as input. The software used specifically includes a natural language processing engine.
[0091] The summary is stored in a database by the server and presented on the terminal when the user resumes consuming information. The terminal receives the summary from the server and presents it to the user visually (display) and aurally (speech synthesis). This allows the user to efficiently resume consuming information from where they left off.
[0092] For example, if a user interrupts their use of self-improvement digital content and tries to resume it the next day, the device receives a summary from the server stating that "basic steps were explained and specific approaches were introduced," and presents this summary to the user, allowing for a smooth resumption of information consumption.
[0093] An example of a prompt would be to input to the generative AI model in the format: "Summarize the following text and include the main points: '[Retrieved Text]'". This allows the generative AI model to efficiently extract the key points and provide a user-friendly summary.
[0094] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0095] Step 1:
[0096] The device detects actions that interrupt the user's consumption of information. Specifically, when the user closes an app or navigates to another page, the device records the current page number and paragraph location. The input is the user's action, and the output is the location information of the interruption.
[0097] Step 2:
[0098] The device sends the recorded location information to the server. The server receives this information and searches its database for related content. The input in this process is the location information sent from the device, and the output is the corresponding content text. Specifically, the server extracts text data from the database that is associated with the interrupted location.
[0099] Step 3:
[0100] The server inputs the extracted text into a generative AI model and generates a summary using a prompt. This prompt might be something like, "Summarize the following text, including the main points: '[Retrieved Text]'". The input is the text data and the prompt, and the output is the summary result. Specifically, the generative AI model uses natural language processing techniques to extract the key points and generate a shortened summary.
[0101] Step 4:
[0102] The server saves the generated summary to the database. In this step, the input is the generated summary, and the output is the saved summary data. Specifically, it records the summary along with related metadata (e.g., user ID, reading position) in the database.
[0103] Step 5:
[0104] When a user resumes consuming information, the device sends a request to the server for a summary. The server searches for the relevant summary and sends it to the device. The input is the request from the device, and the output is the summary data. Specifically, the server quickly searches for and returns summaries that have been stored in the past.
[0105] Step 6:
[0106] The terminal presents the received summary to the user visually and audibly. It displays the text on a screen and reads it aloud using a text-to-speech device. The input is summary data from the server, and the output is the presentation of information to the user. Specifically, the terminal renders the summary in text format and simultaneously provides audio output.
[0107] (Application Example 1)
[0108] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0109] In many information processing tasks, a problem arises where users cannot quickly grasp the workflow after temporarily interrupting and resuming their work, leading to decreased work efficiency. This problem is particularly pronounced in online shopping and e-books, hindering user convenience.
[0110] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0111] In this invention, the server includes means for recording the last operation when a user interrupts an information processing task on an information processing terminal, means for automatically summarizing the work performed up to that point based on the recorded information, and means for saving the summary and presenting it to the user when they resume their work. This allows the user to efficiently resume their work after interruption and to proceed smoothly with their tasks.
[0112] An "information processing terminal" is a device used by users to perform various information processing tasks, and specifically includes personal computers and smartphones.
[0113] "Operation point" refers to the location or positional information that indicates when a user interrupted a specific information processing task, and represents the user's operation history.
[0114] "Work content" refers to the set of data handled and processes performed when a user performs information processing tasks.
[0115] "Automatic summarization" is a process in which the system reconstructs the main information in a concise format based on the input information, and this is done by the system without human intervention.
[0116] "Information display means" refers to functions or devices that visually present necessary information to the user, and specifically includes displays and monitors.
[0117] "Sound generation means" refers to technologies and devices that convey information to users through sound, and includes speakers and speech synthesis software.
[0118] A "data processing algorithm" refers to a set of procedures or methods for manipulating, analyzing, and transforming data for a specific purpose. One example is AI-powered data analysis techniques.
[0119] This invention primarily uses three elements to realize a system: a server, an information processing terminal, and a user.
[0120] The server receives information about the operation location transmitted from the information processing terminal. For example, when a user interrupts product selection while online shopping, information about that product page is sent to the server.
[0121] Based on this information, the server automatically generates a summary using a generative AI model. This summary is tailored to include key features, pricing, and review summaries of the product, and utilizes specific data processing algorithms.
[0122] The generated summary is stored in the server's database and provided to the information processing terminal when the user resumes their work. The terminal uses information display means and sound generation means to present the user with a summary of the resumed position through text display and audio output.
[0123] This system allows users to efficiently resume interrupted tasks and ensure smooth workflow. For example, suppose a user was considering purchasing an electronic appliance online during their commute one day, but interrupted their browsing. The next day, upon resuming, they would be provided with a summary such as, "This product is highly rated for its quiet operation and is 10% cheaper than the market average," enabling them to make a quick purchase decision.
[0124] An example of a prompt used in the generative AI model is: "Briefly summarize the characteristics of the following product: Product name, Key features, Price, User reviews."
[0125] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0126] Step 1:
[0127] The user interrupts their work on the information processing terminal. For example, if the user closes a browser tab on a shopping site, the terminal records the page information at that time (product ID, page URL, timestamp, etc.) and sends it to the server. The input is page information, and the output is the transmission of information to the server.
[0128] Step 2:
[0129] The server analyzes the page information received from the terminal. This analysis uses the received product ID and URL to retrieve detailed product data (name, features, price, reviews, etc.) from the database. The input is page information, and the output is the retrieved product details.
[0130] Step 3:
[0131] The server uses a generative AI model to create a summary based on the retrieved product details. The prompt "Briefly summarize the characteristics of the following product: Product name, Key features, Price, User reviews" is used to input data into the AI model and generate the summary. The input is the product details, and the output is the generated summary.
[0132] Step 4:
[0133] The server saves the generated summary to a database, ready for the user to resume. In the saving step, in addition to the summary, the original page information is also recorded. The input is the generated summary, and the output is the database record of the summary.
[0134] Step 5:
[0135] When a user resumes work, the terminal requests a summary from the server. The server retrieves the corresponding summary from the database and sends it to the terminal. The input is the resume request, and the output is the terminal sending of the summary.
[0136] Step 6:
[0137] The terminal presents the received summary to the user. It displays the information as text using an information display device and plays it back as audio using an audio generation device. The input is summary data from the server, and the output is the presentation of the summary to the user.
[0138] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0139] This invention provides users with an efficient e-book reading experience, and in particular, it not only provides summaries up to the point of interruption, but also features a mechanism to provide customized summaries based on the user's emotions using an emotion engine. This system consists of a server, a terminal, and a user, and the addition of the emotion engine enables flexible information provision according to the emotional state of each individual user.
[0140] System Configuration
[0141] Device: A device that allows users to read ebooks and record their reading progress. The device is equipped with a camera and microphone, which are used to collect the user's voice and facial expressions for emotion recognition.
[0142] Server: Generates an ebook summary based on received data. Utilizing an AI engine, it adjusts the summary to provide the user with the most relevant information, taking into account user sentiment information provided by the sentiment engine.
[0143] Emotion Engine: This software analyzes the user's voice and facial expression data to determine their emotional state. The emotion engine provides information about the user's emotions to the server, which is then used to generate summaries.
[0144] Program operation description
[0145] 1. Acquiring emotional data during reading.
[0146] While the user is reading an e-book, the device acquires voice and facial information through its camera and microphone. This information is analyzed by an emotion engine, which identifies the user's emotional state in real time.
[0147] 2. How to handle interruptions while reading
[0148] When a user interrupts reading, the device sends that location to the server. At the same time, the latest sentiment analysis results are also sent to the server.
[0149] 3. Summary generation and customization
[0150] The server retrieves the content of the ebook up to a specified point and generates a summary using an AI engine. The generated summary is then adjusted in content and presentation style based on emotional information from an emotional engine.
[0151] 4. Presentation to the user
[0152] After generating the summary, the server sends it to the terminal in an appropriate format based on the user's emotions. For example, a bright tone of voice synthesis is used for positive emotional states, and the content is designed to better reflect those emotions.
[0153] Specific example
[0154] When a user reads a self-help book, the device analyzes the user's facial expressions and voice to confirm that the user is relaxed. If the user interrupts reading and resumes reading the next day, the server presents a summary in a calm tone of voice that takes the user's relaxed state into account, providing a peaceful reading experience. In this way, the system dynamically adjusts content based on the user's emotions, providing even more personalized reading support.
[0155] The following describes the processing flow.
[0156] Step 1:
[0157] When a user begins reading an ebook, the device retrieves the ebook's content and user information, and records the start of reading. The device then uses its camera and microphone to begin capturing the user's voice and facial expressions in real time.
[0158] Step 2:
[0159] The device sends the acquired voice and facial expression data to the emotion engine, which analyzes the emotional state. The emotion engine identifies the user's emotions and returns that information to the device. The analysis is performed periodically to create a user emotion log.
[0160] Step 3:
[0161] When a user attempts to interrupt reading, they press the pause button on their device. The device then sends the current page information and the latest sentiment log data to the server.
[0162] Step 4:
[0163] The server retrieves the relevant text from the database based on the received page information. The server then activates an AI engine, analyzes the content up to a specified point, and generates a summary. This summary includes customization based on the user's sentiment data.
[0164] Step 5:
[0165] The server stores the generated summary in a database and associates it with the user's ID. The summary also retains corresponding sentiment data as metadata.
[0166] Step 6:
[0167] When the user operates the device to resume reading, the device sends a request for a summary to the server.
[0168] Step 7:
[0169] The server returns the stored summary and its associated sentiment information to the terminal.
[0170] Step 8:
[0171] The device presents the received summary to the user. When presenting the summary, it reflects the user's previous emotional state; for example, if the user is in a calm emotional state, the tone of the synthesized speech is adjusted to be gentler. The user can also choose to view the summary as text, and the device automatically selects a font and background color appropriate to the user's emotion.
[0172] This process presents users with summaries that match their emotions, further enhancing their reading experience.
[0173] (Example 2)
[0174] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0175] Conventional electronic media viewing systems simply record where the user paused and provide a summary up to that point. Therefore, they are unable to provide individual summaries that take into account the user's emotional state, highlighting the need for more personalized information delivery. Furthermore, while adjusting summaries based on emotions can improve the user experience, such a mechanism has not existed.
[0176] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0177] In this invention, the server includes means for recording the last interruption point when a user interrupts the content of an electronic medium; means for automatically summarizing the content up to that point based on the recorded information; means for identifying the user's emotional state by analyzing voice and facial data; means for adjusting the summary content based on the identified emotional state; and means for saving the summary and presenting it when the user resumes viewing. This makes it possible to provide information in a form that is appropriate to the user's emotions.
[0178] A "user" refers to an individual or group that obtains information using electronic media.
[0179] "Electronic media" refers to media that can provide information in digital format, including ebooks and digital documents.
[0180] "Interruption point" refers to information indicating the point at which the user temporarily stopped using the electronic medium.
[0181] A "summary" refers to information that concisely summarizes the content of an electronic medium and is provided in a format that is easy for the user to understand.
[0182] "Emotional state" refers to a temporary psychological state identified by analyzing the user's voice and facial data.
[0183] A "server" is a part of a system that provides services to user terminals through data processing and storage functions.
[0184] "Voice and facial data" refers to digital information including the user's tone of voice and facial expressions, which is used to identify their emotional state.
[0185] "Adjustment" refers to changing the content and presentation method of a summary based on the user's identified emotional state.
[0186] "Means" refers to a method or apparatus for achieving a specific function.
[0187] This system consists of users, terminals, and servers, and provides users with an efficient and personalized browsing experience of electronic media. Specifically, it is implemented as follows:
[0188] Users use a terminal when browsing electronic media. This terminal is a digital device with a built-in camera and microphone. This allows the terminal to capture data on the user's voice and facial expressions. This digital data is then transferred to an emotion engine to analyze the user's emotional state.
[0189] The server receives information about the interruption point and emotional data from the terminal. Based on this information, the server uses a generative AI model to summarize the content of the electronic medium. At this time, the summary is adjusted according to the user's emotional state, which is analyzed by the emotional engine. For example, the content is adjusted to be presented in a calm tone to a relaxed user.
[0190] As a concrete example, if a user is in a relaxed state while reading a self-help book, the system will use that emotional state to present a summary in a calm tone via speech synthesis when they resume reading the book the next day, providing a peaceful reading experience.
[0191] An example of a prompt message is, "Generate an electronic summary based on the user's current emotional state. The user is relaxed." In this way, the system can provide meaningful information in real time that corresponds to the user's specific emotions.
[0192] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0193] Step 1:
[0194] The device uses its built-in camera and microphone to collect the user's voice and facial expressions while they are browsing electronic media. The collected data is processed in real time within the device to extract attributes necessary for identifying the user's emotional state (e.g., voice tone and facial expression features). This serves as the input, and the output generates the dataset necessary for sentiment analysis.
[0195] Step 2:
[0196] The device sends the collected dataset to the emotion engine. The emotion engine uses an AI model to analyze this dataset and identify the user's current emotional state. In this step, voice tone and facial features are used as input, and the user's emotional state (e.g., relaxed, stressed) is output.
[0197] Step 3:
[0198] When a user interrupts their viewing of electronic media, the device sends information about the interruption point and their current emotional state to the server. This input—the interruption point and emotional state—initiates the server-side summary generation process.
[0199] Step 4:
[0200] The server retrieves the content of the electronic media up to the point of interruption. It then uses an AI engine to invoke a generative AI model, which generates a summary based on the input electronic media content and emotional state. The generative AI model analyzes relevant information to generate a summary, adjusting it in the process to consider the user's emotional state. In this step, the electronic media content and emotional state are the inputs, and a summary tailored to the user's emotions is output.
[0201] Step 5:
[0202] The server sends the generated summary to the terminal. The terminal displays the received summary on its screen and then reads the summary aloud using speech synthesis. Here, the adjusted summary becomes the input and is output as specific visual and auditory information presented to the user.
[0203] (Application Example 2)
[0204] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0205] Modern e-readers lack personalized summaries that take into account the user's emotional state when reading is interrupted, which hinders reading continuity and immersion. Furthermore, to ensure a comfortable reading experience upon resuming, summaries need to be presented not just in a simple format, but in a style that matches the user's emotional state.
[0206] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0207] In this invention, the server includes means for recording the last reading position and emotional state when reading an ebook is interrupted, means for automatically summarizing the content of the ebook up to that point based on the recorded information and the user's emotional state, and adjusting the presentation style, and means for saving the summary and presenting it as an emotionally-based summary when the user resumes reading. This makes it possible to provide the user with personalized reading support that matches their emotional state.
[0208] A "user" is an individual who reads ebooks and is the subject who utilizes the system.
[0209] An "eBook" is a book provided in digital format and is content that can be read on an electronic device.
[0210] "Interrupting reading" refers to the act of a user temporarily stopping the viewing of an e-book.
[0211] "Last read position" refers to the last part of the ebook that the user read before interrupting their reading.
[0212] "Emotional state" refers to information that represents the user's current emotions, and is usually obtained from voice or facial expression data.
[0213] "To record" refers to the act of storing information in a database or storage device in order to preserve specific data.
[0214] A "summary" is a shortened version of an e-book, containing only the key points.
[0215] "Presentation style" refers to the format and method of providing information to users.
[0216] "Saving" refers to maintaining data for later reuse.
[0217] "Means" refers to the methods or devices used to achieve a specific purpose.
[0218] "Personalized" means that it is delivered in a way that is optimized for each individual user.
[0219] This invention is a system for providing a personalized reading experience to users of ebooks. Its embodiments are described below.
[0220] The system consists of a user terminal, a central server, and an emotion engine that performs sentiment analysis. While the user is reading an ebook, the terminal uses its camera and microphone to capture voice and facial expressions in real time. The captured data is analyzed by the emotion engine to identify the user's current emotional state. This sentiment information is simultaneously transmitted to the server.
[0221] When a user pauses reading, the device records the point of interruption and their emotional state, and sends this information to a server. Based on this recorded information, the server uses an AI model to create a summary of the ebook up to the point of interruption. During this process, the content and presentation style of the summary are adjusted based on the emotional information provided by the emotion engine. For example, if the user is in a relaxed state, the summary may be presented in a calm tone of voice, employing an emotion-based approach.
[0222] The generated summary is provided to the device as audio data generated using speech synthesis technology, as well as as text. A modern speech synthesis engine is used to improve speech quality.
[0223] For example, if the system detects that a user is smiling while reading an emotionally moving novel, a summary will be provided in a warm tone of voice when they resume reading the next day, reminding them of their previous reading experience. An example of a prompt for the generative AI model would be: "Generate a summary of the ebook up to where the user left off. The user's emotion is 'relaxed,' so please use a summary style that matches that emotion. In particular, please insert words that evoke a sense of relaxation at the beginning of the summary."
[0224] In this way, the present invention realizes a personalized reading experience based on the user's emotions.
[0225] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0226] Step 1:
[0227] While the user is reading an e-book, the device uses its built-in camera and microphone to acquire audio and facial expression data. This input data is sent to an emotion engine, which analyzes the data to identify the user's emotional state in real time.
[0228] Step 2:
[0229] When a user interrupts reading, the device records the point of interruption and the emotional state provided by the emotion engine. This information, including emotional data, is sent to the server, allowing the server to gain a comprehensive understanding of the interruption.
[0230] Step 3:
[0231] The server, based on the received interruption point and emotional information, forms a prompt sentence for the AI engine and sends it to the generative AI model. The generative AI model uses this prompt sentence to summarize the content of the ebook up to the interruption point. This summary is adjusted according to the emotional state.
[0232] Step 4:
[0233] The generated summary is converted into an audio file using a speech synthesis engine. The speech synthesis engine operates with a tone and speed set based on emotional information. The output audio file and text summary are sent to the terminal.
[0234] Step 5:
[0235] When the user resumes reading, the device presents the received summary in an emotionally appropriate style. The user can choose between audio playback or text display, providing a personalized reading experience. This step ensures the user receives information that matches their current emotional state.
[0236] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0237] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0238] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0239] [Second Embodiment]
[0240] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0241] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0242] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0243] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0244] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0245] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0246] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0247] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0248] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0249] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0250] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0251] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0252] This invention is a system for supporting users in efficiently resuming reading ebooks after interrupting their reading. This system primarily consists of a server, a terminal, and the user.
[0253] System Configuration
[0254] Device: This is the device that a user uses to read ebooks. The device records the current page and position while reading, and saves the position when the user interrupts reading.
[0255] Server: Receives data sent from the terminal during an interruption and uses the AI engine to summarize the text within the specified range. The server stores the generated summary in a database and sends the summary to the terminal when the user resumes.
[0256] Program operation description
[0257] 1. How to handle interruptions while reading
[0258] When a user stops reading an ebook, the device sends its location to the server. This location information includes page numbers and paragraph information.
[0259] 2. Summary generation
[0260] The server retrieves the relevant text from its database based on the received location information and generates a summary using an AI engine. The summary is adjusted to include the main points and the core of the argument.
[0261] 3. Summarize, save, and present
[0262] After the summary is generated, the server saves it to a database. Later, when the user resumes reading, the summary is provided upon request from the terminal.
[0263] When the terminal receives a summary from the server, it displays it as text and also presents it to the user as audio data.
[0264] Specific example
[0265] For example, suppose a user is reading an ebook on self-improvement. The user reads during their daily commute, but today they needed to stop before finishing Chapter 2, "Goal-Setting Techniques." The device sends this interruption point to the server, and the server uses AI to summarize the content up to that point. The next day, when the user resumes reading during their commute, the device provides a summary in both audio and text format, stating that "basic steps for setting goals were explained, and specific approaches to self-actualization were introduced," allowing them to quickly resume reading.
[0266] This system allows users to efficiently utilize their valuable time and learn without spending long hours.
[0267] The following describes the processing flow.
[0268] Step 1:
[0269] When a user starts reading an ebook, the device acquires that information and begins recording the current page and position. This allows the device to track which part of the book the user is reading in real time.
[0270] Step 2:
[0271] When a user pauses reading, they press the pause button on their device. The device detects this action and determines the location information of the last text read (page number, paragraph number, etc.).
[0272] Step 3:
[0273] The device sends confirmed location information to the server. This information includes the page number and paragraph number at the time of interruption.
[0274] Step 4:
[0275] Based on the received location information, the server retrieves the content of ebooks up to that location from its database.
[0276] Step 5:
[0277] The server starts the AI engine and generates a summary of the acquired text. The AI engine performs a process of extracting important points and summarizing them into a short text.
[0278] Step 6:
[0279] The server saves the generated summary in the database and associates it with the user's ID. This enables the summary to be quickly accessed later.
[0280] Step 7:
[0281] When the user resumes reading, the terminal sends a request for the summary to the server.
[0282] Step 8:
[0283] The server searches the database for the saved summary and sends it back to the terminal.
[0284] Step 9:
[0285] The terminal presents the received summary to the user. At this time, options for text display and audio playback are provided so that the user can make a selection.
[0286] Through this series of processes, the user can efficiently interrupt and resume reading and learn without wasting time.
[0287] (Example 1)
[0288] Next, Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0289] When users are forced to interrupt their information consumption, a problem arises in that it is difficult for them to remember where they left off and to efficiently understand the content from where they left off when they resume reading. In particular, with information that is read over a long period of time, reviewing and confirming key points is a time-consuming and laborious task. This invention aims to solve the above problems and provide a means for quickly and efficiently resuming information consumption.
[0290] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0291] In this invention, the server includes means for recording the location where the user interrupted information consumption, a computing device for automatically summarizing relevant information based on the recorded data, means for saving the summary and presenting it upon resumption, and a processing device for generating the summary based on a generation model and prompt statements. This enables the user to efficiently understand what was interrupted and quickly resume information consumption.
[0292] A "user" is an individual or group that consumes and manipulates information and utilizes a system.
[0293] "Information consumption" refers to the act of using and understanding content such as ebooks.
[0294] "The location where the information was stopped" refers to the specific location where the user stopped consuming information.
[0295] "Means of recording" refers to technical elements or methods for saving the location where the operation was stopped as data.
[0296] "Data" refers to recordable content, including information about user behavior and location.
[0297] An "automatic summarizing computing device" refers to hardware or software that uses a specified algorithm or model to extract and simplify the key points of information.
[0298] A "summary" refers to a simplified version of the original information, extracting only the main points.
[0299] "Means of preservation" refers to technical elements or methods for storing summarized information so that it can be used later.
[0300] "Means of presentation" refers to technical elements or methods for providing the saved summary to the user visually or aurally.
[0301] A "generative model" refers to an algorithm or machine learning framework that processes information and creates summaries.
[0302] A "prompt statement" is an instructional statement input into a generative model, referring to text that provides conditions and directions for generating a summary.
[0303] This invention is a system that efficiently supports the interruption and resumption of electronic information. The system mainly consists of a terminal, a server, and a generative AI model, and when a user interrupts the consumption of electronic information, it records the location and presents a summary when it resumes.
[0304] The terminal has a function to automatically record the location where the user interrupted information consumption. Specifically, the terminal sends information such as page number and paragraph position to the server. Hardware used for this includes information input devices (e.g., touchscreen, keyboard).
[0305] The server retrieves relevant content from a database based on location information received from the terminal. It then uses a generative AI model to summarize the relevant text. The generative AI model has the capability to generate summaries using specified prompt sentences as input. The software used specifically includes a natural language processing engine.
[0306] The summary is saved in the database by the server and presented on the terminal when the user resumes consuming information. The terminal receives the summary from the server and presents it to the user visually (display) and auditorily (voice synthesizer). As a result, the user can efficiently resume consuming information from the interruption point.
[0307] As a specific example, when the user interrupts while using self-motivated electronic content and tries to resume it the next day, the terminal receives from the server a summary that "basic steps are explained and specific approaches are introduced" and presents it to the user, enabling a smooth resumption of information consumption.
[0308] As an example of the prompt sentence, it is input in the form of "Summarize the following text and include the main points: '[Retrieved text]'" to the generative AI model. In this way, the generative AI model can efficiently extract the key points and provide a summary that is easy for the user to understand.
[0309] The flow of the specific process in Example 1 will be described using FIG. 11.
[0310] Step 1:
[0311] The terminal detects an operation by the user to interrupt information consumption. Specifically, when the user closes the app or performs a page movement operation, the terminal records the current page number and paragraph position information. The input is the user's operation, and the output is the position information of the interruption point.
[0312] Step 2:
[0313] The terminal sends the recorded position information to the server. The server receives the information and searches the database for relevant content. The input at this time is the position information sent from the terminal, and the output is the corresponding content text. Specifically, the server performs an operation to extract text data associated with the interrupted position from the database.
[0314] Step 3:
[0315] The server inputs the extracted text into a generative AI model and generates a summary using a prompt. This prompt might be something like, "Summarize the following text, including the main points: '[Retrieved Text]'". The input is the text data and the prompt, and the output is the summary result. Specifically, the generative AI model uses natural language processing techniques to extract the key points and generate a shortened summary.
[0316] Step 4:
[0317] The server saves the generated summary to the database. In this step, the input is the generated summary, and the output is the saved summary data. Specifically, it records the summary along with related metadata (e.g., user ID, reading position) in the database.
[0318] Step 5:
[0319] When a user resumes consuming information, the device sends a request to the server for a summary. The server searches for the relevant summary and sends it to the device. The input is the request from the device, and the output is the summary data. Specifically, the server quickly searches for and returns summaries that have been stored in the past.
[0320] Step 6:
[0321] The terminal presents the received summary to the user visually and audibly. It displays the text on a screen and reads it aloud using a text-to-speech device. The input is summary data from the server, and the output is the presentation of information to the user. Specifically, the terminal renders the summary in text format and simultaneously provides audio output.
[0322] (Application Example 1)
[0323] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0324] In many information processing tasks, a problem arises where users cannot quickly grasp the workflow after temporarily interrupting and resuming their work, leading to decreased work efficiency. This problem is particularly pronounced in online shopping and e-books, hindering user convenience.
[0325] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0326] In this invention, the server includes means for recording the last operation when a user interrupts an information processing task on an information processing terminal, means for automatically summarizing the work performed up to that point based on the recorded information, and means for saving the summary and presenting it to the user when they resume their work. This allows the user to efficiently resume their work after interruption and to proceed smoothly with their tasks.
[0327] An "information processing terminal" is a device used by users to perform various information processing tasks, and specifically includes personal computers and smartphones.
[0328] "Operation point" refers to the location or positional information that indicates when a user interrupted a specific information processing task, and represents the user's operation history.
[0329] "Work content" refers to the set of data handled and processes performed when a user performs information processing tasks.
[0330] "Automatic summarization" is a process in which the system reconstructs the main information in a concise format based on the input information, and this is done by the system without human intervention.
[0331] "Information display means" refers to functions or devices that visually present necessary information to the user, and specifically includes displays and monitors.
[0332] "Sound generation means" refers to technologies and devices that convey information to users through sound, and includes speakers and speech synthesis software.
[0333] A "data processing algorithm" refers to a set of procedures or methods for manipulating, analyzing, and transforming data for a specific purpose. One example is AI-powered data analysis techniques.
[0334] This invention primarily uses three elements to realize a system: a server, an information processing terminal, and a user.
[0335] The server receives information about the operation location transmitted from the information processing terminal. For example, when a user interrupts product selection while online shopping, information about that product page is sent to the server.
[0336] Based on this information, the server automatically generates a summary using a generative AI model. This summary is tailored to include key features, pricing, and review summaries of the product, and utilizes specific data processing algorithms.
[0337] The generated summary is stored in the server's database and provided to the information processing terminal when the user resumes their work. The terminal uses information display means and sound generation means to present the user with a summary of the resumed position through text display and audio output.
[0338] This system allows users to efficiently resume interrupted tasks and ensure smooth workflow. For example, suppose a user was considering purchasing an electronic appliance online during their commute one day, but interrupted their browsing. The next day, upon resuming, they would be provided with a summary such as, "This product is highly rated for its quiet operation and is 10% cheaper than the market average," enabling them to make a quick purchase decision.
[0339] An example of a prompt used in the generative AI model is: "Briefly summarize the characteristics of the following product: Product name, Key features, Price, User reviews."
[0340] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0341] Step 1:
[0342] The user interrupts their work on the information processing terminal. For example, if the user closes a browser tab on a shopping site, the terminal records the page information at that time (product ID, page URL, timestamp, etc.) and sends it to the server. The input is page information, and the output is the transmission of information to the server.
[0343] Step 2:
[0344] The server analyzes the page information received from the terminal. This analysis uses the received product ID and URL to retrieve detailed product data (name, features, price, reviews, etc.) from the database. The input is page information, and the output is the retrieved product details.
[0345] Step 3:
[0346] The server uses a generative AI model to create a summary based on the retrieved product details. The prompt "Briefly summarize the characteristics of the following product: Product name, Key features, Price, User reviews" is used to input data into the AI model and generate the summary. The input is the product details, and the output is the generated summary.
[0347] Step 4:
[0348] The server saves the generated summary to a database, ready for the user to resume. In the saving step, in addition to the summary, the original page information is also recorded. The input is the generated summary, and the output is the database record of the summary.
[0349] Step 5:
[0350] When a user resumes work, the terminal requests a summary from the server. The server retrieves the corresponding summary from the database and sends it to the terminal. The input is the resume request, and the output is the terminal sending of the summary.
[0351] Step 6:
[0352] The terminal presents the received summary to the user. It displays the information as text using an information display device and plays it back as audio using an audio generation device. The input is summary data from the server, and the output is the presentation of the summary to the user.
[0353] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0354] This invention provides users with an efficient e-book reading experience, and in particular, it not only provides summaries up to the point of interruption, but also features a mechanism to provide customized summaries based on the user's emotions using an emotion engine. This system consists of a server, a terminal, and a user, and the addition of the emotion engine enables flexible information provision according to the emotional state of each individual user.
[0355] System Configuration
[0356] Device: A device that allows users to read ebooks and record their reading progress. The device is equipped with a camera and microphone, which are used to collect the user's voice and facial expressions for emotion recognition.
[0357] Server: Generates an ebook summary based on received data. Utilizing an AI engine, it adjusts the summary to provide the user with the most relevant information, taking into account user sentiment information provided by the sentiment engine.
[0358] Emotion Engine: This software analyzes the user's voice and facial expression data to determine their emotional state. The emotion engine provides information about the user's emotions to the server, which is then used to generate summaries.
[0359] Program operation description
[0360] 1. Acquiring emotional data during reading.
[0361] While the user is reading an e-book, the device acquires voice and facial information through its camera and microphone. This information is analyzed by an emotion engine, which identifies the user's emotional state in real time.
[0362] 2. How to handle interruptions while reading
[0363] When a user interrupts reading, the device sends that location to the server. At the same time, the latest sentiment analysis results are also sent to the server.
[0364] 3. Summary generation and customization
[0365] The server retrieves the content of the ebook up to a specified point and generates a summary using an AI engine. The generated summary is then adjusted in content and presentation style based on emotional information from an emotional engine.
[0366] 4. Presentation to the user
[0367] After generating the summary, the server sends it to the terminal in an appropriate format based on the user's emotions. For example, a bright tone of voice synthesis is used for positive emotional states, and the content is designed to better reflect those emotions.
[0368] Specific example
[0369] When a user reads a self-help book, the device analyzes the user's facial expressions and voice to confirm that the user is relaxed. If the user interrupts reading and resumes reading the next day, the server presents a summary in a calm tone of voice that takes the user's relaxed state into account, providing a peaceful reading experience. In this way, the system dynamically adjusts content based on the user's emotions, providing even more personalized reading support.
[0370] The following describes the processing flow.
[0371] Step 1:
[0372] When a user begins reading an ebook, the device retrieves the ebook's content and user information, and records the start of reading. The device then uses its camera and microphone to begin capturing the user's voice and facial expressions in real time.
[0373] Step 2:
[0374] The device sends the acquired voice and facial expression data to the emotion engine, which analyzes the emotional state. The emotion engine identifies the user's emotions and returns that information to the device. The analysis is performed periodically to create a user emotion log.
[0375] Step 3:
[0376] When a user attempts to interrupt reading, they press the pause button on their device. The device then sends the current page information and the latest sentiment log data to the server.
[0377] Step 4:
[0378] The server retrieves the relevant text from the database based on the received page information. The server then activates an AI engine, analyzes the content up to a specified point, and generates a summary. This summary includes customization based on the user's sentiment data.
[0379] Step 5:
[0380] The server stores the generated summary in a database and associates it with the user's ID. The summary also retains corresponding sentiment data as metadata.
[0381] Step 6:
[0382] When the user operates the device to resume reading, the device sends a request for a summary to the server.
[0383] Step 7:
[0384] The server returns the stored summary and its associated sentiment information to the terminal.
[0385] Step 8:
[0386] The device presents the received summary to the user. When presenting the summary, it reflects the user's previous emotional state; for example, if the user is in a calm emotional state, the tone of the synthesized speech is adjusted to be gentler. The user can also choose to view the summary as text, and the device automatically selects a font and background color appropriate to the user's emotion.
[0387] This process presents users with summaries that match their emotions, further enhancing their reading experience.
[0388] (Example 2)
[0389] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0390] Conventional electronic media viewing systems simply record where the user paused and provide a summary up to that point. Therefore, they are unable to provide individual summaries that take into account the user's emotional state, highlighting the need for more personalized information delivery. Furthermore, while adjusting summaries based on emotions can improve the user experience, such a mechanism has not existed.
[0391] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0392] In this invention, the server includes means for recording the last interruption point when a user interrupts the content of an electronic medium; means for automatically summarizing the content up to that point based on the recorded information; means for identifying the user's emotional state by analyzing voice and facial data; means for adjusting the summary content based on the identified emotional state; and means for saving the summary and presenting it when the user resumes viewing. This makes it possible to provide information in a form that is appropriate to the user's emotions.
[0393] A "user" refers to an individual or group that obtains information using electronic media.
[0394] "Electronic media" refers to media that can provide information in digital format, including ebooks and digital documents.
[0395] "Interruption point" refers to information indicating the point at which the user temporarily stopped using the electronic medium.
[0396] A "summary" refers to information that concisely summarizes the content of an electronic medium and is provided in a format that is easy for the user to understand.
[0397] "Emotional state" refers to a temporary psychological state identified by analyzing the user's voice and facial data.
[0398] A "server" is a part of a system that provides services to user terminals through data processing and storage functions.
[0399] "Voice and facial data" refers to digital information including the user's tone of voice and facial expressions, which is used to identify their emotional state.
[0400] "Adjustment" refers to changing the content and presentation method of a summary based on the user's identified emotional state.
[0401] "Means" refers to a method or apparatus for achieving a specific function.
[0402] This system consists of users, terminals, and servers, and provides users with an efficient and personalized browsing experience of electronic media. Specifically, it is implemented as follows:
[0403] Users use a terminal when browsing electronic media. This terminal is a digital device with a built-in camera and microphone. This allows the terminal to capture data on the user's voice and facial expressions. This digital data is then transferred to an emotion engine to analyze the user's emotional state.
[0404] The server receives information about the interruption point and emotional data from the terminal. Based on this information, the server uses a generative AI model to summarize the content of the electronic medium. At this time, the summary is adjusted according to the user's emotional state, which is analyzed by the emotional engine. For example, the content is adjusted to be presented in a calm tone to a relaxed user.
[0405] As a concrete example, if a user is in a relaxed state while reading a self-help book, the system will use that emotional state to present a summary in a calm tone via speech synthesis when they resume reading the book the next day, providing a peaceful reading experience.
[0406] An example of a prompt message is, "Generate an electronic summary based on the user's current emotional state. The user is relaxed." In this way, the system can provide meaningful information in real time that corresponds to the user's specific emotions.
[0407] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0408] Step 1:
[0409] The device uses its built-in camera and microphone to collect the user's voice and facial expressions while they are browsing electronic media. The collected data is processed in real time within the device to extract attributes necessary for identifying the user's emotional state (e.g., voice tone and facial expression features). This serves as the input, and the output generates the dataset necessary for sentiment analysis.
[0410] Step 2:
[0411] The device sends the collected dataset to the emotion engine. The emotion engine uses an AI model to analyze this dataset and identify the user's current emotional state. In this step, voice tone and facial features are used as input, and the user's emotional state (e.g., relaxed, stressed) is output.
[0412] Step 3:
[0413] When a user interrupts their viewing of electronic media, the device sends information about the interruption point and their current emotional state to the server. This input—the interruption point and emotional state—initiates the server-side summary generation process.
[0414] Step 4:
[0415] The server retrieves the content of the electronic media up to the point of interruption. It then uses an AI engine to invoke a generative AI model, which generates a summary based on the input electronic media content and emotional state. The generative AI model analyzes relevant information to generate a summary, adjusting it in the process to consider the user's emotional state. In this step, the electronic media content and emotional state are the inputs, and a summary tailored to the user's emotions is output.
[0416] Step 5:
[0417] The server sends the generated summary to the terminal. The terminal displays the received summary on its screen and then reads the summary aloud using speech synthesis. Here, the adjusted summary becomes the input and is output as specific visual and auditory information presented to the user.
[0418] (Application Example 2)
[0419] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".
[0420] Modern e-readers lack personalized summaries that take into account the user's emotional state when reading is interrupted, which hinders reading continuity and immersion. Furthermore, to ensure a comfortable reading experience upon resuming, summaries need to be presented not just in a simple format, but in a style that matches the user's emotional state.
[0421] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0422] In this invention, the server includes means for recording the last reading position and emotional state when reading an ebook is interrupted, means for automatically summarizing the content of the ebook up to that point based on the recorded information and the user's emotional state, and adjusting the presentation style, and means for saving the summary and presenting it as an emotionally-based summary when the user resumes reading. This makes it possible to provide the user with personalized reading support that matches their emotional state.
[0423] A "user" is an individual who reads ebooks and is the subject who utilizes the system.
[0424] An "eBook" is a book provided in digital format and is content that can be read on an electronic device.
[0425] "Interrupting reading" refers to the act of a user temporarily stopping the viewing of an e-book.
[0426] "Last read position" refers to the last part of the ebook that the user read before interrupting their reading.
[0427] "Emotional state" refers to information that represents the user's current emotions, and is usually obtained from voice or facial expression data.
[0428] "To record" refers to the act of storing information in a database or storage device in order to preserve specific data.
[0429] A "summary" is a shortened version of an e-book, containing only the key points.
[0430] "Presentation style" refers to the format and method of providing information to users.
[0431] "Saving" refers to maintaining data for later reuse.
[0432] "Means" refers to the methods or devices used to achieve a specific purpose.
[0433] "Personalized" means that it is delivered in a way that is optimized for each individual user.
[0434] This invention is a system for providing a personalized reading experience to users of ebooks. Its embodiments are described below.
[0435] The system consists of a user terminal, a central server, and an emotion engine that performs sentiment analysis. While the user is reading an ebook, the terminal uses its camera and microphone to capture voice and facial expressions in real time. The captured data is analyzed by the emotion engine to identify the user's current emotional state. This sentiment information is simultaneously transmitted to the server.
[0436] When a user pauses reading, the device records the point of interruption and their emotional state, and sends this information to a server. Based on this recorded information, the server uses an AI model to create a summary of the ebook up to the point of interruption. During this process, the content and presentation style of the summary are adjusted based on the emotional information provided by the emotion engine. For example, if the user is in a relaxed state, the summary may be presented in a calm tone of voice, employing an emotion-based approach.
[0437] The generated summary is provided to the device as audio data generated using speech synthesis technology, as well as as text. A modern speech synthesis engine is used to improve speech quality.
[0438] For example, if the system detects that a user is smiling while reading an emotionally moving novel, a summary will be provided in a warm tone of voice when they resume reading the next day, reminding them of their previous reading experience. An example of a prompt for the generative AI model would be: "Generate a summary of the ebook up to where the user left off. The user's emotion is 'relaxed,' so please use a summary style that matches that emotion. In particular, please insert words that evoke a sense of relaxation at the beginning of the summary."
[0439] In this way, the present invention realizes a personalized reading experience based on the user's emotions.
[0440] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0441] Step 1:
[0442] While the user is reading an e-book, the device uses its built-in camera and microphone to acquire audio and facial expression data. This input data is sent to an emotion engine, which analyzes the data to identify the user's emotional state in real time.
[0443] Step 2:
[0444] When a user interrupts reading, the device records the point of interruption and the emotional state provided by the emotion engine. This information, including emotional data, is sent to the server, allowing the server to gain a comprehensive understanding of the interruption.
[0445] Step 3:
[0446] The server, based on the received interruption point and emotional information, forms a prompt sentence for the AI engine and sends it to the generative AI model. The generative AI model uses this prompt sentence to summarize the content of the ebook up to the interruption point. This summary is adjusted according to the emotional state.
[0447] Step 4:
[0448] The generated summary is converted into an audio file using a speech synthesis engine. The speech synthesis engine operates with a tone and speed set based on emotional information. The output audio file and text summary are sent to the terminal.
[0449] Step 5:
[0450] When the user resumes reading, the device presents the received summary in an emotionally appropriate style. The user can choose between audio playback or text display, providing a personalized reading experience. This step ensures the user receives information that matches their current emotional state.
[0451] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0452] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0453] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0454] [Third Embodiment]
[0455] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0456] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0457] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0458] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0459] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0460] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0461] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0462] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0463] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0464] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0465] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0466] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0467] This invention is a system for supporting users in efficiently resuming reading ebooks after interrupting their reading. This system primarily consists of a server, a terminal, and the user.
[0468] System Configuration
[0469] Device: This is the device that a user uses to read ebooks. The device records the current page and position while reading, and saves the position when the user interrupts reading.
[0470] Server: Receives data sent from the terminal during an interruption and uses the AI engine to summarize the text within the specified range. The server stores the generated summary in a database and sends the summary to the terminal when the user resumes.
[0471] Program operation description
[0472] 1. How to handle interruptions while reading
[0473] When a user stops reading an ebook, the device sends its location to the server. This location information includes page numbers and paragraph information.
[0474] 2. Summary generation
[0475] The server retrieves the relevant text from its database based on the received location information and generates a summary using an AI engine. The summary is adjusted to include the main points and the core of the argument.
[0476] 3. Summarize, save, and present
[0477] After the summary is generated, the server saves it to a database. Later, when the user resumes reading, the summary is provided upon request from the terminal.
[0478] When the terminal receives a summary from the server, it displays it as text and also presents it to the user as audio data.
[0479] Specific example
[0480] For example, suppose a user is reading an ebook on self-improvement. The user reads during their daily commute, but today they needed to stop before finishing Chapter 2, "Goal-Setting Techniques." The device sends this interruption point to the server, and the server uses AI to summarize the content up to that point. The next day, when the user resumes reading during their commute, the device provides a summary in both audio and text format, stating that "basic steps for setting goals were explained, and specific approaches to self-actualization were introduced," allowing them to quickly resume reading.
[0481] This system allows users to efficiently utilize their valuable time and learn without spending long hours.
[0482] The following describes the processing flow.
[0483] Step 1:
[0484] When a user starts reading an ebook, the device acquires that information and begins recording the current page and position. This allows the device to track which part of the book the user is reading in real time.
[0485] Step 2:
[0486] When a user pauses reading, they press the pause button on their device. The device detects this action and determines the location information of the last text read (page number, paragraph number, etc.).
[0487] Step 3:
[0488] The device sends confirmed location information to the server. This information includes the page number and paragraph number at the time of interruption.
[0489] Step 4:
[0490] Based on the received location information, the server retrieves the content of ebooks up to that location from its database.
[0491] Step 5:
[0492] The server activates the AI engine and generates a summary of the acquired text. The AI engine extracts key points and condenses them into a short text.
[0493] Step 6:
[0494] The server saves the generated summary to a database and associates it with the user's ID. This allows for quick access to the summary later.
[0495] Step 7:
[0496] When the user resumes reading, the device sends a request for a summary to the server.
[0497] Step 8:
[0498] The server retrieves the stored summary from the database and sends it back to the terminal.
[0499] Step 9:
[0500] The device displays the received summary to the user. It offers the user the option of displaying the summary as text or playing it as audio.
[0501] This process allows users to efficiently interrupt and resume reading, enabling them to learn without wasting time.
[0502] (Example 1)
[0503] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0504] When users are forced to interrupt their information consumption, a problem arises in that it is difficult for them to remember where they left off and to efficiently understand the content from where they left off when they resume reading. In particular, with information that is read over a long period of time, reviewing and confirming key points is a time-consuming and laborious task. This invention aims to solve the above problems and provide a means for quickly and efficiently resuming information consumption.
[0505] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0506] In this invention, the server includes means for recording the location where the user interrupted information consumption, a computing device for automatically summarizing relevant information based on the recorded data, means for saving the summary and presenting it upon resumption, and a processing device for generating the summary based on a generation model and prompt statements. This enables the user to efficiently understand what was interrupted and quickly resume information consumption.
[0507] A "user" is an individual or group that consumes and manipulates information and utilizes a system.
[0508] "Information consumption" refers to the act of using and understanding content such as ebooks.
[0509] "The location where the information was stopped" refers to the specific location where the user stopped consuming information.
[0510] "Means of recording" refers to technical elements or methods for saving the location where the operation was stopped as data.
[0511] "Data" refers to recordable content, including information about user behavior and location.
[0512] An "automatic summarizing computing device" refers to hardware or software that uses a specified algorithm or model to extract and simplify the key points of information.
[0513] A "summary" refers to a simplified version of the original information, extracting only the main points.
[0514] "Means of preservation" refers to technical elements or methods for storing summarized information so that it can be used later.
[0515] "Means of presentation" refers to technical elements or methods for providing the saved summary to the user visually or aurally.
[0516] A "generative model" refers to an algorithm or machine learning framework that processes information and creates summaries.
[0517] A "prompt statement" is an instructional statement input into a generative model, referring to text that provides conditions and directions for generating a summary.
[0518] This invention is a system that efficiently supports the interruption and resumption of electronic information. The system mainly consists of a terminal, a server, and a generative AI model, and when a user interrupts the consumption of electronic information, it records the location and presents a summary when it resumes.
[0519] The terminal has a function to automatically record the location where the user interrupted information consumption. Specifically, the terminal sends information such as page number and paragraph position to the server. Hardware used for this includes information input devices (e.g., touchscreen, keyboard).
[0520] The server retrieves relevant content from a database based on location information received from the terminal. It then uses a generative AI model to summarize the relevant text. The generative AI model has the capability to generate summaries using specified prompt sentences as input. The software used specifically includes a natural language processing engine.
[0521] The summary is stored in a database by the server and presented on the terminal when the user resumes consuming information. The terminal receives the summary from the server and presents it to the user visually (display) and aurally (speech synthesis). This allows the user to efficiently resume consuming information from where they left off.
[0522] For example, if a user interrupts their use of self-improvement digital content and tries to resume it the next day, the device receives a summary from the server stating that "basic steps were explained and specific approaches were introduced," and presents this summary to the user, allowing for a smooth resumption of information consumption.
[0523] An example of a prompt would be to input to the generative AI model in the format: "Summarize the following text and include the main points: '[Retrieved Text]'". This allows the generative AI model to efficiently extract the key points and provide a user-friendly summary.
[0524] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0525] Step 1:
[0526] The device detects actions that interrupt the user's consumption of information. Specifically, when the user closes an app or navigates to another page, the device records the current page number and paragraph location. The input is the user's action, and the output is the location information of the interruption.
[0527] Step 2:
[0528] The device sends the recorded location information to the server. The server receives this information and searches its database for related content. The input in this process is the location information sent from the device, and the output is the corresponding content text. Specifically, the server extracts text data from the database that is associated with the interrupted location.
[0529] Step 3:
[0530] The server inputs the extracted text into a generative AI model and generates a summary using a prompt. This prompt might be something like, "Summarize the following text, including the main points: '[Retrieved Text]'". The input is the text data and the prompt, and the output is the summary result. Specifically, the generative AI model uses natural language processing techniques to extract the key points and generate a shortened summary.
[0531] Step 4:
[0532] The server saves the generated summary to the database. In this step, the input is the generated summary, and the output is the saved summary data. Specifically, it records the summary along with related metadata (e.g., user ID, reading position) in the database.
[0533] Step 5:
[0534] When a user resumes consuming information, the device sends a request to the server for a summary. The server searches for the relevant summary and sends it to the device. The input is the request from the device, and the output is the summary data. Specifically, the server quickly searches for and returns summaries that have been stored in the past.
[0535] Step 6:
[0536] The terminal presents the received summary to the user visually and audibly. It displays the text on a screen and reads it aloud using a text-to-speech device. The input is summary data from the server, and the output is the presentation of information to the user. Specifically, the terminal renders the summary in text format and simultaneously provides audio output.
[0537] (Application Example 1)
[0538] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0539] In many information processing tasks, a problem arises where users cannot quickly grasp the workflow after temporarily interrupting and resuming their work, leading to decreased work efficiency. This problem is particularly pronounced in online shopping and e-books, hindering user convenience.
[0540] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0541] In this invention, the server includes means for recording the last operation when a user interrupts an information processing task on an information processing terminal, means for automatically summarizing the work performed up to that point based on the recorded information, and means for saving the summary and presenting it to the user when they resume their work. This allows the user to efficiently resume their work after interruption and to proceed smoothly with their tasks.
[0542] An "information processing terminal" is a device used by users to perform various information processing tasks, and specifically includes personal computers and smartphones.
[0543] "Operation point" refers to the location or positional information that indicates when a user interrupted a specific information processing task, and represents the user's operation history.
[0544] "Work content" refers to the set of data handled and processes performed when a user performs information processing tasks.
[0545] "Automatic summarization" is a process in which the system reconstructs the main information in a concise format based on the input information, and this is done by the system without human intervention.
[0546] "Information display means" refers to functions or devices that visually present necessary information to the user, and specifically includes displays and monitors.
[0547] "Sound generation means" refers to technologies and devices that convey information to users through sound, and includes speakers and speech synthesis software.
[0548] A "data processing algorithm" refers to a set of procedures or methods for manipulating, analyzing, and transforming data for a specific purpose. One example is AI-powered data analysis techniques.
[0549] This invention primarily uses three elements to realize a system: a server, an information processing terminal, and a user.
[0550] The server receives information about the operation location transmitted from the information processing terminal. For example, when a user interrupts product selection while online shopping, information about that product page is sent to the server.
[0551] Based on this information, the server automatically generates a summary using a generative AI model. This summary is tailored to include key features, pricing, and review summaries of the product, and utilizes specific data processing algorithms.
[0552] The generated summary is stored in the server's database and provided to the information processing terminal when the user resumes their work. The terminal uses information display means and sound generation means to present the user with a summary of the resumed position through text display and audio output.
[0553] This system allows users to efficiently resume interrupted tasks and ensure smooth workflow. For example, suppose a user was considering purchasing an electronic appliance online during their commute one day, but interrupted their browsing. The next day, upon resuming, they would be provided with a summary such as, "This product is highly rated for its quiet operation and is 10% cheaper than the market average," enabling them to make a quick purchase decision.
[0554] An example of a prompt used in the generative AI model is: "Briefly summarize the characteristics of the following product: Product name, Key features, Price, User reviews."
[0555] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0556] Step 1:
[0557] The user interrupts their work on the information processing terminal. For example, if the user closes a browser tab on a shopping site, the terminal records the page information at that time (product ID, page URL, timestamp, etc.) and sends it to the server. The input is page information, and the output is the transmission of information to the server.
[0558] Step 2:
[0559] The server analyzes the page information received from the terminal. This analysis uses the received product ID and URL to retrieve detailed product data (name, features, price, reviews, etc.) from the database. The input is page information, and the output is the retrieved product details.
[0560] Step 3:
[0561] The server uses a generative AI model to create a summary based on the retrieved product details. The prompt "Briefly summarize the characteristics of the following product: Product name, Key features, Price, User reviews" is used to input data into the AI model and generate the summary. The input is the product details, and the output is the generated summary.
[0562] Step 4:
[0563] The server saves the generated summary to a database, ready for the user to resume. In the saving step, in addition to the summary, the original page information is also recorded. The input is the generated summary, and the output is the database record of the summary.
[0564] Step 5:
[0565] When a user resumes work, the terminal requests a summary from the server. The server retrieves the corresponding summary from the database and sends it to the terminal. The input is the resume request, and the output is the terminal sending of the summary.
[0566] Step 6:
[0567] The terminal presents the received summary to the user. It displays the information as text using an information display device and plays it back as audio using an audio generation device. The input is summary data from the server, and the output is the presentation of the summary to the user.
[0568] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0569] This invention provides users with an efficient e-book reading experience, and in particular, it not only provides summaries up to the point of interruption, but also features a mechanism to provide customized summaries based on the user's emotions using an emotion engine. This system consists of a server, a terminal, and a user, and the addition of the emotion engine enables flexible information provision according to the emotional state of each individual user.
[0570] System Configuration
[0571] Device: A device that allows users to read ebooks and record their reading progress. The device is equipped with a camera and microphone, which are used to collect the user's voice and facial expressions for emotion recognition.
[0572] Server: Generates an ebook summary based on received data. Utilizing an AI engine, it adjusts the summary to provide the user with the most relevant information, taking into account user sentiment information provided by the sentiment engine.
[0573] Emotion Engine: This software analyzes the user's voice and facial expression data to determine their emotional state. The emotion engine provides information about the user's emotions to the server, which is then used to generate summaries.
[0574] Program operation description
[0575] 1. Acquiring emotional data during reading.
[0576] While the user is reading an e-book, the device acquires voice and facial information through its camera and microphone. This information is analyzed by an emotion engine, which identifies the user's emotional state in real time.
[0577] 2. How to handle interruptions while reading
[0578] When a user interrupts reading, the device sends that location to the server. At the same time, the latest sentiment analysis results are also sent to the server.
[0579] 3. Summary generation and customization
[0580] The server retrieves the content of the ebook up to a specified point and generates a summary using an AI engine. The generated summary is then adjusted in content and presentation style based on emotional information from an emotional engine.
[0581] 4. Presentation to the user
[0582] After generating the summary, the server sends it to the terminal in an appropriate format based on the user's emotions. For example, a bright tone of voice synthesis is used for positive emotional states, and the content is designed to better reflect those emotions.
[0583] Specific example
[0584] When a user reads a self-help book, the device analyzes the user's facial expressions and voice to confirm that the user is relaxed. If the user interrupts reading and resumes reading the next day, the server presents a summary in a calm tone of voice that takes the user's relaxed state into account, providing a peaceful reading experience. In this way, the system dynamically adjusts content based on the user's emotions, providing even more personalized reading support.
[0585] The following describes the processing flow.
[0586] Step 1:
[0587] When a user begins reading an ebook, the device retrieves the ebook's content and user information, and records the start of reading. The device then uses its camera and microphone to begin capturing the user's voice and facial expressions in real time.
[0588] Step 2:
[0589] The device sends the acquired voice and facial expression data to the emotion engine, which analyzes the emotional state. The emotion engine identifies the user's emotions and returns that information to the device. The analysis is performed periodically to create a user emotion log.
[0590] Step 3:
[0591] When a user attempts to interrupt reading, they press the pause button on their device. The device then sends the current page information and the latest sentiment log data to the server.
[0592] Step 4:
[0593] The server retrieves the relevant text from the database based on the received page information. The server then activates an AI engine, analyzes the content up to a specified point, and generates a summary. This summary includes customization based on the user's sentiment data.
[0594] Step 5:
[0595] The server stores the generated summary in a database and associates it with the user's ID. The summary also retains corresponding sentiment data as metadata.
[0596] Step 6:
[0597] When the user operates the device to resume reading, the device sends a request for a summary to the server.
[0598] Step 7:
[0599] The server returns the stored summary and its associated sentiment information to the terminal.
[0600] Step 8:
[0601] The device presents the received summary to the user. When presenting the summary, it reflects the user's previous emotional state; for example, if the user is in a calm emotional state, the tone of the synthesized speech is adjusted to be gentler. The user can also choose to view the summary as text, and the device automatically selects a font and background color appropriate to the user's emotion.
[0602] This process presents users with summaries that match their emotions, further enhancing their reading experience.
[0603] (Example 2)
[0604] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0605] Conventional electronic media viewing systems simply record where the user paused and provide a summary up to that point. Therefore, they are unable to provide individual summaries that take into account the user's emotional state, highlighting the need for more personalized information delivery. Furthermore, while adjusting summaries based on emotions can improve the user experience, such a mechanism has not existed.
[0606] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0607] In this invention, the server includes means for recording the last interruption point when a user interrupts the content of an electronic medium; means for automatically summarizing the content up to that point based on the recorded information; means for identifying the user's emotional state by analyzing voice and facial data; means for adjusting the summary content based on the identified emotional state; and means for saving the summary and presenting it when the user resumes viewing. This makes it possible to provide information in a form that is appropriate to the user's emotions.
[0608] A "user" refers to an individual or group that obtains information using electronic media.
[0609] "Electronic media" refers to media that can provide information in digital format, including ebooks and digital documents.
[0610] "Interruption point" refers to information indicating the point at which the user temporarily stopped using the electronic medium.
[0611] A "summary" refers to information that concisely summarizes the content of an electronic medium and is provided in a format that is easy for the user to understand.
[0612] "Emotional state" refers to a temporary psychological state identified by analyzing the user's voice and facial data.
[0613] A "server" is a part of a system that provides services to user terminals through data processing and storage functions.
[0614] "Voice and facial data" refers to digital information including the user's tone of voice and facial expressions, which is used to identify their emotional state.
[0615] "Adjustment" refers to changing the content and presentation method of a summary based on the user's identified emotional state.
[0616] "Means" refers to a method or apparatus for achieving a specific function.
[0617] This system consists of users, terminals, and servers, and provides users with an efficient and personalized browsing experience of electronic media. Specifically, it is implemented as follows:
[0618] Users use a terminal when browsing electronic media. This terminal is a digital device with a built-in camera and microphone. This allows the terminal to capture data on the user's voice and facial expressions. This digital data is then transferred to an emotion engine to analyze the user's emotional state.
[0619] The server receives information about the interruption point and emotional data from the terminal. Based on this information, the server uses a generative AI model to summarize the content of the electronic medium. At this time, the summary is adjusted according to the user's emotional state, which is analyzed by the emotional engine. For example, the content is adjusted to be presented in a calm tone to a relaxed user.
[0620] As a concrete example, if a user is in a relaxed state while reading a self-help book, the system will use that emotional state to present a summary in a calm tone via speech synthesis when they resume reading the book the next day, providing a peaceful reading experience.
[0621] An example of a prompt message is, "Generate an electronic summary based on the user's current emotional state. The user is relaxed." In this way, the system can provide meaningful information in real time that corresponds to the user's specific emotions.
[0622] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0623] Step 1:
[0624] The device uses its built-in camera and microphone to collect the user's voice and facial expressions while they are browsing electronic media. The collected data is processed in real time within the device to extract attributes necessary for identifying the user's emotional state (e.g., voice tone and facial expression features). This serves as the input, and the output generates the dataset necessary for sentiment analysis.
[0625] Step 2:
[0626] The device sends the collected dataset to the emotion engine. The emotion engine uses an AI model to analyze this dataset and identify the user's current emotional state. In this step, voice tone and facial features are used as input, and the user's emotional state (e.g., relaxed, stressed) is output.
[0627] Step 3:
[0628] When a user interrupts their viewing of electronic media, the device sends information about the interruption point and their current emotional state to the server. This input—the interruption point and emotional state—initiates the server-side summary generation process.
[0629] Step 4:
[0630] The server retrieves the content of the electronic media up to the point of interruption. It then uses an AI engine to invoke a generative AI model, which generates a summary based on the input electronic media content and emotional state. The generative AI model analyzes relevant information to generate a summary, adjusting it in the process to consider the user's emotional state. In this step, the electronic media content and emotional state are the inputs, and a summary tailored to the user's emotions is output.
[0631] Step 5:
[0632] The server sends the generated summary to the terminal. The terminal displays the received summary on its screen and then reads the summary aloud using speech synthesis. Here, the adjusted summary becomes the input and is output as specific visual and auditory information presented to the user.
[0633] (Application Example 2)
[0634] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0635] Modern e-readers lack personalized summaries that take into account the user's emotional state when reading is interrupted, which hinders reading continuity and immersion. Furthermore, to ensure a comfortable reading experience upon resuming, summaries need to be presented not just in a simple format, but in a style that matches the user's emotional state.
[0636] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0637] In this invention, the server includes means for recording the last reading position and emotional state when reading an ebook is interrupted, means for automatically summarizing the content of the ebook up to that point based on the recorded information and the user's emotional state, and adjusting the presentation style, and means for saving the summary and presenting it as an emotionally-based summary when the user resumes reading. This makes it possible to provide the user with personalized reading support that matches their emotional state.
[0638] A "user" is an individual who reads ebooks and is the subject who utilizes the system.
[0639] An "eBook" is a book provided in digital format and is content that can be read on an electronic device.
[0640] "Interrupting reading" refers to the act of a user temporarily stopping the viewing of an e-book.
[0641] "Last read position" refers to the last part of the ebook that the user read before interrupting their reading.
[0642] "Emotional state" refers to information that represents the user's current emotions, and is usually obtained from voice or facial expression data.
[0643] "To record" refers to the act of storing information in a database or storage device in order to preserve specific data.
[0644] A "summary" is a shortened version of an e-book, containing only the key points.
[0645] "Presentation style" refers to the format and method of providing information to users.
[0646] "Saving" refers to maintaining data for later reuse.
[0647] "Means" refers to the methods or devices used to achieve a specific purpose.
[0648] "Personalized" means that it is delivered in a way that is optimized for each individual user.
[0649] This invention is a system for providing a personalized reading experience to users of ebooks. Its embodiments are described below.
[0650] The system consists of a user terminal, a central server, and an emotion engine that performs sentiment analysis. While the user is reading an ebook, the terminal uses its camera and microphone to capture voice and facial expressions in real time. The captured data is analyzed by the emotion engine to identify the user's current emotional state. This sentiment information is simultaneously transmitted to the server.
[0651] When a user pauses reading, the device records the point of interruption and their emotional state, and sends this information to a server. Based on this recorded information, the server uses an AI model to create a summary of the ebook up to the point of interruption. During this process, the content and presentation style of the summary are adjusted based on the emotional information provided by the emotion engine. For example, if the user is in a relaxed state, the summary may be presented in a calm tone of voice, employing an emotion-based approach.
[0652] The generated summary is provided to the device as audio data generated using speech synthesis technology, as well as as text. A modern speech synthesis engine is used to improve speech quality.
[0653] For example, if the system detects that a user is smiling while reading an emotionally moving novel, a summary will be provided in a warm tone of voice when they resume reading the next day, reminding them of their previous reading experience. An example of a prompt for the generative AI model would be: "Generate a summary of the ebook up to where the user left off. The user's emotion is 'relaxed,' so please use a summary style that matches that emotion. In particular, please insert words that evoke a sense of relaxation at the beginning of the summary."
[0654] In this way, the present invention realizes a personalized reading experience based on the user's emotions.
[0655] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0656] Step 1:
[0657] While the user is reading an e-book, the device uses its built-in camera and microphone to acquire audio and facial expression data. This input data is sent to an emotion engine, which analyzes the data to identify the user's emotional state in real time.
[0658] Step 2:
[0659] When a user interrupts reading, the device records the point of interruption and the emotional state provided by the emotion engine. This information, including emotional data, is sent to the server, allowing the server to gain a comprehensive understanding of the interruption.
[0660] Step 3:
[0661] The server, based on the received interruption point and emotional information, forms a prompt sentence for the AI engine and sends it to the generative AI model. The generative AI model uses this prompt sentence to summarize the content of the ebook up to the interruption point. This summary is adjusted according to the emotional state.
[0662] Step 4:
[0663] The generated summary is converted into an audio file using a speech synthesis engine. The speech synthesis engine operates with a tone and speed set based on emotional information. The output audio file and text summary are sent to the terminal.
[0664] Step 5:
[0665] When the user resumes reading, the device presents the received summary in an emotionally appropriate style. The user can choose between audio playback or text display, providing a personalized reading experience. This step ensures the user receives information that matches their current emotional state.
[0666] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0667] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0668] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0669] [Fourth Embodiment]
[0670] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0671] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0672] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0673] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0674] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0675] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0676] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0677] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0678] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0679] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0680] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0681] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0682] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0683] This invention is a system for supporting users in efficiently resuming reading ebooks after interrupting their reading. This system primarily consists of a server, a terminal, and the user.
[0684] System Configuration
[0685] Device: This is the device that a user uses to read ebooks. The device records the current page and position while reading, and saves the position when the user interrupts reading.
[0686] Server: Receives data sent from the terminal during an interruption and uses the AI engine to summarize the text within the specified range. The server stores the generated summary in a database and sends the summary to the terminal when the user resumes.
[0687] Program operation description
[0688] 1. How to handle interruptions while reading
[0689] When a user stops reading an ebook, the device sends its location to the server. This location information includes page numbers and paragraph information.
[0690] 2. Summary generation
[0691] The server retrieves the relevant text from its database based on the received location information and generates a summary using an AI engine. The summary is adjusted to include the main points and the core of the argument.
[0692] 3. Summarize, save, and present
[0693] After the summary is generated, the server saves it to a database. Later, when the user resumes reading, the summary is provided upon request from the terminal.
[0694] When the terminal receives a summary from the server, it displays it as text and also presents it to the user as audio data.
[0695] Specific example
[0696] For example, suppose a user is reading an ebook on self-improvement. The user reads during their daily commute, but today they needed to stop before finishing Chapter 2, "Goal-Setting Techniques." The device sends this interruption point to the server, and the server uses AI to summarize the content up to that point. The next day, when the user resumes reading during their commute, the device provides a summary in both audio and text format, stating that "basic steps for setting goals were explained, and specific approaches to self-actualization were introduced," allowing them to quickly resume reading.
[0697] This system allows users to efficiently utilize their valuable time and learn without spending long hours.
[0698] The following describes the processing flow.
[0699] Step 1:
[0700] When a user starts reading an ebook, the device acquires that information and begins recording the current page and position. This allows the device to track which part of the book the user is reading in real time.
[0701] Step 2:
[0702] When a user pauses reading, they press the pause button on their device. The device detects this action and determines the location information of the last text read (page number, paragraph number, etc.).
[0703] Step 3:
[0704] The device sends confirmed location information to the server. This information includes the page number and paragraph number at the time of interruption.
[0705] Step 4:
[0706] Based on the received location information, the server retrieves the content of ebooks up to that location from its database.
[0707] Step 5:
[0708] The server activates the AI engine and generates a summary of the acquired text. The AI engine extracts key points and condenses them into a short text.
[0709] Step 6:
[0710] The server saves the generated summary to a database and associates it with the user's ID. This allows for quick access to the summary later.
[0711] Step 7:
[0712] When the user resumes reading, the device sends a request for a summary to the server.
[0713] Step 8:
[0714] The server retrieves the stored summary from the database and sends it back to the terminal.
[0715] Step 9:
[0716] The device displays the received summary to the user. It offers the user the option of displaying the summary as text or playing it as audio.
[0717] This process allows users to efficiently interrupt and resume reading, enabling them to learn without wasting time.
[0718] (Example 1)
[0719] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0720] When users are forced to interrupt their information consumption, a problem arises in that it is difficult for them to remember where they left off and to efficiently understand the content from where they left off when they resume reading. In particular, with information that is read over a long period of time, reviewing and confirming key points is a time-consuming and laborious task. This invention aims to solve the above problems and provide a means for quickly and efficiently resuming information consumption.
[0721] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0722] In this invention, the server includes means for recording the location where the user interrupted information consumption, a computing device for automatically summarizing relevant information based on the recorded data, means for saving the summary and presenting it upon resumption, and a processing device for generating the summary based on a generation model and prompt statements. This enables the user to efficiently understand what was interrupted and quickly resume information consumption.
[0723] A "user" is an individual or group that consumes and manipulates information and utilizes a system.
[0724] "Information consumption" refers to the act of using and understanding content such as ebooks.
[0725] "The location where the information was stopped" refers to the specific location where the user stopped consuming information.
[0726] "Means of recording" refers to technical elements or methods for saving the location where the operation was stopped as data.
[0727] "Data" refers to recordable content, including information about user behavior and location.
[0728] An "automatic summarizing computing device" refers to hardware or software that uses a specified algorithm or model to extract and simplify the key points of information.
[0729] A "summary" refers to a simplified version of the original information, extracting only the main points.
[0730] "Means of preservation" refers to technical elements or methods for storing summarized information so that it can be used later.
[0731] "Means of presentation" refers to technical elements or methods for providing the saved summary to the user visually or aurally.
[0732] A "generative model" refers to an algorithm or machine learning framework that processes information and creates summaries.
[0733] A "prompt statement" is an instructional statement input into a generative model, referring to text that provides conditions and directions for generating a summary.
[0734] This invention is a system that efficiently supports the interruption and resumption of electronic information. The system mainly consists of a terminal, a server, and a generative AI model, and when a user interrupts the consumption of electronic information, it records the location and presents a summary when it resumes.
[0735] The terminal has a function to automatically record the location where the user interrupted information consumption. Specifically, the terminal sends information such as page number and paragraph position to the server. Hardware used for this includes information input devices (e.g., touchscreen, keyboard).
[0736] The server retrieves relevant content from a database based on location information received from the terminal. It then uses a generative AI model to summarize the relevant text. The generative AI model has the capability to generate summaries using specified prompt sentences as input. The software used specifically includes a natural language processing engine.
[0737] The summary is stored in a database by the server and presented on the terminal when the user resumes consuming information. The terminal receives the summary from the server and presents it to the user visually (display) and aurally (speech synthesis). This allows the user to efficiently resume consuming information from where they left off.
[0738] For example, if a user interrupts their use of self-improvement digital content and tries to resume it the next day, the device receives a summary from the server stating that "basic steps were explained and specific approaches were introduced," and presents this summary to the user, allowing for a smooth resumption of information consumption.
[0739] An example of a prompt would be to input to the generative AI model in the format: "Summarize the following text and include the main points: '[Retrieved Text]'". This allows the generative AI model to efficiently extract the key points and provide a user-friendly summary.
[0740] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0741] Step 1:
[0742] The device detects actions that interrupt the user's consumption of information. Specifically, when the user closes an app or navigates to another page, the device records the current page number and paragraph location. The input is the user's action, and the output is the location information of the interruption.
[0743] Step 2:
[0744] The device sends the recorded location information to the server. The server receives this information and searches its database for related content. The input in this process is the location information sent from the device, and the output is the corresponding content text. Specifically, the server extracts text data from the database that is associated with the interrupted location.
[0745] Step 3:
[0746] The server inputs the extracted text into a generative AI model and generates a summary using a prompt. This prompt might be something like, "Summarize the following text, including the main points: '[Retrieved Text]'". The input is the text data and the prompt, and the output is the summary result. Specifically, the generative AI model uses natural language processing techniques to extract the key points and generate a shortened summary.
[0747] Step 4:
[0748] The server saves the generated summary to the database. In this step, the input is the generated summary, and the output is the saved summary data. Specifically, it records the summary along with related metadata (e.g., user ID, reading position) in the database.
[0749] Step 5:
[0750] When a user resumes consuming information, the device sends a request to the server for a summary. The server searches for the relevant summary and sends it to the device. The input is the request from the device, and the output is the summary data. Specifically, the server quickly searches for and returns summaries that have been stored in the past.
[0751] Step 6:
[0752] The terminal presents the received summary to the user visually and audibly. It displays the text on a screen and reads it aloud using a text-to-speech device. The input is summary data from the server, and the output is the presentation of information to the user. Specifically, the terminal renders the summary in text format and simultaneously provides audio output.
[0753] (Application Example 1)
[0754] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0755] In many information processing tasks, a problem arises where users cannot quickly grasp the workflow after temporarily interrupting and resuming their work, leading to decreased work efficiency. This problem is particularly pronounced in online shopping and e-books, hindering user convenience.
[0756] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0757] In this invention, the server includes means for recording the last operation when a user interrupts an information processing task on an information processing terminal, means for automatically summarizing the work performed up to that point based on the recorded information, and means for saving the summary and presenting it to the user when they resume their work. This allows the user to efficiently resume their work after interruption and to proceed smoothly with their tasks.
[0758] An "information processing terminal" is a device used by users to perform various information processing tasks, and specifically includes personal computers and smartphones.
[0759] "Operation point" refers to the location or positional information that indicates when a user interrupted a specific information processing task, and represents the user's operation history.
[0760] "Work content" refers to the set of data handled and processes performed when a user performs information processing tasks.
[0761] "Automatic summarization" is a process in which the system reconstructs the main information in a concise format based on the input information, and this is done by the system without human intervention.
[0762] "Information display means" refers to functions or devices that visually present necessary information to the user, and specifically includes displays and monitors.
[0763] "Sound generation means" refers to technologies and devices that convey information to users through sound, and includes speakers and speech synthesis software.
[0764] A "data processing algorithm" refers to a set of procedures or methods for manipulating, analyzing, and transforming data for a specific purpose. One example is AI-powered data analysis techniques.
[0765] This invention primarily uses three elements to realize a system: a server, an information processing terminal, and a user.
[0766] The server receives information about the operation location transmitted from the information processing terminal. For example, when a user interrupts product selection while online shopping, information about that product page is sent to the server.
[0767] Based on this information, the server automatically generates a summary using a generative AI model. This summary is tailored to include key features, pricing, and review summaries of the product, and utilizes specific data processing algorithms.
[0768] The generated summary is stored in the server's database and provided to the information processing terminal when the user resumes their work. The terminal uses information display means and sound generation means to present the user with a summary of the resumed position through text display and audio output.
[0769] This system allows users to efficiently resume interrupted tasks and ensure smooth workflow. For example, suppose a user was considering purchasing an electronic appliance online during their commute one day, but interrupted their browsing. The next day, upon resuming, they would be provided with a summary such as, "This product is highly rated for its quiet operation and is 10% cheaper than the market average," enabling them to make a quick purchase decision.
[0770] An example of a prompt used in the generative AI model is: "Briefly summarize the characteristics of the following product: Product name, Key features, Price, User reviews."
[0771] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0772] Step 1:
[0773] The user interrupts their work on the information processing terminal. For example, if the user closes a browser tab on a shopping site, the terminal records the page information at that time (product ID, page URL, timestamp, etc.) and sends it to the server. The input is page information, and the output is the transmission of information to the server.
[0774] Step 2:
[0775] The server analyzes the page information received from the terminal. This analysis uses the received product ID and URL to retrieve detailed product data (name, features, price, reviews, etc.) from the database. The input is page information, and the output is the retrieved product details.
[0776] Step 3:
[0777] The server uses a generative AI model to create a summary based on the retrieved product details. The prompt "Briefly summarize the characteristics of the following product: Product name, Key features, Price, User reviews" is used to input data into the AI model and generate the summary. The input is the product details, and the output is the generated summary.
[0778] Step 4:
[0779] The server saves the generated summary to a database, ready for the user to resume. In the saving step, in addition to the summary, the original page information is also recorded. The input is the generated summary, and the output is the database record of the summary.
[0780] Step 5:
[0781] When a user resumes work, the terminal requests a summary from the server. The server retrieves the corresponding summary from the database and sends it to the terminal. The input is the resume request, and the output is the terminal sending of the summary.
[0782] Step 6:
[0783] The terminal presents the received summary to the user. It displays the information as text using an information display device and plays it back as audio using an audio generation device. The input is summary data from the server, and the output is the presentation of the summary to the user.
[0784] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0785] This invention provides users with an efficient e-book reading experience, and in particular, it not only provides summaries up to the point of interruption, but also features a mechanism to provide customized summaries based on the user's emotions using an emotion engine. This system consists of a server, a terminal, and a user, and the addition of the emotion engine enables flexible information provision according to the emotional state of each individual user.
[0786] System Configuration
[0787] Device: A device that allows users to read ebooks and record their reading progress. The device is equipped with a camera and microphone, which are used to collect the user's voice and facial expressions for emotion recognition.
[0788] Server: Generates an ebook summary based on received data. Utilizing an AI engine, it adjusts the summary to provide the user with the most relevant information, taking into account user sentiment information provided by the sentiment engine.
[0789] Emotion Engine: This software analyzes the user's voice and facial expression data to determine their emotional state. The emotion engine provides information about the user's emotions to the server, which is then used to generate summaries.
[0790] Program operation description
[0791] 1. Acquiring emotional data during reading.
[0792] While the user is reading an e-book, the device acquires voice and facial information through its camera and microphone. This information is analyzed by an emotion engine, which identifies the user's emotional state in real time.
[0793] 2. How to handle interruptions while reading
[0794] When a user interrupts reading, the device sends that location to the server. At the same time, the latest sentiment analysis results are also sent to the server.
[0795] 3. Summary generation and customization
[0796] The server retrieves the content of the ebook up to a specified point and generates a summary using an AI engine. The generated summary is then adjusted in content and presentation style based on emotional information from an emotional engine.
[0797] 4. Presentation to the user
[0798] After generating the summary, the server sends it to the terminal in an appropriate format based on the user's emotions. For example, a bright tone of voice synthesis is used for positive emotional states, and the content is designed to better reflect those emotions.
[0799] Specific example
[0800] When a user reads a self-help book, the device analyzes the user's facial expressions and voice to confirm that the user is relaxed. If the user interrupts reading and resumes reading the next day, the server presents a summary in a calm tone of voice that takes the user's relaxed state into account, providing a peaceful reading experience. In this way, the system dynamically adjusts content based on the user's emotions, providing even more personalized reading support.
[0801] The following describes the processing flow.
[0802] Step 1:
[0803] When a user begins reading an ebook, the device retrieves the ebook's content and user information, and records the start of reading. The device then uses its camera and microphone to begin capturing the user's voice and facial expressions in real time.
[0804] Step 2:
[0805] The device sends the acquired voice and facial expression data to the emotion engine, which analyzes the emotional state. The emotion engine identifies the user's emotions and returns that information to the device. The analysis is performed periodically to create a user emotion log.
[0806] Step 3:
[0807] When a user attempts to interrupt reading, they press the pause button on their device. The device then sends the current page information and the latest sentiment log data to the server.
[0808] Step 4:
[0809] The server retrieves the relevant text from the database based on the received page information. The server then activates an AI engine, analyzes the content up to a specified point, and generates a summary. This summary includes customization based on the user's sentiment data.
[0810] Step 5:
[0811] The server stores the generated summary in a database and associates it with the user's ID. The summary also retains corresponding sentiment data as metadata.
[0812] Step 6:
[0813] When the user operates the device to resume reading, the device sends a request for a summary to the server.
[0814] Step 7:
[0815] The server returns the stored summary and its associated sentiment information to the terminal.
[0816] Step 8:
[0817] The device presents the received summary to the user. When presenting the summary, it reflects the user's previous emotional state; for example, if the user is in a calm emotional state, the tone of the synthesized speech is adjusted to be gentler. The user can also choose to view the summary as text, and the device automatically selects a font and background color appropriate to the user's emotion.
[0818] This process presents users with summaries that match their emotions, further enhancing their reading experience.
[0819] (Example 2)
[0820] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0821] Conventional electronic media viewing systems simply record where the user paused and provide a summary up to that point. Therefore, they are unable to provide individual summaries that take into account the user's emotional state, highlighting the need for more personalized information delivery. Furthermore, while adjusting summaries based on emotions can improve the user experience, such a mechanism has not existed.
[0822] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0823] In this invention, the server includes means for recording the last interruption point when a user interrupts the content of an electronic medium; means for automatically summarizing the content up to that point based on the recorded information; means for identifying the user's emotional state by analyzing voice and facial data; means for adjusting the summary content based on the identified emotional state; and means for saving the summary and presenting it when the user resumes viewing. This makes it possible to provide information in a form that is appropriate to the user's emotions.
[0824] A "user" refers to an individual or group that obtains information using electronic media.
[0825] "Electronic media" refers to media that can provide information in digital format, including ebooks and digital documents.
[0826] "Interruption point" refers to information indicating the point at which the user temporarily stopped using the electronic medium.
[0827] A "summary" refers to information that concisely summarizes the content of an electronic medium and is provided in a format that is easy for the user to understand.
[0828] "Emotional state" refers to a temporary psychological state identified by analyzing the user's voice and facial data.
[0829] A "server" is a part of a system that provides services to user terminals through data processing and storage functions.
[0830] "Voice and facial data" refers to digital information including the user's tone of voice and facial expressions, which is used to identify their emotional state.
[0831] "Adjustment" refers to changing the content and presentation method of a summary based on the user's identified emotional state.
[0832] "Means" refers to a method or apparatus for achieving a specific function.
[0833] This system consists of users, terminals, and servers, and provides users with an efficient and personalized browsing experience of electronic media. Specifically, it is implemented as follows:
[0834] Users use a terminal when browsing electronic media. This terminal is a digital device with a built-in camera and microphone. This allows the terminal to capture data on the user's voice and facial expressions. This digital data is then transferred to an emotion engine to analyze the user's emotional state.
[0835] The server receives information about the interruption point and emotional data from the terminal. Based on this information, the server uses a generative AI model to summarize the content of the electronic medium. At this time, the summary is adjusted according to the user's emotional state, which is analyzed by the emotional engine. For example, the content is adjusted to be presented in a calm tone to a relaxed user.
[0836] As a concrete example, if a user is in a relaxed state while reading a self-help book, the system will use that emotional state to present a summary in a calm tone via speech synthesis when they resume reading the book the next day, providing a peaceful reading experience.
[0837] An example of a prompt message is, "Generate an electronic summary based on the user's current emotional state. The user is relaxed." In this way, the system can provide meaningful information in real time that corresponds to the user's specific emotions.
[0838] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0839] Step 1:
[0840] The device uses its built-in camera and microphone to collect the user's voice and facial expressions while they are browsing electronic media. The collected data is processed in real time within the device to extract attributes necessary for identifying the user's emotional state (e.g., voice tone and facial expression features). This serves as the input, and the output generates the dataset necessary for sentiment analysis.
[0841] Step 2:
[0842] The device sends the collected dataset to the emotion engine. The emotion engine uses an AI model to analyze this dataset and identify the user's current emotional state. In this step, voice tone and facial features are used as input, and the user's emotional state (e.g., relaxed, stressed) is output.
[0843] Step 3:
[0844] When a user interrupts their viewing of electronic media, the device sends information about the interruption point and their current emotional state to the server. This input—the interruption point and emotional state—initiates the server-side summary generation process.
[0845] Step 4:
[0846] The server retrieves the content of the electronic media up to the point of interruption. It then uses an AI engine to invoke a generative AI model, which generates a summary based on the input electronic media content and emotional state. The generative AI model analyzes relevant information to generate a summary, adjusting it in the process to consider the user's emotional state. In this step, the electronic media content and emotional state are the inputs, and a summary tailored to the user's emotions is output.
[0847] Step 5:
[0848] The server sends the generated summary to the terminal. The terminal displays the received summary on its screen and then reads the summary aloud using speech synthesis. Here, the adjusted summary becomes the input and is output as specific visual and auditory information presented to the user.
[0849] (Application Example 2)
[0850] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0851] Modern e-readers lack personalized summaries that take into account the user's emotional state when reading is interrupted, which hinders reading continuity and immersion. Furthermore, to ensure a comfortable reading experience upon resuming, summaries need to be presented not just in a simple format, but in a style that matches the user's emotional state.
[0852] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0853] In this invention, the server includes means for recording the last reading position and emotional state when reading an ebook is interrupted, means for automatically summarizing the content of the ebook up to that point based on the recorded information and the user's emotional state, and adjusting the presentation style, and means for saving the summary and presenting it as an emotionally-based summary when the user resumes reading. This makes it possible to provide the user with personalized reading support that matches their emotional state.
[0854] A "user" is an individual who reads ebooks and is the subject who utilizes the system.
[0855] An "eBook" is a book provided in digital format and is content that can be read on an electronic device.
[0856] "Interrupting reading" refers to the act of a user temporarily stopping the viewing of an e-book.
[0857] "Last read position" refers to the last part of the ebook that the user read before interrupting their reading.
[0858] "Emotional state" refers to information that represents the user's current emotions, and is usually obtained from voice or facial expression data.
[0859] "To record" refers to the act of storing information in a database or storage device in order to preserve specific data.
[0860] A "summary" is a shortened version of an e-book, containing only the key points.
[0861] "Presentation style" refers to the format and method of providing information to users.
[0862] "Saving" refers to maintaining data for later reuse.
[0863] "Means" refers to the methods or devices used to achieve a specific purpose.
[0864] "Personalized" means that it is delivered in a way that is optimized for each individual user.
[0865] This invention is a system for providing a personalized reading experience to users of ebooks. Its embodiments are described below.
[0866] The system consists of a user terminal, a central server, and an emotion engine that performs sentiment analysis. While the user is reading an ebook, the terminal uses its camera and microphone to capture voice and facial expressions in real time. The captured data is analyzed by the emotion engine to identify the user's current emotional state. This sentiment information is simultaneously transmitted to the server.
[0867] When a user pauses reading, the device records the point of interruption and their emotional state, and sends this information to a server. Based on this recorded information, the server uses an AI model to create a summary of the ebook up to the point of interruption. During this process, the content and presentation style of the summary are adjusted based on the emotional information provided by the emotion engine. For example, if the user is in a relaxed state, the summary may be presented in a calm tone of voice, employing an emotion-based approach.
[0868] The generated summary is provided to the device as audio data generated using speech synthesis technology, as well as as text. A modern speech synthesis engine is used to improve speech quality.
[0869] For example, if the system detects that a user is smiling while reading an emotionally moving novel, a summary will be provided in a warm tone of voice when they resume reading the next day, reminding them of their previous reading experience. An example of a prompt for the generative AI model would be: "Generate a summary of the ebook up to where the user left off. The user's emotion is 'relaxed,' so please use a summary style that matches that emotion. In particular, please insert words that evoke a sense of relaxation at the beginning of the summary."
[0870] In this way, the present invention realizes a personalized reading experience based on the user's emotions.
[0871] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0872] Step 1:
[0873] While the user is reading an e-book, the device uses its built-in camera and microphone to acquire audio and facial expression data. This input data is sent to an emotion engine, which analyzes the data to identify the user's emotional state in real time.
[0874] Step 2:
[0875] When a user interrupts reading, the device records the point of interruption and the emotional state provided by the emotion engine. This information, including emotional data, is sent to the server, allowing the server to gain a comprehensive understanding of the interruption.
[0876] Step 3:
[0877] The server, based on the received interruption point and emotional information, forms a prompt sentence for the AI engine and sends it to the generative AI model. The generative AI model uses this prompt sentence to summarize the content of the ebook up to the interruption point. This summary is adjusted according to the emotional state.
[0878] Step 4:
[0879] The generated summary is converted into an audio file using a speech synthesis engine. The speech synthesis engine operates with a tone and speed set based on emotional information. The output audio file and text summary are sent to the terminal.
[0880] Step 5:
[0881] When the user resumes reading, the device presents the received summary in an emotionally appropriate style. The user can choose between audio playback or text display, providing a personalized reading experience. This step ensures the user receives information that matches their current emotional state.
[0882] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0883] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0884] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0885] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0886] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0887] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0888] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0889] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0890] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0891] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0892] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0893] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0894] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0895] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0896] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0897] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0898] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0899] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0900] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0901] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0902] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0903] The following is further disclosed regarding the embodiments described above.
[0904] (Claim 1)
[0905] A means of recording the last reading position when a user interrupts reading an ebook,
[0906] A means for automatically summarizing the content of an e-book up to the relevant section based on the recorded information,
[0907] A means for saving the aforementioned summary and presenting it when the user resumes reading,
[0908] A system that includes this.
[0909] (Claim 2)
[0910] The system according to claim 1, wherein the summary is presented to the user by text display and speech synthesis.
[0911] (Claim 3)
[0912] The system according to claim 1, wherein a specific algorithm is used to extract key points when generating the summary.
[0913] "Example 1"
[0914] (Claim 1)
[0915] A means of recording the location where a user interrupts information consumption,
[0916] A computing device that automatically summarizes relevant information based on the recorded data,
[0917] Means for saving the aforementioned summary and presenting it when the user resumes consuming information,
[0918] A processing unit that generates a summary based on a generative model and prompt statements,
[0919] A system that includes this.
[0920] (Claim 2)
[0921] The system according to claim 1, wherein the summary is presented to the user by a display device and a speech synthesis device.
[0922] (Claim 3)
[0923] The system according to claim 1, wherein in generating the summary, a specific mathematical method is used to extract important elements.
[0924] "Application Example 1"
[0925] (Claim 1)
[0926] A means for recording the last operation performed when a user interrupts an information processing task on an information processing terminal,
[0927] A means for automatically summarizing the work performed up to the relevant section based on the recorded information,
[0928] Means for saving the aforementioned summary and presenting it when the user resumes work,
[0929] A system that includes this.
[0930] (Claim 2)
[0931] The system according to claim 1, wherein the summary is presented to the user by an information display means and an audio generation means.
[0932] (Claim 3)
[0933] The system according to claim 1, wherein a specific data processing algorithm is used to extract important elements when generating the summary.
[0934] "Example 2 of combining an emotion engine"
[0935] (Claim 1)
[0936] A means of recording the last interruption point when a user interrupts the content of an electronic medium,
[0937] A means for automatically summarizing the content of the electronic media up to the relevant portion based on the recorded information,
[0938] In generating the aforementioned summary, means for analyzing voice and facial data to identify the user's emotional state,
[0939] A means of adjusting the summary content based on identified emotional states,
[0940] Means for saving the aforementioned summary and presenting it when the user resumes browsing,
[0941] A system that includes this.
[0942] (Claim 2)
[0943] The system according to claim 1, wherein the summary is presented to the user by visual display and audio generation functions.
[0944] (Claim 3)
[0945] The system according to claim 1, which uses a specific method for extracting important information units when generating the summary.
[0946] "Application example 2 of combining emotional engines"
[0947] (Claim 1)
[0948] A means to record the last reading position and emotional state when a user interrupts reading an ebook,
[0949] A means for automatically summarizing the content of the ebook up to the relevant section and adjusting the presentation style based on the recorded information and the user's emotional state,
[0950] A means for saving the aforementioned summary and presenting it as an emotion-based summary when the user resumes reading,
[0951] A system that includes this.
[0952] (Claim 2)
[0953] The system according to claim 1, wherein the summary is presented to the user by text display and speech synthesis with an emotionally appropriate presentation style.
[0954] (Claim 3)
[0955] The system according to claim 1, wherein when generating the summary, it uses a specific algorithm for extracting key points and a generative AI model that takes sentiment information into account. [Explanation of Symbols]
[0956] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of recording the last reading position when a user interrupts reading an ebook, A means for automatically summarizing the content of an e-book up to the relevant section based on the recorded information, A means for saving the aforementioned summary and presenting it when the user resumes reading, A system that includes this.
2. The system according to claim 1, wherein the summary is presented to the user by text display and speech synthesis.
3. The system according to claim 1, wherein a specific algorithm is used to extract key points when generating the summary.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A